Detection of biological states through fragmentomic analysis

WO2026169862A1PCT designated stage Publication Date: 2026-08-13REALSEQ BIOSCIENCES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-08-13

Smart Images

  • Figure US2026014078_13082026_PF_FP_ABST
    Figure US2026014078_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are methods of detecting a biological state (such as a disease, infection and severity of disease or infection) utilizing nucleic acid fragments from a biological sample which sequencing reads are generated, a machine learning model may be used to detect the biological state from a quantitative measure of the sequencing reads.
Need to check novelty before this filing date? Find Prior Art

Description

WSGR Docket No. 57767-715.601DETECTION OF BIOLOGICAL STATES THROUGH FRAGMENTOMIC ANALYSIS CROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 755,129, filed February 6, 2025, which application is incorporated herein by reference in its entirety.STATEMENT AS TO FEDERALLY SPONSORED RESEARCH

[0002] This present disclosure was made with government support under R43HG013284 awarded by National Human Genome Research Institute (NHGRI / NIH). The government has certain rights in the disclosure.BACKGROUND

[0003] Cell-free DNA fragments (cfDNA) found in blood and other biofluids are promising biomarkers with diagnostic potential for cancer and other diverse pathologies. DNA fragmentomics is focused on studying the fragmentation of DNA molecules. However, current methods exploiting cfDNA-based biomarkers are not sensitive enough to avoid false positive or negative results in diagnosing cancer at early stages and / or monitoring minimal residual disease for the cancer recurrence.SUMMARY

[0004] In an aspect the present disclosure describes or provides a method of detecting a biological state comprising preparing a plurality of RNA fragments from a biological sample, sequencing the plurality of RNA fragments to obtain sequencing reads, aligning the sequencing reads to a nucleic acid sequence to obtain aligned sequencing reads, determining a quantitative measure of aligned sequencing reads that begin or end at a first genomic position and processing the quantitative measure of the aligned sequencing reads against a reference derived from a reference subject or using a trained machine learning model. In some embodiments, the present disclosure provides that the method further comprises adjusting the quantitative measure based at least in part on a count of sequencing reads that align to a range in a nucleic acid sequence. In some embodiments, the present disclosure provides that the quantitative measure is normalized to sequencing depth. In some embodiments, the present disclosure provides that the quantitative measure comprises a vector of values. In some embodiments, the present disclosure provides that the vector of values comprises one of; read coverage values, pseudocount values, or probability of coverage. In some embodiments, the present disclosure provides that the machine learning model comprises a clustering method. In some embodiments, the present disclosureWSGR Docket No. 57767-715.601provides that the clustering method is fuzzy c-means. In some embodiments, the present disclosure provides that the processing comprises detecting a signature. In some embodiments, the present disclosure provides that the processing produces an indication of a biological state. In some embodiments, the present disclosure provides that the biological state is a disease state. In some embodiments, the present disclosure provides that the disease state comprises at least one of; cancer, presence of a pathogen, organ failure, presence of an autoimmune disease, presence of precancerous lesions, presence of metastasis, presence of type 2 diabetes, inflammation, schizophrenia, Alzheimer’s, Lewy body dementia, asthma, infection with a pathogen, or any combination thereof. In some embodiments, the disease state comprises cancer. In some embodiments, the method further comprises administering a cancer therapeutic to a subject wherein the biological sample was obtained from the subject. In some embodiments, the disease state comprises presence of a pathogen or infection with a pathogen. In some embodiments, the pathogen comprises a fungal disease. In some embodiments, the method further comprises administering an anti-fungal agent to a subject, wherein the biological sample was obtained from the subject. In some embodiments, the present disclosure provides that the biological state comprises an indication of severity. In some embodiments, the present disclosure provides that the processing produces an indication of an affected region of the body. In some embodiments, the present disclosure provides that the affected region of the body comprises a tissue, an organ, or a biological system. In some embodiments, the present disclosure provides that the biological sample comprises one of venous blood, peripheral blood, stool, nasal swab, mucus, urine, blood, a blood fraction, plasma, serum, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid (CSF), peritoneal fluid, or a combination thereof. In some embodiments, the present disclosure provides that the preparing of the plurality of RNA fragments is untargeted. In some embodiments, the present disclosure provides that the preparing of the plurality of RNA fragments is targeted. In some embodiments, the present disclosure provides that the preparing of the plurality of RNA fragments comprises an enrichment of an RNA sequences of interest. In some embodiments, the present disclosure provides that the preparing of the plurality of RNA fragments comprises an enrichment of an RNA fragment having specific phosphorylation states of the ends. In some embodiments, the present disclosure provides that the preparing of the plurality of RNA fragments comprises a selection for an RNA sequences of interest. In some embodiments, the present disclosure provides that the preparing of the plurality of RNA fragments comprises a selection for an RNA fragment having specific phosphorylation states of the ends. In some embodiments, the present disclosure provides that the preparing of the plurality of RNA fragments comprises a selectionWSGR Docket No. 57767-715.601of an RNA fragment having specific phosphorylation states of the ends. In some embodiments, the present disclosure provides that the phosphorylation state is at the 5’ end of the RNA fragment or RNA sequence. In some embodiments, the present disclosure provides that the phosphorylation state of the 5’ end of the RNA fragment or RNA sequence is monophosphorylated. In some embodiments, the present disclosure provides that the phosphorylation state of the 5’ end of the RNA fragment or RNA sequence is diphosphorylated. In some embodiments, the present disclosure provides that the phosphorylation state of the 5’ end of the RNA fragment or RNA sequence is triphosphorylated. In some embodiments, the present disclosure provides that the phosphorylation state is at the 3’ end of the RNA fragment or RNA sequence. In some embodiments, the present disclosure provides that the phosphorylation state of the 3’ end of the RNA fragment or RNA sequence is 3 ’-phosphate (3’-P). In some embodiments, the present disclosure provides that the phosphorylation state of the 3’ end of the RNA fragment or RNA sequence is 2’, 3 ’-cyclic phosphate (cyc-P).

[0005] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.

[0006] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.

[0007] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE

[0008] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.WSGR Docket No. 57767-715.601BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0010] FIG. 1 shows an illustration of how RNA fragments arise.

[0011] FIG. 2 shows fragmentomic data being used for cluster analysis.

[0012] FIG. 3 shows an illustration of fragments being made from a healthy and cancerous cell.

[0013] FIG. 4 shows fragmentomic from two sources being used for cluster analysis.

[0014] FIG. 5 shows abundance data for different classes of samples.

[0015] FIG. 6 shows fragmentomic data, a consensus sequence and a cluster analysis from the same samples.

[0016] FIG. 7A shows an illustration of fragments being made from two fungal tissues.

[0017] FIG. 7B shows percent of transcripts that are morphology based.

[0018] FIG. 7C shows fragmentomic data, a consensus sequence and a cluster analysis from the same samples for the Met-CAU-1 transcript.

[0019] FIG. 7D shows fragmentomic data, a consensus sequence and a cluster analysis from the same samples for the U4 transcript.

[0020] FIG. 7E shows fragmentomic data, a consensus sequence, and a cluster analysis from the same samples for the sRNAlocus_4698 transcript.

[0021] FIG. 8 shows a computer system that is programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION

[0022] Cell-free nucleic acids (cfDNA and cfRNA) found in blood and other biofluids are promising biomarkers with diagnostic potential for cancer and other diverse pathologies.However, current methods exploiting cfDNA analytes are not sensitive enough to avoid false positive or negative results in diagnosing cancer at early stages (when it is more treatable) and monitoring minimal residual disease for recurrence when tumor-associated cfDNA are in several orders of magnitude lower abundance than background cfDNA that originated from non-cancerous cells. Cell-free RNA (cfRNA) imay be an important class of biomarkers for cancerWSGR Docket No. 57767-715.601and infectious disease detection. Because RNA is transcribed in multiple copies from the genomic and intergenic DNA templates (including the ones that are normally silent) and may contribute to higher tumor-associated RNA abundance than tumor-derived DNA both in cells and in circulation. cfRNA species may contain information relating to biological phenotypes that could provide information about tissues of cancer origin and cancer subtype specificity. Changes in the RNA expression profile in tumors, dysregulated RNA post-transcriptional events (including alternative splicing and formation of chimeric RNAs that are detectable in the transcriptome but not in the genome) contribute to a higher complexity of the cfRNA landscape. RNA biomarkers may be useful in the detection of many diverse pathologies such as, but not limited to cancer, microbial (viral, bacterial and fungi) infections and genetic disorders.

[0023] Approximately 95% of total cfRNA are small RNAs (sRNAs) and RNA fragments (RFs), which are shorter than 42 nucleotides (nt) in length. As shown in FIG. 1, sRNA and RFs may comprise products of cleavages of larger RNAs of various classes by intracellular and / or circulating ribonucleases which may yield a variety of RNA molecules that also differ in state of phosphorylation at their ends. Protective RNA secondary structures, RNA-protein complexes, or encapsulation into lipid EVs may aid the RFs released into circulation to survive further degradation. The protected sRNAs and RFs may be detected and analyzed by next-generation sequencing (NGS). The RNA fragmentome may include RFs cut from precursors and mature mRNAs, IncRNAs, tRNAs, rRNAs, snRNAs and other RNA classes. The small RFs derived from specific regions of these RNAs, as opposed to products of random RNA fragmentation, may be useful biomarkers.. Analysis of cfRNAs has been primarily focused on miRNA.However, there are a limited number of miRNA with tissue-specific expression patterns, and these miRNA represent a small variety of the transcriptome, whereas some other RNA classes (e.g., mRNA and IncRNA) have much greater diversity than miRNA. As a result, the potential to obtain biomarkers that reliably assess the state of a disease using sRNs and RFs derived from these more diverse and abundant RNAs is much higher.

[0024] Sequencing analysis of the RNA fragmentome (also known as the cfRNA transcriptome) may improve understanding of the roles of diverse sRNAs and RFs in many tasks, such as monitoring disease development and treatment resistance, discovering novel RNA biomarkers, improving disease diagnostics, prognostics, treatment monitoring, and detecting minimal residual disease re-occurrence. RNA fragmentomic data may be analyzed individually or in combination with other data such as DNA fragmentomic data. Multimodal nucleic acids analysis utilizing both cell-free RNA and DNA sequencing data may be analyzed simultaneously from the same biofluid sample (liquid biopsy) providing a more comprehensiveWSGR Docket No. 57767-715.601understanding of cellular processes and disease mechanisms. This approach would enable researchers to link genetic and DNA fragmentation variations with their functional consequences at the transcriptional and RNA fragmentation level.

[0025] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0026] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0027] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0028] Fragmentomic patterns may be used to infer health states of a patient. Many efforts to infer biological states from fragmentomics utilize cell-free DNA cfDNA. Circulating free DNA (cfDNA) is released into the bloodstream primarily through the breakdown of cells, including apoptotic, necrotic, lysed, or damaged cells. Due to cfDNA being derived from cells that have died, the resulting fragmentation patterns may carry signatures that are indicative of the cell death rather than a biological state. Such noise may be confounding or may mask relevant information to the biological state. As shown in FIG. 1, Cell-free RNA (cfRNA) may be derived from living cells, such as through RNA shedding, or dying cells. The combination of both living and dying cell sources means signatures of disease may be more easily distinguished from signatures of cell death, improving the ability to discovery relevant patterns of cfRNA fragmentation.

[0029] While promising, current technologies are limited by the amount of cfRNA they may see. Many methods for small RNA-Seq library preparation, which enable analysis of absolute majority of sRNAs and RFs found in biofluids, can capture only miRNA and some other RNA molecules having 5’-P and 3’-OH ends. However, these types of sRNAs and RFs account for less than 10% of the whole RNA fragmentome, while the remaining RNA typesWSGR Docket No. 57767-715.601(accounting for >90%) having different phosphorylation statuses of their termini are hidden and, therefore, non-detected. The hidden sRNAs may possess ends that may impede the ligation of sequencing adapters or polyadenylation of the 3’ end used in these protocols causing them to be unable to detect these molecules. Methods that are capable of detecting these molecules may provide a more complete picture of the fragmentomic patterns in a sample thus providing data with more information that may be utilized by a machine learning method such as those described herein.

[0030] Herein, machine learning is used to identify cfRNA fragmentomic patterns and utilize those patterns to detect a biological state.

[0031] Various aspects of the present disclosures provide methods and systems for detecting a biological state.

[0032] In some aspects, the present disclosure provides a method for detecting a biological state in a sample. The sample may be from an individual. In some embodiments, the method may comprise preparing a plurality of nucleic acids from a biological sample of an individual. Nucleic acids may be RNA or DNA. The plurality of nucleic acids may be fragments of RNA or DNA. The plurality of nucleic acids may be sequenced to obtain sequencing reads. The sequencing reads may be aligned to a nucleic acid sequence to obtain aligned sequencing reads. In some embodiments, the method comprises determining a quantitative measure of the aligned sequencing reads that begin or end at a first genomic position. In some embodiments, the method comprises processing the quantitative measure of the aligned reads against a reference derived from a reference subject. In some embodiments, the method comprises processing the quantitative measure of the aligned reads using a trained machine learning model. In some embodiments, the method comprises processing the quantitative measure of the aligned reads against a reference derived from a reference subject and using a trained machine learning model. In some embodiments the trained machine learning model detects a biological state. In some embodiments, the trained machine learning model predicts the biological state. In some embodiments, the trained machine learning model classifies the biological state.

[0033] In some embodiments, the method further comprises adjusting the quantitative measure. In some embodiments, the adjusting may be based at least in part on a count of sequencing reads that align to a range in the nucleic acid sequence.

[0034] In some embodiments, the quantitative measure is normalized to sequencing depth. Sequencing depth may be average sequencing depth, sequencing depth at a nucleic acid sequence, or sequencing depth at a single nucleotide.

[0035] In some embodiments, the quantitative measure may comprise a vector of values. InWSGR Docket No. 57767-715.601some embodiments the vector of values may comprise read coverage values, psuedocount values, probability of coverage, or any combination thereof.

[0036] In some embodiments, the machine learning model may comprise a clustering method.

[0037] In some embodiments, the clustering method may comprise fuzzy c-means.

[0038] In some embodiments, the processing may comprise detecting a signature. A signature may be a portion of a quantitative measure which is useful in distinguishing a biological state from other biological states. A signature may be learned by a machine learning model. A signature may be quantifiable through a comparison of the quantitative measure to a reference. A signature may be a pattern of the quantitative measure. A pattern may be a configuration or approximate configuration of the quantitative measure that repeats. A pattern may repeat within a sample. A pattern may repeat across samples, such as a configuration or approximate configuration of a portion of the quantitative measure which appears when a condition, such as a biological state, is present in a sample. A pattern may occur in a portion of the quantitative measure.

[0039] A pattern or signature may be a repeated or predictable arrangement of element, structures, or behaviors that emerge over time or across a dataset. A pattern may be a regularity in space, time, or data. A pattern may reveal underlying order or structure. A pattern may be simple, or complex.

[0040] In some embodiments, the processing may produce an indication of a biological state. A biological state may be indicated by the output of a machine learning model such as a classification, indication, or a probability. The indication may be detection. The output may be a continuous value. The output may be a discrete value. The machine learning model may output an indication of more than one biological state, such as a classifier with multiple binary outputs each corresponding to the presence or absence of a particular biological state or a clustering method where a sample is predicted as having membership in multiple clusters, each cluster indicative of a particular biological state. The machine learning model may indicate a single biological state out of many biological states, such as a machine learning model which outputs a vector of numbers each indicating the probability of a particular biological state which is configured to maximize one single number in the vector that will be the single biological state being predicted. The machine learning model may indicate the presence or absence of a biological state, such as in a binary classification method. The machine learning model may indicate the likelihood of a biological state, such as through a probability or logit output where a logit may be a real valued output of a machine learning model such as a neural network prior toWSGR Docket No. 57767-715.601a transformation such as a sigmoid or tanh.

[0041] The biological state may comprise a healthy state, cancer, presence of a pathogen, organ failure, autoimmune disease, precancerous lesions, metastasis, type 2 diabetes, inflammation, schizophrenia, Alzheimer’s disease, lewy body dementia, asthma, or any combination thereof.

[0042] The biological state may comprise an indication of the severity of a disease or condition. Severity may be indicated by a score or classification. A score may indicate severity through a continuous output within a range (such as 0 to 1) where a higher score indicates greater severity. A score may indicate severity through a classification when discrete severities exists (such as low, medium, and high) or where severity may be indicated by clinical classes (such as liver fibrosis classes 0 - 4).Nucleic acid

[0043] In some embodiments a nucleic acid comprises a deoxyribonucleotide, a ribonucleotide, a deoxyribonucleotide analog, chemically modified canonical deoxyribonucleotides, ribonucleotides, and / or ribonucleotide analog, nucleic acids with modified backbones, or any combination thereof. A nucleic acid may be a polynucleotide or an oligonucleotide.

[0044] In some embodiments, a nucleic acid may be cell-free. Cell-free may be the condition of the nucleic acid outside a cell, viral particle or virion as it appeared in the body immediately before the sample is obtained from the body. For example, circulating cell-free nucleic acids in a sample may have originated as cell -free nucleic acids circulating in the bloodstream of a subject. In contrast, nucleic acids that are extracted post-collection from an intact microorganism, such as a blood-borne pathogen, or removed post-collection from intact virions in a plasma sample, are generally not considered to be “cell-free.”

[0045] The nucleic acids may be analyzed to obtain various types of information including genomic, epigenetic (e.g., methylation), and RNA expression. Methylation analysis may be performed by, for example, conversion of methylated bases followed by DNA sequencing. RNA expression analysis may be performed, for example, by polynucleotide array hybridization, by RNA sequencing techniques, or by sequencing cDNA produced from RNA.

[0046] The following are non -limiting examples of nucleic acids: coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), long non coding RNA (IncWSGR Docket No. 57767-715.601RNA), small non coding RNAs such as but not restricted to piwi RNAs and enhancer RNAs, circ RNA (circular RNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, primers, mitochondrial DNA, circulating nucleic acids, cell-free nucleic acids, cfDNA, cfNA, host cfNA, non-host cfNA, circulating cfNA, fungal cell free nucleic acids, microbial cell free nucleic acids, viral nucleic acid, bacterial nucleic acid, genomic DNA, pathogen nucleic acids, fungal nucleic acid, parasitic nucleic acid, exosomal nucleic acid, intercellular signal nucleic acid, exogenous nucleic acids, nucleic acid therapeutics, and DNA enzymes. A nucleic acid may comprise one or more modified nucleotides, such as methylated nucleotides or methylated nucleotide analogs. If present, modifications to the structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A nucleic acid may be further modified after polymerization, such as by conjugation with a labeling component. A nucleic acid may be single-stranded, double-stranded, have higher numbers of strands (e.g., triple- stranded), and / or have a higher order of structure (e.g., tertiary, or quaternary structure). A target nucleic acid may be any type, category, or subcategory of nucleic acids.

[0047] In some embodiments a target polynucleotide comprises a nucleic acid molecule or polynucleotide in a starting population of nucleic acid molecules having a target sequence whose presence, amount, and / or nucleotide sequence, or changes in one or more of these, are desired to be determined. A target sequence may comprise a nucleic acid sequence on a single strand of nucleic acid. The target sequence may be a portion of a gene, a regulatory sequence, genomic DNA, cDNA, RNA including mRNA, miRNA, rRNA, or others. The target sequence may be a target sequence from a sample or a secondary target such as a product of an amplification reaction.

[0048] A plurality of RNA molecules can be prepared from biological samples in any way known in the art that preserves the nucleotide information and is compatible with sequencing library preparation. For example, methods described in U.S. Pat. No. 11,014,957, which is hereby incorporated in its entirety by reference, may be used to isolate and purify RNA sequences and prepare sequencing libraries therefrom.Targeted / Untargeted

[0049] In some embodiments, sequencing of the plurality of RNA fragments is untargeted. Untargeted sequencing may use high throughput sequencing that sequences nucleic acids without prior knowledge or bias towards specific genomic regions of interest. UntargetedWSGR Docket No. 57767-715.601sequencing may comprise the use of non-sequence specific sequencing methods such as the ligation of nonspecific oligonucleotides to a nucleic acid regardless of the nucleic acids sequence and subsequent sequencing and / or amplification of that nucleic acid for further downstream applications such as alignment, quantification or both. Such methods may be used in the processing of fragmented nucleic acids (such as cfRNA, or cfDNA fragments) as a nucleic acid may be fragmented in ways that interfere with the ability of a sequence specific primer to bind to its target sequence making those primers less likely to yield reliable results when a specific fragment is not well characterized as being useful to analysis. Untargeted analysis yields a large volume of sequencing data and may provide contextual information that is useful in some machine learning applications.

[0050] In some embodiments, preparing of the plurality of RNA fragments is targeted, wherein a set of nucleic acids arising from specified regions of interest are sequenced. Such method may use oligonucleotides designed to bind to a target sequence allowing for the subsequent amplification and / or quantification of the target sequence. Targeted sequencing may be used in fragmentomic analysis when a particular sequence or fragment is well characterized, and others may be of lesser value to the analysis.

[0051] In some embodiments a fragment or segment comprises a portion of nucleic acid. A polynucleotide, for example, may be broken up, or fragmented into, a plurality of segments, either through natural processes, as is the case with, e.g., cfDNA fragments that may naturally occur within a biological sample, or through in vitro manipulation.Subject / individual

[0052] A subject (e.g., a patient or individual) may be a human or non-human animal. The subject may be known to have, or potentially have, a medical condition, biological state, or disorder, such as, e.g., a cancer. A subject may be an individual. An individual may be a human individual or non-human individual. A subject may be an agricultural subject (e.g., a plant), an environmental subject (such as a soil sample, a water sample, an air sample, etc.)

[0053] A biological sample may be derived from any subject. The subject may be healthy. In some embodiments, the subject is a human patient having, suspected of having, or at risk of having, a disease or infection. In some embodiments, the disease or infection is pathogen related.

[0054] A human subject may be a male or female. In some embodiments, the sample may be from a human embryo or a human fetus. In some embodiments, the human may be an infant, child, teenager, adult, or elderly person. In some embodiments, the subject is a female subjectWSGR Docket No. 57767-715.601who is pregnant, suspected of being pregnant, or planning to become pregnant.

[0055] In some embodiments, the subject is a human subject who has undergone an organ transplant or who is planning to undergo organ transplant.

[0056] In some embodiments, the subject is a farm animal, a lab animal, a wild animal, or a domestic pet. In some embodiments, the animal may be an insect, a dog, a cat, a horse, a cow, a mouse, a rat, a pig, a fish, a bird, a chicken, or a monkey.

[0057] The subject may be an organism, such as a single-celled or multi-cellular organism. In some embodiments, the sample may be obtained from a plant, fungi, eubacteria, archeabacteria, protist, or any multicellular organism. The subject may be cultured cells, which may be primary cells or cells from an established cell line.

[0058] In some embodiments, the subject has a genetic disease or disorder, is affected by a genetic disease or disorder, or is at risk of having a genetic disease or disorder. A genetic disease or disorder may be linked to a genetic variation such as mutations, insertions, additions, deletions, translocations, point mutations, trinucleotide repeat disorders, single nucleotide polymorphisms (SNPs), or a combination of genetic variations. A genetic disease or disorder may be linked to an expression related disorder such as variations in abundance of a RNA isoform, alterations in alternative splicing, alterations in post-translational modifications of RNA, or any combination thereof.

[0059] In some aspects, the subject is healthy or asymptomatic, or exhibits mild or nonspecific clinical symptoms. In some cases, a subject may be infected or suspected of being infected by a particular pathogen. In other cases, the subject is suspected of having an infection of unknown origin. In some cases, the subject has been exposed to a pathogen, or suspected to have been exposed to a pathogen such as by living conditions, by travel to a particular geographic region or by interaction or sexual interaction with an infected individual.Biological sample

[0060] In some embodiments, the biological sample comprises a solid or a body fluid such as blood, plasma, serum, cerebrospinal fluid, synovial fluid, bronchoalveolar lavage, urine, stool, saliva, abdominal fluid, ascites fluid, peritoneal lavage, gastric fluid, interstitial fluid, lymph fluid, bile, abscess fluid, tissue, amniotic fluid, meconium, sinus aspirate, lymph node, bone marrow, hair, nails, cheek swab, skin swab, urethral swab, cervical swab, nasopharyngeal swab, nasopharyngeal aspirate, vaginal swab, epithelial cells, semen, vaginal discharge, intercellular fluid, pericardial fluid, rectal swab, bone, skin tissue, soft tissue, tears, and / or a nasal sample. In some embodiments, the biological sample comprises, consists of, or consists essentially ofWSGR Docket No. 57767-715.601plasma. In some embodiments, the biological sample is from a human subject. In some embodiments, the biological sample is from a non-human subject.

[0061] In some embodiments, a biological sample may be made up of, in whole or in part, cells and / or tissue. The biological sample may be cell-free or cell-depleted. The biological cell-free sample may comprise, consist of, or consist essentially of nucleic acids that originated from a different site in the body, such as a site of pathogenic infection. In the case of blood, serum, lymph, or plasma, the cell-free sample or cell-depleted biological sample may contain “circulating” cell-free nucleic acids that originated at anatomic locations other than the site of bodily fluid collection of the fluid in question. In the case of urine, the cell-free nucleic acids may be cell-free nucleic acids that originated in a different site in the body. The cell-free samples or cell-depleted biological samples can be obtained by depleting or removing cells, cell fragments, or exosomes, such as by centrifugation or filtration.

[0062] In some embodiments an invasive disease comprises a disease based, in part, on the ability of particular pathogens to compromise the health of infected subjects, as opposed to merely colonizing other infected subjects, either as a commensal or infection with no or minor symptoms. For example, certain microbes can locally colonize tissues without causing any health problems in some hosts, while, in other hosts, they may invade tissues to the point where they cause serious inflammation, tissue or organ damage, sepsis, cancer, and other serious health issues. Microbes may also colonize a subject who is asymptomatic at one time point, but at a later point develops serious symptoms when the microbe translocates and / or becomes “active.”

[0063] The biological sample can be from a subject who has a specific disease, condition, or infection, or is suspected of having (or at risk of having) a specific disease, condition, or infection. For example, the biological sample can be from a cancer patient, a patient suspected of having cancer or a patient at risk of having cancer. In some embodiments, the biological sample can be from a patient with an infection, a patient suspected of an infection, or a patient at risk of having an infection. In some embodiments, the biological sample is from a subject who has undergone, or will undergo, an organ transplant.

[0064] In some embodiments, a biological sample comprises circulating tumor or fetal nucleic acids. In some embodiments, the biological sample comprises, consists of, or consists essentially of circulating donor nucleic acids.Biological state

[0065] A biological state may be a medical condition of the subject. A biological state may include conditions such as, but not limited to, a healthy state, an infection, a disease state, a severity of a disease state, the likelihood of a disease state, A biological state may be a specificWSGR Docket No. 57767-715.601or configuration of an organism (such as the subject) a given point in time. A biological state may comprise various physiological, biochemical, and molecular characteristics or configurations. A biological state may describe the status of a single cell, an organ, a tissue, or an entire organism, and may reflect the organism’s development, response to environmental stimuli, health, or disease condition. A biological state may be dynamic and may change in response to internal or external factors such as gene expression, nutrient availability, stress, hormonal changes, or pathogens.

[0066] The disease state may comprise at least one of cancer, presence of a pathogen, organ failure, presence of an autoimmune disease, presence of precancerous lesions, presence of metastasis, presence of type 2 diabetes, inflammation, schizophrenia, Alzheimer’s, Lewy body dementia, asthma, infection with a pathogen, or any combination thereof. In some embodiments, the disease state comprises cancer. In some embodiments, the disease state comprises presence of a pathogen. In some embodiments, the disease state comprises organ failure. In some embodiments, the disease state comprises presence of an autoimmune disease. In some embodiments, the disease state comprises presence of precancerous lesions. In some embodiments, the disease state comprises presence of metastasis. In some embodiments, the disease state comprises presence of type 2 diabetes. In some embodiments, the disease state comprises schizophrenia. In some embodiments, the disease state comprises Alzheimer’s. In some embodiments, the disease state comprises Lewy body dementia. In some embodiments, the disease state comprises asthma. In some embodiments, the disease state comprises infection with a pathogen.

[0067] A biological state may be measured at multiple levels, such as molecular, cellular, tissue, organ, biological system (such as integumentary, skeletal, muscular, nervous, endocrine, cardiovascular, lymphatic (immune), respiratory, digestive, urinary, or reproductive systems), and / or whole-body levels.

[0068] In some embodiments a biological state may cause alterations in cfRNA fragments due to alteration in splicing, protein binding or tertiary structure of an RNA which may alter fragment length and / or distribution of fragments over a nucleic acid sequence, such as a transcript.Treatment

[0069] In some embodiments, the method further comprises a step of treating following identification of a disease state. The method may comprise administering a subject a therapeutic, wherein the biological sample was obtained from the subject.WSGR Docket No. 57767-715.601

[0070] In some embodiments, the disease state is cancer, and the method comprises administering a cancer therapeutic. In some embodiments, the cancer therapeutic comprises a chemotherapeutic agent. In some embodiments, the chemotherapeutic agent comprises cyclophosphamide, chlorambucil, melphalan, busulfan, 5 -fluorouracil (5-FU), methotrexate, gemcitabine, cytarabine, doxorubicin, daunorubicin, epirubicin, idarubicin, paclitaxel, docetaxel, , vincristine, vinblastine, vinorelbine, topotecan, irinotecan, imatinib (Gleevec), erlotinib (Tarceva), lapatinib (Tykerb),, olaparib (Lynparza), niraparib (Zejula), or a combination thereof. In some embodiments, the anti-cancer agent comprises pembrolizumab (Keytruda®), nivolumab (Opdivo®), atezolizumab (Tecentriq®), durvalumab (Imfinzi®), ipilimumab (Yervoy®), tremelimumab (Imjudo®), rituximab (Rituxan® / Mab Thera®), daratumumab (Darzalex®), brentuximab vedotin (Adcetris®), alemtuzumab (Campath®), obinutuzumab (Gazyva®), elotuzumab (Empliciti®), isatuximab (Sarclisa®), gemtuzumab ozogamicin (Mylotarg®), trastuzumab (Herceptin®), bevacizumab (Avastin®), cetuximab (Erbitux®), pertuzumab (Perjeta®), ramucirumab (Cyramza®), panitumumab (Vectibix®), ado-trastuzumab emtansine (Kadcyla®), dinutuximab (Unituxin®), denosumab (Xgeva®), teclistamab (Tecvayli®), mosunetuzumab (Lunsumio®), epcoritamab (Epkinly®), glofitamab (Columvi®), talquetamab (Talvey®), or a combination thereof.

[0071] In some embodiments, the disease state comprises presence of a pathogen or infection with a pathogen. In some embodiments, the pathogen comprises a fungus. In some embodiments, the method comprises administering an anti-fungal agent. In some embodiments, the anti-fungal agent comprises clotrimazole, exonazole, terbinafine, fluconazole, ketoconazole, nystatin, anidulafungin, caspofungin, micafungin, rezafungin, or amphotericin. In some embodiments, the anti-fungal agent comprises an amphotericin B compound (AmB). In some embodiments the AmB comprises intrathecal AmBd or L-AmB. In some embodiments, the antifungal agent comprises an azole. In some embodiments, the azole comprises fluconazole, itraconazole, SUBA-itraconazole, voriconazole, posaconazole, isavuconazonium, or a combination thereof. In some embodiments, the anti-fungal agent is olorofim.Region of the body

[0072] In some embodiments, the processing comprises producing an indication of an affected region of the body. An affected region of the body may be indicated by the output of a machine learning model such as a classification, indication, or a probability. The indication may be detection. The output may be a continuous value. The output may be a discrete value. The machine learning model may output an indication of more than one region of the body, such as aWSGR Docket No. 57767-715.601classifier with multiple binary outputs each corresponding to the different regions of the body or a clustering method where a sample is predicted as having membership in multiple clusters, each cluster indicative of a particular region of the body. The machine learning model may indicate a single region of the body out of many regions of the body, such as a machine learning model which outputs a vector of numbers each indicating the probability of a particular region of the body which is configured to maximize one single number in the vector that will be the single region of the body being predicted. The machine learning model may indicate the presence or absence of a biological state, such as in a binary classification method. The machine learning model may indicate the likelihood of a region of the body, such as through a probability or logit output.

[0073] In some embodiments, the affected region of the body may be a tissue, an organ or a biological system.

[0074] In some embodiments, the biological sample may comprise venous blood, peripheral blood, stool, nasal swab, mucus, urine, saliva, pleural fluid, cerebrospinal fluid (CSF), peritoneal fluid, or a combination thereof.

[0075] The present application provides methods of determining a site of localization in a subject. A site of localization is a region of the body. Nucleic acids from microbes or microorganisms from different sites within a subject may exhibit different fragment length profiles. The fragment length profile of a nucleic acid library or a subset of the nucleic acid library containing microbial nucleic acids differs if the microbial infection is circulating rather than located at one or more sites of localization. Thus, comparing a fragment length profile to a reference fragment length profile of one or more source sites may predict a site of localization if the fragment length profile from the sample is similar to a reference fragment length profile from a source site.

[0076] A site of localization may be any source site within a subject wherea microbe occurs, persists, survives or proliferates. Source sites include, but are not limited to the bloodstream, blood, deep tissue, such as but not limited to the kidneys, liver, stomach, bladder, digestive organs, nerve cells, lung, bone, brain, heart, heart lining, sinus, GI tract, spleen, skin, joint, ear, nose, and mouth. It is envisioned that a subject may have more than one site of localization for a particular microbe. It is further understood that some sites of localization for a particular microbe may not contribute to a disease state or condition. Rather, some sites of localization for a particular microbe may indicate a commensal relationship between the microbe and host, while other sites of localization for a particular microbe may indicate a parasitic or amensal relationship between the microbe and host. It is furtherWSGR Docket No. 57767-715.601recognized that the occurrence of multiple sites of localization for a particular microbe may indicate a systemic infection of the host. Additionally, it is recognized that site of localization for a particular microbe or pathogen of interest may impact a decision to treat or not to treat and may impact selection of appropriate treatment options. For example, and without being limited by mechanism, a fungal pathogen localized to the skin may be treated differently than a fungal pathogen localized to the lung and a bacterial microbe localized to heart tissue including but not limited to the lining of the hear may be treated differently than a bacterial microbe localized to the blood or blood stream.

[0077] The present application provides methods for determining tissue of origin (TOO). Tissue of origin is a region of the body. Tissue of origin may be the organ, organ group, body region or cell type that a nucleotide sequence originates from. The nucleotide sequence may indicate a disease state, such as cancer. The tissue of origin may be a microbe. The tissue of origin may be a fungal tissue. The identification of a tissue of origin allows for identification of the most appropriate next steps in the care continuum of cancer to further diagnose, stage, severity, prognosis, and decide on treatment.Quantitative measure

[0078] The quantitative measure may comprise a size distribution of sequencing reads which align to a nucleic acid sequence. The quantitative measurement may further comprise a size distribution ratio of sequencing reads which align to a nucleic acid sequence. The data set may further comprise values that indicate, for a plurality of nucleic acid sequences, the location in an intron or exon. The quantitative measure may comprise one of alignment score, identity percentage, mismatch count, gap count, read depth, base quality scores, mapping quality scores, reference allele frequency, variant allele frequency, base composition, insert size, sequence depth, heterozygosity, allele balance, sequencing error rate, read pairs, mapping distance, variant call, coverage uniformity, or any combination thereof.

[0079] A pseudocount may be a small, arbitrary value added to data to prevent the occurrence of zero values that may cause undefined or biased results in some model types. A pseudocount may provide more stable and / or reliable estimates, particularly in situations where certain observations are rare or absent in the dataset, such as in sequence alignment or frequency estimation.

[0080] Ranges of the nucleic acid sequence that are not present in the biological sample may be imputed. Imputation may comprise mean imputation, median imputation, mode imputation, regression imputation, k-nearest neighbors (KNN) imputation, multiple imputation,WSGR Docket No. 57767-715.601interpolation, extrapolation, expectation-maximization (EM), random forests imputation, decision tree imputation, deep learning imputation, Bayesian imputation, knn-based smoothing, hot-deck imputation, cold-deck imputation, principal component analysis (PCA) imputation, imputation by similarity, zero imputation, constant value imputation. A probability of coverage may be an estimate of coverage for a range of the nucleotide sequence which is not present in the biological sample.

[0081] In some embodiments, the quantitative measure may comprise a vector of values. The vector of values may comprise channels where each channel contains a different type of value such as read coverage values, pseudocount values, or probability of coverage. The different channels may be dimensions such that the vector of values is an N-dimensional vector where n is the number of channels. The vector of values may have a length where each position in the length of the vector comprises the channels. For example, image data may be a 3-dimensional vector of length K vector where an individual position, Ki, in the vector comprises 3 channels (or dimensions) such as red, green blue. In a similar manner, the vector of values may be of length K with a dimensionality of N, where each position, Ki, is a value of the quantitative measure and each dimension, Nj, is a channel, the machine learning model may be configured to take as input an N dimensional vector of length K so as to process the entire vector as a single input or may be configured to take each channel of a vector of length k as individual inputs or to take as input each position of the length K for all channels N or take as input each position of the length K for a single channel N of the vector.Microbe

[0082] In some embodiments a microbe (i.e., microbial or microorganism) comprises an organism, such as, a microscopic or macroscopic organism, which may exist as a single cell or as a colony of cells, capsids, spores, filaments, or multicellular organisms. Microbes include all unicellular organisms and some multicellular organisms, such as, for example, those from archaea, bacteria, protozoa, fungi, nematodes, viruses and eukaryotes. Microbes may be pathogens responsible for disease, or may exist in a non-pathogenic, symbiotic relationship with a host, such as the subject. A commensal microorganism may be microbes that exist in a non-pathogenic, symbiotic relationship with a host. A host organism may harbor multiple types of non-host organisms simultaneously. In co-infection a host organism harbors multiple types of non-host organisms. The multiple types of non-host organisms may include one or more pathogens, one or more commensal microorganisms, or at least one pathogen and at least one commensal microorganism. The methods of the current application may be used to distinguishWSGR Docket No. 57767-715.601between closely related microorganisms, distinguish between microbes present as a pathogen, a commensal microorganism, as incidental but clinically unimportant microbes, morphological features of a microbe, or tissues of a microbe.

[0083] Microbes or pathogens may include archaea, bacteria, yeast, fungi, molds, protozoans, nematodes, eukaryotes, and / or viruses. Microbes or pathogens may also include DNA viruses, RNA viruses, culturable bacteria, additional fastidious and unculturable bacteria, mycobacteria, and eukaryotic pathogens (See, Bennett J. E., D., R., Blaser, M. J. Mandell, Douglas, and Bennett's Principles and Practice of Infectious Diseases; Saunders, Philadelphia, Pa., 2014; and Netter's Infectious Disease, 1st Edition, edited by Elaine C. Jong, M D and Dennis L. Stevens, M D, PhD (2015)). Microbes or pathogens may also include any of the microbes set forth in https: / / www.ncbi.nlm.nih.gov / genome / microbes / or https: / / www.ncbi.nlm.nih.gov / biosample / .

[0084] Examples of microbes are one or more of the species or strains from one or more of the following genera: Coniosporium, Hantavirus, Talaromyces, Machlomovirus, Betatetravirus, Raoultella, Aeromonas, Ephemerovirus, Empedobacter, Loa, Macluravirus, Stenotrophomonas, Alfamovirus, Rosavirus, Emmonsia, Aggregatibacter, Orthopneumovirus, Weeksella, Nairovirus, Salivirus, Weissella, Mosavirus, Gammapar titivirus, Strongyloides, Passerivirus, Erysipelatoclostridium, Bacillarnavirus, lotatorquevirus, Taenia, Trypanosoma, Olsenella, Cladosporium, Rhizobium, Prevotella, Leclercia, Paracoccus, liarvirus, Lagovirus, Rasamsonia, Plasmodium, Acremonium, Chlamydia, Clonorchis, Vibrio, Bartonella, Nakazawaea, Franconibacter, Anisakis, Norovirus, Nocardia, Solobacterium, Parechovirus, Avenavirus, Orthohepevirus, Aphthovirus, Hepandensovirus, Microbacterium, Lichtheimia, Lomentospora, Achromobacter, Ipomovirus, Tsukamurella, Elizabethkingia, Hepevirus, Seadornavirus, Alternaria, Trueperella, Gammatorquevirus, Bifidobacterium, Chrysosporium, Thogotovirus, Curtovirus, Deltatorquevirus, Balamuthia, Mastrevirus, Bdellomicrovirus, Mupapillomavirus, Pseudozyma, Wicker hamiella, Aquamavirus, Alloscardovia, Thielavia, Idaeovirus, Henipavirus, Coxiella, Haemophilus, Gammacoronavirus, Negevirus, Brevibacterium, Peptoniphilus, Alphacarmotetravirus, Nosema, Trichovirus, Arenavirus, Thermomyces, Necator, Waikavirus, Blosnavirus, Jonesia, Tetraparvovirus, Emaravirus, Plectrovirus, Sclerodarnavirus, Toxocara, Umbravirus, Burkholderia, Chromobacterium, Paracoccidioides, Brugia, Eragrovirus, Macrococcus, Absidia, Colletotrichum, Inovirus, Phycomyces, Wickerhamomyces, Acidaminococcus, Moraxella, Rothia, Phlebovirus, Slackia, Purpureocillium, Betapapillomavirus, Tupavirus, Cryspovirus, Saksenaea, Erysipelothrix, Kobuvirus, Mimoreovirus, Echinococcus, Mannheimia, Bergeyella, Cyclospora, Xylanimonas,WSGR Docket No. 57767-715.601Leptospira, Finegoldia, Curvularia, Cryptosporidium, Babuvirus, Pecluvirus, Lambdatorquevirus, Pythium, Carlavirus, Entomobirnavirus, Kocuria, Anaplasma, Ampelovirus, Avihepatovirus, Nepovirus, Rhodococcus, Bordetella, Mischivirus, Scedosporium, Gardnerella, Maculavirus, Trichoderma, Aveparvovirus, Salmonella, Avastrovirus, Copiparvovirus, Trachipleistophora, Clostridioides, Nanovirus, Siccibacter, Leptotrichia, Citrivirus, Odoribacter, Sanguibacter, Novirhabdovirus, Acremonium, Hafnia, Chaetomium, Tenuivirus, Yokenella, Rubulavirus, Varicellovirus, Alphamesonivirus, Sicinivirus, Leuconostoc, Microvirus, Gallantivirus, Morbillivirus, Lolavirus, Pantoea, Hepatovirus, Nupapillomavirus, Metschnikowia, Barnavirus, Kytococcus, Tritimovirus, Tannerella, Respirovirus, Pneumocystis, Dirofdaria, Pediococcus, Lactococcus, Blastomyces, Dianthovirus, Actinobacillus, Teschovirus, Oscivirus, Begomovirus, Potyvirus, Byssochlamys, Alphacoronavirus, Molluscipoxvirus, Lymphocryptovirus, Sapelovirus, Parabacteroides, Pyrenochaeta, Listeria, Senecavirus, Brevidensovirus, Potexvirus, Parvimonas, Flavivirus, Recovirus, Toxoplasma, Yatapoxvirus, Opisthorchis, Trichuris, Cyphellophora, Morganella, Perhabdovirus, Micrococcus, Pequenovirus, Mastadenovirus, Anaeroglobus, Tropheryma, Dolosigranulum, Wolbachia, Lelliottia, Mycoplasma, Tobravirus, Shewanella, Paeniclostridium, Erythroparvovirus, Sutterella, Sporopachydermia, Narnavirus, Nyavirus, Francisella, Arthroderma, Epsilontorquevirus, Sigmavirus, Amdoparvovirus, Actinomyces, Alphapermutotetravirus, Cardiobacterium, Influenzavirus C, Orthopoxvirus, Poacevirus, Phialophora, Lactobacillus, Polyomavirus, Debaryomyces, Foveavirus, Bymovirus, Mycoflexivirus, Grimontia, Mucor, Rhytidhysteron, Quadrivirus, Thermoascus, Aureusvirus, Trichosporon, Myceliophthora, Dermacoccus, Dysgonomonas, Pseudoramibacter, Becurtovirus, Gordonia, Sapovirus, Orthobunyavirus, Spiromicrovirus, Pomovirus, Exophiala, Sneathia, Helicobacter, Photorhabdus, Mogibacterium, Betapartitivirus, Avibirnavirus, Ambidensovirus, Oleavirus, Orientia, Deltacoronavirus, Anulavirus, Trichomonasvirus, Budvicia, Geotrichum, Enamovirus, Lachnoclostridium, Schistosoma, Paecilomyces, Panicovirus, Rhizoctonia, Brevibacillus, Beauveria, Pestivirus, Tombusvirus, Cilevirus, Cokeromyces, Peptostreptococcus, Phanerochaete, Proteus, Idnoreovirus, Aspergillus, Pasteurella, Malassezia, Hanseniaspora, Endor navirus, Azospirillum, Velar ivir us, Cystovirus, Avisivirus, Bacteroides, Picobirnavirus, Myroides, Circovirus, Arterivirus, Aquaparamyxovirus, Onchocerca, Cosavirus, Kluyveromyces, Fijivirus, Candida, Hepacivirus, Dermabacter, Ourmiavirus, Allexivirus, Enterobacter, Acidovorax, Bracorhabdovirus, Carmovirus, Pluralibacter, Coltivirus, Fonsecaea, Streptobacillus, Corynebacterium, Macrophomina, Marburgvirus, Comovirus, Fabavirus, Alphanodavirus, Cellulomonas, Enter obius, Catabacter, Moellerella, Nakaseomyces,WSGR Docket No. 57767-715.601Cucumovirus, Valsa, Deltapartitivirus, Plesiomonas, Pseudomonas, Torovirus, Cuevavirus, Hypovirus, Trichomonas, Influenzavirus D, Giardiavirus, Crinivirus, Tepovirus, Sakobuvirus, Cyberlindnera, Paenalcaligenes, Bafmivirus, Rymovirus, Pegivirus, Yarrow ia, Treponema, Borreliella, Rubivirus, Aureobasidium, Angiostrongylus, Filobasidium, Photobacterium, Rhizopus, Orthoreovirus, Ustilago, Simplexvirus, Aquareovirus, Protoparvovirus, Propionibacterium, Sprivivirus, Hunnivirus, Apophy somyces, Meyerozyma, Alphapapillomavirus, Candida, Brucella, Gallivirus, Dinovernavirus, Anaerobiospirillum, Eubacterium, Tatlockia, Terri sporobacter, Quaranjavirus, Sobemovirus, Dicipivirus, Arcanobacterium, Macanavirus, Atopobium, Vesivirus, Lodderomyces, Dinornavirus, Betatorquevirus, Kerstersia, Aparavirus, Neisseria, Agrobacterium, Edwardsiella, Labyrnavirus, Totivirus, Actinomadura, Tobamovirus, Influenzavirus B, Mandarivirus, Anaerococcus, Kunsagivirus, Naegleria, Campylobacter, Veillonella, Yamadazyma, Filobasidiella, Oerskovia, Penicillium, Anncaliia, Leptosphaeria, Pneumovirus, Psychrobacter, Isavirus, Granulicatella, Torradovirus, Cladophialophora, Influenzavirus A, Ophiostoma, Aerococcus, Ureaplasma, Etatorquevirus, Bocaparvovirus, Megasphaera, Reptarenavirus, Comamonas, Capnocytophaga, Alphatorquevirus, Syncephalastrum, Wallemia, Betacoronavirus, Hyphopichia, Nocardiopsis, Legionella, Trichinella, Paraburkholderia, Mammarenavirus, Echinostoma, Sphingobacterium, Enterovirus, Methanobrevibacter, Ochroconis, Cheravirus, Pasivirus, Enterococcus, Mycoreovirus, Tospovirus, Betanodavirus, Phytoreovirus, Enterocytozoon, Ferlavirus, Stemphylium, Filifactor, Leishmaniavirus, Gemella, Bromovirus, Alloiococcus, Cunninghamella, Cronobacter, Oribacterium, Orbivirus, Chrysovirus, Cripavirus, Tatumella, Pandoraea, Ogataea, Dracunculus, Volvariella, Iflavirus, Benyvirus, Rhadinovirus, Histoplasma, Rahnella, Morococcus, Verticillium, Janibacter, Gyrovirus, Alphapartitivirus, Mycobacterium, Roseomonas, Varicosavirus, Chryseobacterium, Parapoxvirus, Rhizomucor, Aureimonas, Levivirus, Leishmania, Luteovirus, Cypovirus, Ochrobactrum, Microsporum, Piscihepevirus, Ceratocystis, Sporothrix, Vesiculovirus, Cupriavidus, Cryptococcus, Metapneumovirus, Alphanecrovirus, Eikenella, Brevundimonas, Escherichia, Leifsonia, Schizophyllum, Granulibacter, Gordonibacter, Lachancea, Madurella, Ophiovirus, Phellinus, Nebovirus, Acanthamoeba, Fusobacterium, Pichia, Verruconis, Ehrlichia, Tibrovirus, Higrevirus, Wohlfahrtiimonas, Rhinocladiella, Neorickettsia, Sadwavirus, Roseobacter, Sequivirus, Pannonibacter, Rotavirus, Turicella, Cardiovirus, Propionimicrobium, Furovirus, Naumovozyma, Closterovirus, Fluoribacter, Zeavirus, Clavispora, Megrivirus, Gammapapillomavirus, Rickettsia, Polemovirus, Corynespora, Encephalitozoon, Shimwellia, Fusarium, Yersinia, Capronia, Delftia, Victorivirus, Marafivirus, Kluyvera, Iteradensovirus,WSGR Docket No. 57767-715.601Isoptericola, Vitivirus, Roseolovirus, Conidiobolus, Abiotrophia, Babesia, Phoma, Sanguibacteroides, Staphylococcus, Rhodotorula, Zetatorquevirus, Hymenolepis, Fasciola, Cytorhabdovirus, Cardoreovirus, Memnoniella, Trichophyton, Mitovirus, Phaeoacremonium, Providencia, Lysinibacillus, Giardia, Oligella, Streptomyces, Paraclostridium, Ralstonia, Coccidioides, Brambyvirus, Biatriospora, Allolevivirus, Acinetobacter, Starmerella, Omegatetravirus, Porphyromonas, Avulavirus, Streptococcus, Arcobacter, Topocuvirus, Mamastrovirus, Ancylostoma, Bornavirus, Capillovirus, Alphavirus, Tymovirus, Nucleorhabdovirus, Diaporthe, Chlamydiamicrovirus, Turneurtovirus, Saccharomyces, Riemerella, Betanecrovirus, Clostridium, Mobiluncus, Cercospora, Marnavirus, Mortierella, Aquabirnavirus, Xanthomonas, Dependoparvovirus, Ebolavirus, Neofusicoccum, Borrelia, Leminorella, Klebsiella, Blastocystis, Alcaligenes, Citrobacter, Eggerthella, Cedecea, Serratia, Penstyldensovirus, Bacillus, Laribacter, Wuchereria, Hordeivirus, Cytomegalovirus, Actinomucor, Ascaris, Shigella, Vittaforma, Torulaspora, Kingella, Oryzavirus, Polerovirus, Tremovirus, Erbovirus, Entamoeba, Lyssavirus, Paenibacillus, Facklamia, Kappatorquevirus, Metarhizium, Stachybotrys, Okavirus, Botrexvirus, Thetatorquevirus, and Basidiobolus.Sequencing

[0085] Sequencing methods include, but are not limited to, Maxam-Gilbert sequencingbased techniques, chain-termination-based techniques, shotgun sequencing, bridge PCR sequencing, single-molecule real-time sequencing, ion semiconductor sequencing (e.g., Ion Torrent sequencing), nanopore sequencing, pyrosequencing (454), sequencing by synthesis, sequencing by ligation (SOLiD sequencing), sequencing by electron microscopy, dideoxy sequencing reactions (Sanger method), massively parallel sequencing, polony sequencing, and DNA nanoball sequencing. The term “Next Generation Sequencing (NGS)” herein refers to sequencing methods that allow for massively parallel sequencing of nucleic acid molecules during which a plurality, e.g., millions, of nucleic acid fragments from a single sample or from multiple different samples are sequenced simultaneously. Non-limiting examples of NGS include sequencing-by-synthesis, sequencing-by-ligation, real-time sequencing, and nanopore sequencing. In some embodiments, sequencing involves hybridizing a primer to the template to form a template / primer duplex, contacting the duplex with a polymerase enzyme in the presence of detectably labeled or unlabeled nucleotides under conditions that permit the polymerase to add labeled or unlabeled nucleotides to the primer in a template-dependent manner, detecting a signal from the incorporated labeled nucleotide or detecting a signal resulting from the process of incorporating labeled or unlabeled nucleotide (e.g., proton release), and sequentially repeatingWSGR Docket No. 57767-715.601the contacting and / or detecting steps at least once, wherein sequential detection of incorporated labeled or unlabeled nucleotide determines the sequence of the nucleic acid.

[0086] Exemplary detectable labels include radiolabels, fluorescent labels, protein labels, dye labels, enzymatic labels, etc. In some embodiments, the detectable label may be an optically detectable label, such as a fluorescent label. Exemplary fluorescent labels include cyanine, rhodamine, fluorescein, coumarin, BODIPY, alexa, or conjugated multi -dyes.

[0087] In some embodiments, the sequencing comprises, consists of, or consists essentially of obtaining paired end reads. In some embodiments, the sequencing comprises, consists of, or consists essentially of obtaining consensus reads.

[0088] The accuracy or average accuracy of the sequence information may be greater than about 80%, about 90%, about 95%, about 99%, about 99.98%, or about 99.99%. The sequence accuracy or average accuracy may be greater than about 95% or about 99%. The sequence coverage may be greater than about 0.00001 fold, 0.0001 fold, 0.001 fold, about 0.01 fold, about 0.1 fold, about 0.5 fold, about 0.7 fold, or about 0.9 fold. The sequence coverage may be less than about 200,000 fold, about 100,000 fold, about 10,000 fold, about 1,000 fold, or about 500 fold.

[0089] In some embodiments, the sequence information obtained per nucleic acid template is more than about 10 base pairs, about 15 base pairs, about 20 base pairs, about 50 base pairs, about 100 base pairs, or about 200 base pairs. The sequence information may be obtained in less than 1 month, 2 weeks, 1 week, 2 days, 1 day, 14 hours, 10 hours, 3 hours, 1 hour, 30 minutes, 10 minutes, or 5 minutes.

[0090] A consensus sequence may represent a common or typical sequence, signature or pattern found across a set of similar sequences. A consensus sequence may be derived from multiple aligned sequencing reads.

[0091] In some embodiments a sequence read comprises a string of nucleotides from part of, or all of, a nucleic acid molecule from a sample obtained from a subject. A sequence read may be a short string of nucleotides (e.g., 10-150) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained through various methods. For example, a sequence read may be obtained in a variety of ways, e.g., using sequencing techniques or using probes, e.g., in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.

[0092] In some embodiments a read segment (i.e., read) comprises any nucleotideWSGR Docket No. 57767-715.601sequences, including sequence reads obtained from a subject and / or nucleotide sequences, derived from a biological sequence read from a sample. For example, a read segment may refer to an aligned sequence read, a collapsed sequence read, or a stitched read. Furthermore, a read segment may refer to an individual nucleotide base, such as a single nucleotide variant.

[0093] In some embodiments, sequencing depth comprises the count of the number of times a given target nucleic acid within a sample has been sequenced (e.g., the count of sequence reads at a given target region). Increasing sequencing depth may reduce required amounts of nucleic acids required to assess a disease state (e.g., cancer or cancer tissue of origin).Targeted / Untargeted

[0094] In some embodiments, sequencing of the plurality of RNA fragments is untargeted. Untargeted sequencing may use high throughput sequencing that sequences nucleic acids without prior knowledge or bias towards specific genomic regions of interest. Untargeted sequencing may comprise the use of non-sequence specific sequencing methods such as the ligation of nonspecific oligonucleotides to a nucleic acid regardless of the nucleic acids sequence and subsequent sequencing and / or amplification of that nucleic acid for further downstream applications such as alignment, quantification or both. Such methods may be used in the processing of fragmented nucleic acids (such as cfRNA, or cfDNA fragments) as a nucleic acid may be fragmented in ways that interfere with the ability of a sequence specific primer to bind to its target sequence making those primers less likely to yield reliable results when a specific fragment is not well characterized as being useful to analysis. Untargeted analysis yields a large volume of sequencing data and may provide contextual information that is useful in some machine learning applications.

[0095] In some embodiments, sequencing of the plurality of RNA fragments is targeted, wherein a set of nucleic acids arising from specified regions of interest are sequenced. Such method may use oligonucleotides designed to bind to a target sequence allowing for the subsequent amplification and / or quantification of the target sequence. Targeted sequencing may be used in fragmentomic analysis when a particular sequence or fragment is well characterized, and others may be of lesser value to the analysis.

[0096] In some embodiments, the preparing of the plurality of RNA fragments comprises an enrichment an RNA sequence of interest. In some embodiments, the preparing of the plurality of RNA fragments comprises an enrichment an RNA fragment having specific phosphorylation states of the ends. In some embodiments, the preparing of the plurality of RNA fragments comprises a selection of an RNA sequences of interest. In some embodiments, the preparing ofWSGR Docket No. 57767-715.601the plurality of RNA fragments comprises a selection of an RNA fragment having specific phosphorylation states of the ends. The phosphorylation state may be at the 5’ end of the RNA fragment or RNA sequence. The phosphorylation state of the 5’ end of the RNA fragment or RNA sequence may be monophosphorylated. The phosphorylation state of the 5’ end of the RNA fragment or RNA sequence may be diphosphorylated. The phosphorylation state of the 5’ end of the RNA fragment or RNA sequence may be triphosphorylated. The phosphorylation state may be at the 3’ end of the RNA fragment or RNA sequence. The phosphorylation state of the 3’ end of the RNA fragment or RNA sequence may be 3 ’-phosphate (3’-P). The phosphorylation state of the 3’ end of the RNA fragment or RNA sequence may be 2’, 3 ’-cyclic phosphate (cyc-P).Machine learning

[0097] A machine learning model, as disclosed herein, may be a trained machine learning model, when used the term “machine learning model” may be used to describe a trained machine learning model, an untrained machine learning model, a machine learning algorithm (whether trained or untrained).

[0098] A machine learning model may be trained in a supervised setting, Supervised learning may involve training a machine learning model on labeled data, where the input-output pairs guide the learning process. Examples of supervised learning include but are not limited to linear models such as linear regression, logistic regression, ridge regression, lasso regression, tree-based models include decision trees, random forests, and gradient boosting machines (GBM), XGBoost, LightGBM, and CatBoost. Support Vector Machines (SVM), Support Vector Regression (SVR), Bayesian models, Naive Bayes classifiers, Gaussian Process models, and some neural networks such as convolutional neural networks, long short-term memory (LSTM), and feed forward networks.

[0099] A machine learning model may be trained in an unsupervised setting, Unsupervised learning models may identify structures within data without relying on labeled examples.Examples of unsupervised learning include, but are not limited to, Clustering models such as K-Means, DBSCAN, hierarchical clustering, Gaussian Mixture Models (GMM), OPTICS, Affinity Propagation, Fuzzy C-Means, Principal Component Analysis (PCA), t-SNE, UMAP, Independent Component Analysis (ICA), One-Class SVM, Isolation Forest, Local Outlier Factor (LOF), Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Apriori, Eclat, and FP-Growth.

[0100] Ensemble learning methods combine multiple models to improve predictive accuracyWSGR Docket No. 57767-715.601and robustness. Bagging methods like Random Forest and Bagged Decision Trees may reduce variance by training multiple models on different subsets of the data. Boosting techniques, such as AdaBoost, Gradient Boosting, and their advanced variants such as XGBoost, LightGBM, and CatBoost, focus on reducing bias by iteratively refining predictions through weighted models. Stacking models may combine different types of algorithms. Voting classifiers may aggregate predictions from multiple models to make final decisions, using methods like majority voting or weighted voting.

[0101] Clustering may group similar items or data points into clusters based on shared characteristics, such as patterns or signatures. Clustering may comprise partitioning algorithms, hierarchical algorithms, density -based algorithms, or fuzzy clustering algorithms. Partitioning methods(such as K-means) may divide data into a predetermined number of clusters based on distance metrics. Hierarchical methods (such as agglomerative ward) may create a tree-like structure of nested clusters. Density -based methods (such as DBSCAN) may focus on regions of high data point concentration, identifying clusters of arbitrary shapes. Fuzzy clustering techniques (such as Fuzzy C-Means (FCM)) may allow for samples to belong to multiple clusters with varying degrees of membership. Examples of clustering methods include K-means, K-medoids, Fuzzy C-Means (FCM), DBSCAN (Density-Based Spatial Clustering of Applications with Noise), Agglomerative Hierarchical Clustering, Divisive Hierarchical Clustering, Gaussian Mixture Models (GMM), Spectral Clustering, Self-Organizing Maps (SOM), Mean-Shift Clustering, DBSCAN with k-nearest neighbors (K-NN), Affinity Propagation, Bisecting K-means, OPTICS (Ordering Points To Identify the Clustering Structure), Latent Dirichlet Allocation (LDA), Birch (Balanced Iterative Reducing and Clustering using Hierarchies), HDBSCAN (Hierarchical DBSCAN), Fuzzy C-Means with spatial constraints (Fuzzy C-Means-SC), SOM-K-means hybrid, and Deep Clustering Methods using neural networks. These clustering algorithms, when appropriately applied, serve diverse fields such as image recognition, genomics, and market segmentation, enhancing the ability to discover patterns and structures within complex data.

[0102] In some embodiments, the methods disclosed herein may further comprise assessing the performance of the machine learning model. The performance of the machine learning model or any portion of the model may be assessed using a metric (performance metric).Assessment of the models performance may comprise a calculation. The calculation may comprise accuracy. The calculation may comprise specificity. The calculation may comprise sensitivity. The calculation may comprise f-measure. The calculation may comprise fl -measure. The calculation may comprise f2 -measure. The calculation may comprise area under the curve.WSGR Docket No. 57767-715.601The calculation may comprise a confusion matrix. The calculation may comprise true negative. The calculation may comprise true positive. The calculation may comprise false negative. The calculation may comprise false positive. The calculation may comprise precision. The calculation may comprise recall. The calculation may comprise f-beta score. The calculation may comprise Receiver operating curve. The calculation may comprise area under the receiver operating curve (ROCAUC).

[0103] ROCAUC may be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99.

[0104] F-measure may be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99

[0105] Accuracy may be at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99Computer systems

[0106] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 8 shows a computer system 801 that is programmed or otherwise configured to detect a biological state. The computer system 801 may regulate various aspects of the methods such as processing, aligning, sequencing, storing of outputs, assessment of the present disclosure. The computer system 801 may be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device may be a mobile electronic device.

[0107] The computer system 801 includes a central processing unit (CPU, also “processor”WSGR Docket No. 57767-715.601and “computer processor” herein) 805, which may be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 801 also includes memory or memory location 810 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 815 (e.g., hard disk), communication interface 820 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 825, such as cache, other memory, data storage and / or electronic display adapters. The memory 810, storage unit 815, interface 820 and peripheral devices 825 are in communication with the CPU 805 through a communication bus (solid lines), such as a motherboard. The storage unit 815 may be a data storage unit (or data repository) for storing data. The computer system 801 may be operatively coupled to a computer network (“network”) 830 with the aid of the communication interface 820. The network 830 may be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 830 in some cases is a telecommunication and / or data network. The network 830 may include one or more computer servers, which may enable distributed computing, such as cloud computing. The network 830, in some cases with the aid of the computer system 801, may implement a peer-to-peer network, which may enable devices coupled to the computer system 801 to behave as a client or a server.

[0108] The CPU 805 may execute a sequence of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 810. The instructions may be directed to the CPU 805, which may subsequently program or otherwise configure the CPU 805 to implement methods of the present disclosure. Examples of operations performed by the CPU 805 may include fetch, decode, execute, and writeback.

[0109] The CPU 805 may be part of a circuit, such as an integrated circuit. One or more other components of the system 801 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0110] The storage unit 815 may store files, such as drivers, libraries and saved programs. The storage unit 815 may store user data, e.g., user preferences and user programs. The computer system 801 in some cases may include one or more additional data storage units that are external to the computer system 801, such as located on a remote server that is in communication with the computer system 801 through an intranet or the Internet.[OHl] The computer system 801 may communicate with one or more remote computer systems through the network 830. For instance, the computer system 801 may communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung®WSGR Docket No. 57767-715.601Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user may access the computer system 801 via the network 830.

[0112] Methods as described herein may be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 801, such as, for example, on the memory 810 or electronic storage unit 815. The machine executable or machine-readable code may be provided in the form of software. During use, the code may be executed by the processor 805. In some cases, the code may be retrieved from the storage unit 815 and stored on the memory 810 for ready access by the processor 805. In some situations, the electronic storage unit 815 may be precluded, and machine-executable instructions are stored on memory 810.

[0113] The code may be pre-compiled and configured for use with a machine having a processer adapted to execute the code or may be compiled during runtime. The code may be supplied in a programming language that may be selected to enable the code to execute in a precompiled or as-compiled fashion.

[0114] Aspects of the systems and methods provided herein, such as the computer system 801, may be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code may be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk.“Storage” type media may include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium”WSGR Docket No. 57767-715.601refer to any medium that participates in providing instructions to a processor for execution.

[0115] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0116] The computer system 801 may include or be in communication with an electronic display 835 that comprises a user interface (UI) 840. Examples of UFs include, without limitation, a graphical user interface (GUI) and web-based user interface.

[0117] Methods and systems of the present disclosure may be implemented by way of one or more algorithms. An algorithm may be implemented by way of software upon execution by the central processing unit 805. The algorithm can, for example, detect a biological state.

[0118] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend uponWSGR Docket No. 57767-715.601a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.EXAMPLESExample 1: Detection of cancer

[0119] Fragmentomic patterns may be used to infer health states of a patient. RNA fragments in circulation, collected through liquid biopsy, such as cell free RNA (cfRNA) may be used to detect a fragmentomic pattern indicative of a healthy state, or a disease state. Many current efforts to infer health states from fragmentomic utilize cell-free DNA. DNA holds promise but suffers several issues arising from high volumes and multiple sources of noise. For example, technologies that utilize DNA derived from cells that have died, such as having undergone apoptosis or cell lysis, may have to contend with DNA fragmentation patterns that are indicative of cell death rather than a disease, which may mask relevant information to the disease state by overlapping with or altering patterns that result from the disease state. RNA fragmentomics may be derived from living cells or dying cells effectively allowing a model to learn signatures of disease independent of cell death. While promising, current technologies are limited by the amount of cfRNA they may see. Here, a method to use cfRNA fragments for disease detection is discussed that is capable of seeing much larger amount of cfRNA which provides for a more complete look at the RNA fragmentomic patterns. Herein, a model is developed and tested that utilizes fragmentomic RNA patterns to detect disease.

[0120] Fragmentomic sequencing reads may be obtained using a variety of assay methods, such as digital droplet polymerase chain reaction (ddPCR), quantitative polymerase chain reaction (qPCR), and array-based comparative genomic hybridization ( CGH). Once obtained, the sequencing reads may be aligned to a nucleic acid sequence such as a genome or a portion of a genome like an intron or exon.

[0121] Once aligned, a range for the profile, from the start to the end, is defined. This range could span an individual RNA transcript, a combination of multiple transcripts, a genomic interval, the entire genome, or even multiple genomes. For each position, or nucleotide, within the specified range, the number of reads overlapping each nucleotide is counted, thereby generating a vector of read coverage values where each value in the vector is an indication of read coverage (i.e., the number of sequencing reads that align to a specific region of theWSGR Docket No. 57767-715.601nucleotide sequence, providing an indication of the depth of sequencing for that region) at a position in the defined range of the nucleic acid sequence.

[0122] Examples of ranges of the nucleic acid sequence may include Nucleotide, codon, exon, intron, promoter, enhancer, silencer, untranslated region (UTR), gene, operon, intergenic region, repetitive sequence, splice site, chromatin, chromosome, telomere, centromere, heterochromatin, euchromatin, locus, regulatory region, origin of replication, transcription factor binding site, cis-regulatory element, transposon, microRNA, long non-coding RNA (IncRNA), ribosomal RNA (rRNA), messenger RNA (mRNA), transfer RNA (tRNA), genome, multiple genomes, or any combination thereof. Ranges may be of contiguous lengths of nucleotides of at least 1, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 11000, 21000, 31000, 41000, 51000, 61000, 71000, 81000, 91000, 1,000,000, 1,100,000, 2,100,000, 3,100,000, 4,100,000, 5,100,000, 6,100,000, 7,100,000, 8,100,000, 9,100,000, 1,000,000,000, 2,000,000,000, 3,000,000,000 nucleotides, or any combination thereof.

[0123] Optionally, the coverage the quantitative measure, such as a vector of values, may be adjusted. The adjustment may comprise normalization. Normalization may account for variations in sequencing depth and library size across different samples. Normalization may be carried out by dividing each value within the range by the total number of reads for the defined range. This step ensures that the coverage data may be accurately compared between samples or experimental conditions by transforming the raw coverage counts into a standardized scale. However, normalization of the coverage vectors is not mandatory and without normalization the model would also factor in the total abundance of the pattern, rather than simply the pattern itself.

[0124] Adjustment may comprise normalization, standardization, min-max scaling, z-score transformation, whitening, log transformation, binning, one-hot encoding, label encoding, principal component analysis (pea), feature selection, feature extraction, missing value imputation, outlier detection and removal, data augmentation, smoothing, robust scaling, polynomial feature expansion, discretization, noise reduction, or an combination thereof.

[0125] The normalized coverage vectors may be subjected to clustering analysis to group similar coverage profiles. This process utilizes the Fuzzy C-Means (FCM) clustering method, which is designed to allow each data point to belong to multiple clusters, rather than being exclusively assigned to one cluster.

[0126] Briefly, FCM is parameterized with an initial number of clusters (N), e.g., if you have three experimental conditions (Day 1, Day 5, Day 10) you may want to use three clusters. This generates N initial cluster centers, which are represented as vectors, spanning the definedWSGR Docket No. 57767-715.601data range.

[0127] For each initial coverage vector and each cluster, the degree of membership of each coverage vector to each cluster is calculated. The position of each cluster center is recalculated based on minimizing an objective function that represents the distance from any given coverage vector to a cluster center weighted by the membership of that coverage vector in the cluster.

[0128] Once trained, the model will output probabilities of each coverage vector belonging to , or having membership in, each cluster

[0129] As shown in FIG. 2, cfRNA coverage patterns for a transcript, or Ribomarker, differ across BRCA and control samples with low variability within the BRCA and control sample groups but high variability between the two groups of samples. Coverage patterns such as these may be used to train a machine learning model which, once trained, may be used to cluster the BRCA and control samples resulting in clusters that are mostly homogeneous with respect to the sample sources (BRCA and Control). As shown, BRCA samples were clustered into cluster l with high probabilities (e.g., membership scores), while control samples were clustered into cluster O with high probabilities.

[0130] Performance of the model is primarily due to the realization that healthy and cancer cells have different patterns of RNA fragmentation. As illustrated in FIG. 3, a healthy cell and cancerous cell may release an RNA, RNA X, into the extracellular space where it may be degraded or fragmented by various biological means, such as RNase degradation. The RNAs from both cells may have a different suit of proteins bound to them and / or a different tertiary structure, causing the RNAs from each cell to fragment differently. This difference in fragmentation may be sequenced and aligned resulting in coverage vectors which may be visualized as shown in in FIG. 4 where the healthy and cancerous fragment coverage graphs show % coverage from the 5’ to 3’ end of the transcript. The sequencing data shown in FIG. 4 was obtained with the RiboMarker library preparation kit with 10M reads for each patient (3 healthy donors, and 3 stage IV breast cancer patients). When performed on multiple healthy and cancerous samples a consensus sequence may be identified, such as through a machine learning model, which may be used to predict a disease state, in this case healthy or cancerous, for samples. The prediction and identification of a consensus sequence may happen in the same model (e.g., a transformer, a convolutional neural network, a recurrent neural network, or a long short-term memory model, a clustering model, a support vector machine, a logistic regression model, etc.), or in more than one model. By quantitatively comparing a samples coverage vector to a consensus pattern, a probability of a pattern match corresponding to a diagnosis may be given. The method described in this example utilize fuzzy c-means which outputs a membershipWSGR Docket No. 57767-715.601probability for each cluster. Shown in FIG. 4, each sample is assigned a probability to either pattern l or pattern_2 which may be assigned a label of cancer or healthy, this means that the method not only gives a real-valued probability of a cancerous or healthy state but may give the probability of many states for each sample.

[0131] One possible explanation for the difference in fragment levels is a difference in expression among cancer and healthy samples which may lead to a larger population of an RNA in one sample population or another. FIG. 4 shows expression data for healthy and cancer patients (CRC and BRCA). The log2 normalized read count (read counts that have been transformed to a log2 scale and normalized for sequencing depth) of the three sample groups indicates that expression levels are largely the same among healthy and cancer groups. However, in the same samples FIG. 6 shows that fragment patterns are discernable for each group and may be used to train a model capable of distinguishing between CRC, BRCA, and healthy samples.Example 2: Detection of pathogens

[0132] Fragmentomics may be used to identify a pathogen, furthermore fragmentomics may be used to assess the progression of a fungal infection. Morphological changes are a very common and effective strategy for pathogens to survive in a host. During interactions with their host, pathogenic fungi undergo an array of morphological changes that are tightly associated with virulence. Candida albicans switches between yeast cells and hyphae during infection. Thermally dimorphic pathogens, such as Histoplasma capsulatum and Blastomyces species transform from hyphal growth to yeast cells in response to host stimuli. Coccidioides and Pneumocystis species produce spherules and cysts, respectively, which allow for the production of offspring in a protected environment. Finally, Cryptococcus species suppress hyphal growth and instead produce an array of yeast cells — from large polyploid titan cells to micro cells.While the morphology changes produced by human fungal pathogens are diverse, they all allow for the pathogens to evade, manipulate, and overcome host immune defenses to cause disease. As such, they are directly related to the severity and progression of infection and, if detectable, may be used as indicators of the progression of a fungal infection. Pathogens such as fungi, have fragmentomic patterns that may be useful in identifying them when they are present as pathogens. Additionally, fragmentomic patterns of a particular transcript may change according to the tissues of origin of a fungus, such as hyphae, mycelia, and / or arthroconidia. FIG. 7A shows an example of an RNA, RNAX, which is fragmented differently in the Arthroconidia and mycelia and Spherule tissues of a fungus. FIG. 7B shows the percentage of RNAs that wereWSGR Docket No. 57767-715.601found to have different patterns of fragmentation in the different morphologies (morphologically- biased). 90% of the rRNAs were morphologically biased, 40% of tRNAs were morphologically biased, 100% of the snRNA were morphologically biases, 81% of snoRNA were morphologically biased, and 39 % of the unannotated RNA were morphologically biased. This demonstrates that the patterns of fragmentation change across biological states such as morphologies.. FIG. 7C-7E shows fragmentation patterns of Arthroconidia and my celia and Spherule, a consensus sequence for each, shown as pattern 1, pattern 2 and pattern 3 and cluster membership of each sample across three transcripts (Met-CAU-1, U4, and sRNAlocus_4698. The fragmentomic profiles show low variability among within tissues but high variability among tissues, this allows for the generation of distinct consensus sequences. The resulting trained clustering model assigns cluster membership similarly within the tissues and assigns different tissues to different clusters each. These results show that an identifiable fragmentomic pattern is present in different tissues and serve as indicator of both infection, severity of infection, and / or progression of infection.Example 3: tissue of origin detection

[0133] Circulating cfRNAs may be made up of fragments which originate from multiple sources, such as different tissues, pathogens, parasites, and / or individual (such as may be the case during pregnancy). Separating sources may lead to more powerful diagnoses and better information provided by fragmentomic analysis. Tissue of origin analysis through source separation of cfRNa may provide a fine-grained look at the health status of multiple systems and organs in an individual at once. Sources may be separated through a deconvolution process. By generating a database of fragmentation patterns for each specific tissue and pathology (for example a specific cancer), we will be able to deconvolute the patterns of fragmentation in liquid biopsies to identify cell or tissue of origin. Alternatively, a machine learning model may be used to attend to features of the overall fragmentomic profile and identify multiple patterns at once, with this information the model may output information about the patterns, their sources and the health implications of those patterns. This may be achieved through a neural network especially one with a spatial or longitudinal bias, such as a convolutional neural network, a recurrent neural network, a long-short term memory model, a transformer model, a vision transformer, or some combination thereof. Furthermore, a model may be trained to understand contextual information in a fragmentomic profile which may provide improved pattern recognition and detection. It shall be understood that different aspects of the invention may be appreciated individually, collectively, or in combination with each other. Various aspects of the invention described herein may be applied to any of the particular applications disclosed herein.WSGR Docket No. 57767-715.601The compositions of matter disclosed herein in the composition section of the present disclosure may be utilized in the method section including methods of use and production disclosed herein, or vice versa.

[0134] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. WSGR Docket No. 57767-715.601CLAIMS WHAT IS CLAIMED IS:

1. A method of detecting a biological state comprising:a. preparing a plurality of RNA fragments from a biological sample; b. sequencing the plurality of RNA fragments to obtain sequencing reads; c. aligning the sequencing reads to a nucleic acid sequence to obtain aligned sequencing reads;d. determining a quantitative measure of aligned sequencing reads that begin or end at a first genomic position; ande. processing the quantitative measure of the aligned sequencing reads against a reference derived from a reference subject or using a trained machine learning model.

2. The method of claim 1, further comprising adjusting the quantitative measure based at least in part on a count of sequencing reads that align to a range in a nucleic acid sequence.

3. The method of claim 1 or 2, wherein the quantitative measure is normalized to sequencing depth.

4. The method of any one of claims 1-3, wherein the quantitative measure comprises a vector of values.

5. The method of claim 4, wherein the vector of values comprises one of: read coverage values, pseudocount values, or probability of coverage.

6. The method of any one of claims 1-5, wherein the machine learning model comprises a clustering method.

7. The method of claim 6, wherein the clustering method is fuzzy c-means.

8. The method of any one of claims 1-7, wherein the processing comprises detecting a signature.

9. The method of any one of claims 1-8, wherein the processing produces an indication of a biological state.

10. The method of any one of claims 1-9, wherein the biological state is a disease state.

11. The method of claim 10, wherein the disease state comprises at least one of: cancer, presence of a pathogen, organ failure, presence of an autoimmune disease, presence of precancerous lesions, presence of metastasis, presence of type 2 diabetes, inflammation, schizophrenia, Alzheimer’s, Lewy body dementia, asthma, infection with a pathogen, or any combination thereof.WSGR Docket No. 57767-715.60112. The method of claim 11, wherein the disease state comprises cancer.

13. The method of claim 12, further comprising administering a cancer therapeutic to a subject wherein the biological sample was obtained from the subject.

14. The method of claim 11, wherein the disease state comprises presence of a pathogen or infection with a pathogen.

15. The method of claim 14, wherein the pathogen comprises a fungal disease.

16. The method of claim 15, further comprising administering an anti-fungal agent to a subject, wherein the biological sample was obtained from the subject.

17. The method of any one of claims 1-16, wherein the biological state comprises an indication of severity.

18. The method of any one of claims 1-17, wherein the processing produces an indication of an affected region of the body.

19. The method of any one of claims 1-18, wherein the affected region of the body comprises a tissue, an organ, or a biological system.

20. The method of any one of claims 1-19, wherein the biological sample comprises one of venous blood, peripheral blood, stool, nasal swab, mucus, urine, blood, a blood fraction, plasma, serum, saliva, pleural fluid, pericardial fluid, cerebrospinal fluid (CSF), peritoneal fluid, or a combination thereof.

21. The method of any one of claims 1-20, wherein the preparing of the plurality of RNA fragments is untargeted.

22. The method of any one of claims 1-21, wherein the preparing of the plurality of RNA fragments is targeted23. The method of any one of claims 1-22, wherein the preparing of the plurality of RNA fragments comprises an enrichment of an RNA sequences of interest.

24. The method of any one of claims 1-23, wherein the preparing of the plurality of RNA fragments comprises an enrichment of an RNA fragment having specific phosphorylation states of the ends.

25. The method of any one of claims 1-24, wherein the preparing of the plurality of RNA fragments comprises a selection for an RNA sequences of interest.

26. The method of any one of claims 1-25, wherein the preparing of the plurality of RNA fragments comprises a selection for an RNA fragment having specific phosphorylation states of the ends.

27. The method of any one of claims 1-26, wherein the preparing of the plurality of RNA fragments comprises a selection of an RNA fragment having specific phosphorylationWSGR Docket No. 57767-715.601states of the ends.

28. The method of claim 27, wherein the phosphorylation state is at the 5’ end of the RNA fragment or RNA sequence.

29. The method of claim 27, wherein the phosphorylation state of the 5’ end of the RNA fragment or RNA sequence is monophosphorylated.

30. The method of claim 27, wherein the phosphorylation state of the 5’ end of the RNA fragment or RNA sequence is diphosphorylated.

31. The method of claim 27, wherein the phosphorylation state of the 5’ end of the RNA fragment or RNA sequence is triphosphorylated.

32. The method of claim 27, wherein the phosphorylation state is at the 3’ end of the RNA fragment or RNA sequence.

33. The method of claim 27, wherein the phosphorylation state of the 3’ end of the RNA fragment or RNA sequence is 3 ’-phosphate (3’-P).

34. The method of claim 27, wherein the phosphorylation state of the 3’ end of the RNA fragment or RNA sequence is 2’, 3 ’-cyclic phosphate (cyc-P).