RNA liquid biopsy and nanopore sequencing
Long-read nanopore sequencing and machine learning models improve the sensitivity and specificity of esophageal cancer detection by identifying differentially expressed cell-free RNAs, addressing the limitations of current liquid biopsy technologies.
Patent Information
- Application Number
- PCT/US2025/030676
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-05-14
- Filing Date
- 2025-05-22
- Publication Date
- 2025-11-27
AI Technical Summary
Current liquid biopsy technologies for esophageal cancer exhibit lower detection sensitivity and are constrained by short-read sequencing, which cannot sequence full-length cell-free RNAs, necessitating the development of novel approaches with improved sensitivity and specificity for early detection.
A method involving single-molecule, long-read nanopore sequencing to detect cell-free RNAs, followed by annotation and combination with a custom reference transcriptome to identify differentially expressed RNAs, trained using machine learning models to predict disease status.
Enhances the sensitivity and specificity of esophageal cancer detection by identifying novel cell-free RNAs, enabling accurate disease classification and treatment prediction.
Smart Images

Figure US2025030676_27112025_PF_FP_ABST
Abstract
Description
[0001] RNA LIQUID BIOPSY AND NANOPORE SEQUENCING
[0002] CROSS REFERENCE TO RELATED APPLICATIONS
[0003] The present application claims the benefit of U.S. Provisional Application No. 63 / 650,485 filed on May 22, 2024, U.S. Provisional Application No. 63 / 710,427 filed on October 22, 2024, U.S. Provisional Application No. 63 / 805,904 filed on May 15. 2025, and U.S. Provisional Application No. 63 / 775,936 filed on March 21, 2025, the entire disclosures of which are hereby incorporated herein by reference in their entirety for all purposes.
[0004] BACKGROUND
[0005] Esophageal cancer is the sixth leading cause of cancer-related death around the world, with a 5-year survival rate of only 20%, underscoring the urgent need for early detection methods. Current liquid biopsy technologies based on cell-free RNA show promise for the early detection of certain cancers and precancerous conditions including lung, liver, colorectal, and stomach cancers. However, these methods exhibit lower detection sensitivity for esophageal cancer, necessitating the development of novel approaches with improved sensitivity and specificity. Moreover, current cell-free RNA liquid biopsy technologies are also constrained by the use of short-read sequencing technologies, which are unable to sequence full-length cell-free RNAs that are longer than a few hundred nucleotides in length.
[0006] SUMMARY
[0007] Provided herein is a method for determining a disease status in a subject by training one or more machine learning models to predict the disease status based on cell-free RNAs. The method includes (a) collecting a biological sample from the subject comprising cell-free ribonucleic acids (RNAs); (b) detecting, using a sequencing method, the cell-free RNAs in the biological sample, wherein the sequencing method generates sequencing read data for the cell-free RNAs; (c) identifying a subset of cell-free RNAs from the cell-free RNAs in the biological sample by annotating the sequencing read data for the cell-free RNAs to sequencing read data for a standard reference transcriptome; (d) combining, based on the annotated sequencing read data for the identified subset of cell-free RNAs, the subset of cell- free RNAs with the standard reference transcriptome to generate a custom reference transcriptome comprising a custom set of cell-free RNAs: (e) quantify ing transcript abundance for the custom set of cell-free RNAs in the custom reference transcriptome; (f) identifying a list of differentially expressed cell-free RNAs by comparing the transcript abundance of the custom set of cell-free RNAs in the custom reference transcriptome to cell- free RNA transcript abundance values from healthy subjects; (g) training, using at least the list of differentially expressed RNA transcripts, one or more machine learning models to predict the disease status of the subject. The method can also include step (h) outputting the one or more trained machine learning models that predict the disease status of the subject into a production environment. Optionally, the identification of the subset of cell-free RNAs from the cell-free RNAs in the biological sample are not found in the standard reference transcriptome. Thus, for example, step (c) can be identifying a subset of cell-free RNAs from the cell-free RNAs in the biological sample by annotating the sequencing read data for the cell-free RNAs that are not present in sequencing read data from a standard reference transcriptome.
[0008] Also provided herein is a method for determining a disease status in a subject by using one or more trained machine learning models that predict the disease status based on the cell-free RNA profile of the subject. The method includes (a) collecting a biological sample from the subject comprising cell-free ribonucleic acids (RNAs); (b) detecting, using a sequencing method, the cell-free RNAs in the biological sample, wherein the sequencing method generates sequencing read data for the cell-free RNAs; (c) identify ing a subset of cell-free RNAs from the cell-free RNAs in the biological sample by annotating the sequencing read data for the cell-free RNAs to sequencing read data for a standard reference transcriptome; (d) combining, based on the annotated sequencing read data for the identified subset of cell-free RNAs, the subset of cell-free RNAs with the standard reference transcriptome to generate a custom reference transcriptome comprising a custom set of cell- free RNAs; (e) quantifying transcript abundance for the custom set of cell-free RNAs in the custom reference transcriptome; (f) identifying a list of differentially expressed cell-free RNAs by comparing the transcript abundance of the custom set of cell-free RNAs in the custom reference transcriptome to cell-free RNA transcript abundance values from healthy subjects; (g) predicting, using one or more trained machine learning models, the disease status of the subject based at least on the list of differentially expressed cell-free RNAs, wherein the disease status is a negative disease status, a pre-disease status, or a positive disease status. The method can also include step (h) outputting the predicted disease status of the subject in a production environment. Optionally, the identification of the subset of cell- free RNAs from the cell-free RNAs in the biological sample are not found in the standard reference transcriptome. Thus, for example, step (c) can be identifying a subset of cell-free RNAs from the cell-free RNAs in the biological sample by annotating the sequencing read data for the cell-free RNAs that are not present in sequencing read data from a standard reference transcriptome.
[0009] Also provided herein are methods of treating a subject having the precancerous status or the cancer status. The method comprises utilizing the method above for using one or more trained machine learning models to predict the diseases status of the subject and administering a treatment (e.g., a therapeutic agent) to the subject. Optionally the method above for using one or more trained machine learning models to predict the diseases status of the subj ect may be repeated to determine the efficacy of treatment.
[0010] Also provided herein is a method for determining one or more drug targets to treat a disease in a subject. The method includes (a) collecting a blood sample from the subject comprising cell-free ribonucleic acids (RNAs); (b) detecting, using a sequencing method, the cell-free RNAs in the blood sample, wherein the sequencing method generates sequencing read data for the cell-free RNAs; (c) quantifying, using the sequencing read data, transcript abundance for the cell-free RNAs in the blood sample; (d) identifying one or more differentially expressed cell-free RNA by comparing the transcript abundance of the cell-free RNAs in the blood sample to cell-free RNA transcript abundance values from healthy subjects; and (e) providing, based on the one or more differentially expressed cell-free RNAs, one or more drug targets to treat the disease in a subject.
[0011] The details of one or more embodiments are set forth in the description below. Other features, objects, and advantages will be apparent from the description and from the claims.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The present application includes the following figures. The figures are intended to illustrate certain embodiments and / or features of the compositions and methods, and to supplement any description(s) of the compositions and methods. The figures do not limit the scope of the compositions and methods, unless the written description expressly indicates that such is the case.
[0014] FIG. 1 shows a computing environment in accordance with various embodiments.
[0015] FIG. 2 shows a block diagram of an exemplary machine learning pipeline comprising several subsystems that work together to train, validate, and implement one or more machine learning models in accordance with various embodiments. FIG. 3 illustrates a flowchart for training a machine learning model in accordance with various embodiments.
[0016] FIG. 4A shows a schematic of the LOCATE-seq workflow. FIGs 4B-4E show gene tracks of examples of novel cell-free RNA transcripts identified using the LOCATE-seq workflow'.
[0017] FIGs. 4F-4H are box plots showing the number of protein-coding RNAs detected across patient samples and transcriptome references (FIG 4F); the number of long noncoding RNAs detected across patient samples and transcriptome references (FIG. 4G); and the number of novel cell-free RNAs detected across patient samples (FIG. 4H). FIG. 41 shows a bar graph of the genomic distribution of novel cell-free RNAs. FIG. 4J show s the size distributions of novel and GENCODE cell-free RNAs.
[0018] FIGs. 5A-5G show mitochondrial RNA enrichment in dysplasia and cancer. FIGs. 5A-5C are volcano plots illustrating the differential expression (DE) analysis of healthy vs. dysplasia cell-free RNA (FIG. 5A); healthy vs. cancer cell-free RNA (FIG. 5B); and healthy vs dysplasia and cancer cell-free RNA (FIG. 5C). FIG. 5D and 5E are UpSet plots of common upregulated DE genes and downregulaled DE genes, respectively. FIG. 5F shows a box plot of the number of detected mitochondrial genes. FIG. 5G is a heatmap showing the hierarchical clustering of mitochondrial RNA expression.
[0019] FIGs. 6A-6F show the differential abundance of novel cell-free RNA in dysplasia and cancer. FIGs. 6A-6C are volcano plots shows differential expression (DE) analysis of healthy vs. dysplasia cell-free RNA (FIG. 6A); healthy vs. cancer cell-free RNA (FIG. 6B); and healthy vs dysplasia and cancer cell-free RNA (FIG. 6C). FIGs. 6D and 6E are UpSet plots of the common upregulated DE genes and downregulated DE genes, respectively. FIG. 6F is a heat map showing the hierarchical clustering of novel cell-free RNA expression.
[0020] FIGs. 7A-7D show that novel cell-free RNA features enable accurate disease classification. FIG. 7A is a box plot of the number of detected unique novel cell-free RNAs. FIG. 7B is a bar plot of the number of detected unique novel cell-free RNAs in each healthy individual or patient. FIG. 7C is a receiver operating characteristic curve for a logistic regression model trained using known and novel cell-free RNA features for classification of high-grade dysplasia. FIG. 7D is a receiver operating characteristic curve for a logistic regression model trained using know n and novel cell-free RNA features for classification of early-stage esophageal cancer. FIGs. 8A-8G show dysregulated pathways as potential therapeutics in dysplasia and cancer. FIG. 8 A is a bar graph showing the top 10 most significantly enriched pathways identified from gene set enrichment analysis for pathways enriched in dysplasia patient plasma. FIGs. 8B-8G are box plots of normalized counts for HER2 cell-free RNA (FIG. 8B), CTLA-4 cell-free RNA (FIG. 8C), LAG-3 cell-free RNA (FIG. 8D), PD-L1 cell-free RNA (FIG. 8E), TIGIT cell-free RNA (FIG. 8F), and MDM2 cell-free RNA (FIG. 8G).
[0021] DETAILED DESCRIPTION
[0022] The following description recites various aspects and embodiments of the present compositions and methods. No particular embodiment is intended to define the scope of the compositions and methods. Rather, the embodiments merely provide non-limiting examples that are at least included within the scope of the disclosed compositions and methods. The description is to be read from the perspective of one of ordinary skill in the art; therefore, information well known to the skilled artisan is not necessarily included.
[0023] Esophageal cancer is the sixth leading cause of cancer-related death around the world, with a 5-year survival rate of only 20%. Current projections estimate that 880,000 people will die from esophageal cancer in 2040, highlighting the urgent need for better methods to detect esophageal cancer at the earliest signs of disease. Gastroesophageal reflux disease can lead to precancerous Barrett’s esophagus with increasing dysplasia, which progressively increases the risk of developing esophageal adenocarcinoma. However, currently available liquid biopsy methods based on cell-free DNA are unable to detect cancer at the earliest stages.
[0024] Liquid biopsy technologies based on cell-free RNA show promise for the early detection of certain cancers and precancerous conditions. Cell-free RNA is actively secreted by cells into the circulation and protected from degradation by being encapsulated in extracellular vesicles (EVs), which are nanometer-sized lipid bilayer particles that reflect their healthy or diseased cell of origin. Cell-free RNAs are comprised of various classes of short and long RNAs, including protein-coding messenger RNAs (mRNAs), long noncoding RNAs (IncRNAs), and repeat RNAs. Although repeat-aware RNA liquid biopsy technologies such as COMPLETE-seq enable the accurate classification of certain cancer types, including lung, liver, colorectal, and stomach cancers, other cancers such as esophageal cancer exhibit lower detection sensitivity, highlighting the need for novel approaches that enable cancer early detection with higher sensitivity and specificity. Current RNA liquid biopsy approaches are also constrained by the use of short-read sequencing technologies, which are unable to sequence full-length cell-free RNAs that are longer than a few hundred nucleotides in length. Single-molecule, long-read nanopore sequencing enables full-length cell-free RNA sequencing.
[0025] Thus, provided herein is a method for determining a disease status in a subject. The method includes collecting a biological sample from the subject comprising cell-free ribonucleic acids (RNAs). A sequencing method is used to detect the cell-free RNAs in the biological sample, wherein the sequencing method generates sequencing read data for cell- free RNA in the biological sample. A subset of cell-free RNAs from the cell-free RNAs in the biological sample are identified by annotating the sequencing read data for the cell-free RNAs to sequencing read data for a standard reference transcriptome. Then, based on the annotated sequencing read data for the identified subset of cell-free RNAs, the subset of cell- free RNAs are combined with the standard reference transcriptome to generate a custom reference transcriptome comprising a custom set of cell-free RNAs. Then, the custom set of cell-free RNAs in the custom reference transcriptome are quantified. By comparing the transcript abundance of the custom set of cell-free RNAs in the custom reference transcriptome to cell -free RNA transcript abundance values from healthy subjects, a list of differentially expressed cell-free RNAs are identified. The list of differentially expressed cell- free RNAs are used to train one or more machine learning models to predict the disease status of the subject and one or more trained machine learning models that predict the disease status of the subject are output into a production environment.
[0026] Computing Environment
[0027] Certain processes and methods described herein are performed within a computing environment comprising a computer, microprocessor, software, module, other machines such as sequencers, or combinations thereof. The methods described herein typically are computer-implemented methods, and one or more portions or steps of the method are performed by one or more processors (e.g., microprocessors), computers, systems, apparatuses, or machines (e.g.. microprocessor-controlled machine). Computers, systems, apparatuses, machines, and computer program products suitable for use often include, or are utilized in conjunction with, computer readable storage media. Non-limiting examples of computer readable storage media include memory', hard disk, CD-ROM, flash memory' device and the like. Computer readable storage media generally are computer hardware, and often are non-transitory computer-readable storage media. Computer readable storage media are not computer readable transmission media, the latter of which are transmission signals per se.
[0028] FIG. 1 shows a computing environment 100 in accordance with aspects of the present disclosure. Computing environment 100 is only one example of a suitable computing environment and is not intended to suggest any limitation as to the scope of use or functionality of the systems, methods, and data structures described herein. Neither should computing environment 100 be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in computing environment 100. A subset of systems, methods, and data structures shown in FIG. 1 can be utilized in certain embodiments. Systems, methods, and data structures described herein are operational with numerous other general purpose or special purpose computing system environments or configurations. Examples of know n computing systems, environments, and / or configurations that may be suitable include, but are not limited to, personal computers, server computers, thin clients, thick clients, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like
[0029] Computing environment 100 includes a client device 105, data repositories 110, and a liquid biopsy platform 115 connected to each other by a network 120. Although FIG. 1 illustrates a particular arrangement of a client device 105. data repositories 110, and a liquid biopsy platform 115, this disclosure contemplates any suitable arrangement of a client device 105, data repositories 110, and a liquid biopsy platform 115. As an example, and not by way of limitation, two or more client devices 105, a data repository 110, and liquid biopsy platform 115 may be connected to each other directly, bypassing network 120. As another example, two or more client devices 105, a data repository 110, and a liquid biopsy platform 115 may be physically or logically co-located with each other in whole or in part. Moreover, although FIG. 1 illustrates a particular number of a client device 105, a data repository 110, a liquid biopsy platform 115, and network 120, this disclosure contemplates any suitable number of client devices 105, data repositories 110. liquid biopsy platforms 115. and networks 120. As an example, and not by way of limitation, computing environment 100 may include multiple client devices 105, data repositories 110, liquid biopsy platforms 115, and networks 120. A client device 105 is an electronic device including hardware, software, or embedded logic components or a combination of two or more such components and capable of interacting with data repository 110 and the liquid biopsy platform 1 15 with respect to identifying nucleic acids in a sample from a subject to determine a disease status in accordance with techniques of the disclosure. The client device 105 may be a computing device such as a conventional computer, a distributed computer, or any other type of computer (e.g.. portable handheld devices, general purpose computers such as personal computers and laptops, workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, and the like). The computing device may execute and run various types and versions of software applications and systems (e.g., Internet-related apps, communication applications (e.g., E-mail applications, short message service (SMS) applications), LIMS applications, auto review system 135) and operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX® or UNIX-like operating systems. Linux or Linux-like operating systems such as Google Chrome™ OS) including various mobile operating systems (e.g., Microsoft Windows Mobile®, iOS®, Windows Phone®, Android™, BlackBerry®, Palm OS®) using one or more communication protocols. Portable handheld devices may include cellular phones, smartphones, (e.g., an iPhone), tablets (e.g., iPad®), personal digital assistants (PDAs), and the like. Wearable devices may include Google Glass® head mounted display, and other devices. This disclosure contemplates any suitable client device 105 configured to generate and output product target discovery content to a user. For example, users may use client device 105 to execute one or more applications, which may generate one or more discovery or storage requests that may then be serviced in accordance with the teachings of this disclosure. A client device 105 may provide an interface 125 (e.g., a graphical user interface) that enables a user of the client device 105 to interact with the client device 105. The client device 105 may also output information to the user via this interface 125. Although FIG. 1 depicts only one client device 105, any number of client devices 105 may be supported.
[0030] In some aspects, the client device 105 includes a processing unit, a system memory, and a system bus that operatively couples various system components including the system memory to the processing unit. There may be only one or there may be more than one processing unit, such that the processor of computing device includes a single centralprocessing unit (CPU), or a plurality of processing units, commonly referred to as a parallel processing environment. The system bus may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. The system memory’ may also be referred to as simply the memory and includes read only memory (ROM) and random access memory (RAM). A basic input / output system (BIOS), containing the basic routines that help to transfer information between elements within the computing device, such as during start-up, is stored in ROM. The computing device may further include a hard disk drive for reading from and writing to a hard disk, a magnetic disk drive for reading from or writing to a removable magnetic disk, and an optical disk drive for reading from or writing to a removable optical disk such as a CD ROM or other optical media.
[0031] The hard disk drive, magnetic disk drive, and optical disk drive may be connected to the system bus by a hard disk drive interface, a magnetic disk drive interface, and an optical disk drive interface, respectively. The drives and their associated computer-readable media provide non-volatile storage of computer-readable instructions, data structures, program modules and other data for the client device 105. Any ty pe of computer-readable media that can store data that is accessible by a computer, such as magnetic cassettes, flash memory cards, digital video disks, Bernoulli cartridges, random access memories (RAMs), read only memories (ROMs), and the like, may be used in the operating environment.
[0032] A number of program modules may be stored on the hard disk, magnetic disk, optical disk, ROM. or RAM. including an operating system, one or more application programs, other program modules, and program data. A user may enter commands and information into the client device 105 through input devices such as a keyboard and pointing device (e.g., mouse). Other input devices may include a microphone joystick, game pad, satellite dish, scanner, or the like. These and other input devices are often connected to the processing unit through a serial port interface that is coupled to the system bus. but may be connected by other interfaces, such as a parallel port, game port, or a universal serial bus (USB). A monitor or other type of display’ device is also connected to the system bus via an interface, such as a video adapter. In addition to the monitor, computing devices typically include other peripheral output devices, such as speakers and printers.
[0033] Network 120 can be any type of network familiar to those skilled in the art that may support data communications using any of a variety of available protocols including without limitation TCP / IP (transmission control protocol / Intemet protocol), SNA (systems network architecture). IPX (Internet packet exchange), AppleTalk®, and the like. Merely by way of example, network(s) 120 may be a local area network (LAN), networks based on Ethernet. Token-Ring, a wide-area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infra-red network, a wireless network (e.g., a network operating under any of the Institute of Electrical and Electronics (IEEE) 1002.11 suite of protocols, Bluetooth®, and / or any other wireless protocol), and / or any combination of these and / or other networks.
[0034] Links 130 may connect a client device 105, LIMS 125, data repositories 110, and liquid biopsy platform 115 to network 120 or to each other. This disclosure contemplates any suitable links 130. In particular embodiments, one or more links 130 include one or more wireline (such as for example Digital Subscriber Line (DSL) or Data Over Cable Sen ice Interface Specification (DOCSIS)), wireless (such as for example Wi-Fi or Worldwide Interoperability for Microwave Access (WiMAX)), or optical (such as for example Synchronous Optical Network (SONET) or Synchronous Digital Hierarchy (SDH)) links. In particular embodiments, one or more links 130 each include an ad hoc network, an intranet, an extranet, a VPN, a LAN, a WLAN, a WAN, a WWAN, a MAN, a portion of the Internet, a portion of the PSTN, a cellular technology-based network, a satellite communications technology-based network, another link 130, or a combination of two or more such links 130. Links 130 need not necessarily be the same throughout a computing environment 100. One or more first links 130 may differ in one or more respects from one or more second links 130.
[0035] Servers 135 are computers or computing devices designed to manage, store, process, and deliver data or services (e.g., data and services related to LIMS 125 and liquid biopsy platform 115) to other devices, such as client 105, over network 120. The servers 135 can be physical machines or virtual instances (e.g., virtual machines) created through virtualization technologies. The servers 135 operate using specialized hardware, such as high- performance CPUs, large amounts of RAM. and storage arrays, to handle intensive workloads. They are equipped with server-grade operating systems (e.g., Windows Server, Linux-based distributions) and software that facilitate specific functions, such as hosting websites, managing databases, running applications, or handling email services. To provide data or services such as those related to LIMS 125 and liquid biopsy platform 115, servers 135 listen for incoming requests from clients (e.g., client device 105) via network protocols (e.g., HTTP for web services, SMTP for email, or FTP for file transfers). When a request is received, a server processes it using its resources and returns the appropriate response, such as delivering lab results, analyzing lab orders and reports, and generating reports. Servers 135 can also support multi-user environments, enabling simultaneous access to shared resources, and can be integrated into a distributed system or cloud infrastructure to ensure scalability, reliability, and high availability . For example, infrastructure as a service (laaS) is one particular type of cloud computing. laaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In an laaS model, a cloud computing provider can host the infrastructure components (e.g., servers 135, storage devices such as data repositories 1 10, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., a hypervisor layer), or the like). In some cases, an laaS provider may also supply a variety7of sendees to accompany those infrastructure components (example sendees include billing software, monitoring software, logging software, load balancing software, clustering software, etc.). Thus, as these services may be policy-driven, laaS users may be able to implement policies to drive load balancing to maintain application availability7and performance.
[0036] In some instances, laaS users may access resources and services through network 120 and use the cloud provider's services to install the remaining elements of an application stack. For example, the user can log in to the laaS platform to create virtual machines such as servers 135, install operating systems (OSs) on each virtual machine, deploy middleware including databases such as those that may be included in data repositories 110, create storage buckets for workloads and backups, and even install enterprise software such as software 150 into one or more virtual machines. Users can then use the provider's services to perform various functions, including troubleshooting laboratory equipment issues, monitoring performance of laboratory7equipment, reviewing laboratory7orders, results, and reports, generating reports, analyzing laboratory results, generating laboratory7results, etc.
[0037] A data repository such as one of the data repositories 110 is a centralized location for storing, managing, and maintaining data. It functions as a digital warehouse that enables efficient organization, retrieval, and analysis of data. Data repositories 110 are used to store data (e.g., structured, semi-structured, or unstructured data) and other information for use byclient device 105 and / or liquid biopsy platform 115. The data repositories 110 are characterized by their storage infrastructure, which can include on-premise servers, cloudbased systems, or hybrid solutions. They include access control mechanisms to ensure secure data management, supporting various formats such as relational, semi-structured, or unstructured data. Scalability- is a key attribute, allowing data repositories 110 to handle growing data volumes, while integration capabilities enable seamless connection with external systems and data pipelines. Querying tools, such as SQL, APIs, or other interfaces, are used to facilitate efficient data retrieval and manipulation. Although the data repositories 110 are shown on the same server in FIG. 1 it should be understood that they could be spread across multiple servers in various configurations without departing from the spirit and scope of the present disclosure.
[0038] In some instances, a data repository such as one of the data repositories 110 is a database which is a specialized form of data store designed to manage structured data efficiently. While all databases are data stores, not all data stores are databases. For example, in other instances, a data repository such as one of the data repositories 110 is a data store which includes various systems for storing data, such as file systems, key-value stores, and object stores. This broad category covers any technology used to persist data, whether structured, semi-structured, or unstructured. A client device 105 may interact with a data store through a structured process that involves communication over a network using APIs, query languages, or other protocols. The client device 105 sends requests to the data store to perform operations such as retrieving, updating, inserting, or deleting data. These requests may be formatted in a query language (e.g., SQL for relational databases) or through API calls (e.g., REST or GraphQL for web-based systems). The data store, equipped with access control mechanisms, authenticates the client and verifies permissions before executing the requested operations. Once processed, the data store returns the requested data or a confirmation of the operation's success back to the client device, for example in a structured format such as JSON or XML for easy parsing and utilization. This interaction is governed by network protocols such as HTTP or HTTPS and optimized for efficiency, scalability, and security, ensuring that the data exchange is reliable and compliant with organizational or legal standards.
[0039] The liquid biopsy platform 1 15 is an integrated technological system or suite of tools that enable a set of laboratory analyses, assays, or experiments. The liquid biopsy platform 115 comprises a combination of instrumentation 140, reagents 145, software 150, and artificial intelligence models 155 that facilitate the identification of nucleic acids (e.g.. cell-free RNAs) in a sample obtained from a subject to determine a disease status. One of the primary components of the liquid biopsy platform 115 is the instrumentation 140 or hardware. Instrumentation 140 includes the physical devices used to conduct one or more assays, such as PCR machines, sequencing devices like the Illumina NovaSeq or Thermo Fisher Ion Tonent. mass spectrometers, microarray readers, and the like. The reagents 145 used by the liquid biopsy platform 115 can include kits, chemicals, enzy mes, primers, probes, and other materials that are specifically designed or validated for use with the liquid biopsy platform 115. Software 150 and / or data analysis tools can include integrated or compatible software for instrument control, data acquisition, and analysis. For example, the liquid biopsy platform 115 can integrate base calling and alignment software for sequencing data and / or image analysis tools for high-throughput screening. In addition, software 150 may utilize proprietary data formats, cloud-based analysis environments, or API integrations to facilitate advanced data processing and interoperability. The liquid biopsy platform 115 can further leverage artificial intelligence models 155 to enhance the capabilities and efficiencies during data analysis. For example, artificial intelligence can be used for data acquisition and preprocessing, instrument control and automation, data analysis and interpretation, and the like. The liquid biopsy platform 115 may reside in a variety of locations including servers 135. For example, a liquid biopsy platform 115 executed by server 135 may be local to server 135 or may be remote from server 135 (e.g., on a virtual machine) and in communication with server 135 via a network-based or dedicated connection of network 120.
[0040] In some instances, instrumentation 140 includes at least a sequencing device which is any machine capable of sequencing one or more nucleic acid molecules to generate raw sequencing data (e.g., reads). Library7prepared nucleic acid samples may be pooled and loaded into lanes of a sequencing flow cell. The flow cell may be loaded into the sequencing device and imaged to generate sequence data. To achieve this, reagents (i. e. , reagents 145) that interact with the nucleic acid samples fluoresce at particular wavelengths in response to an excitation beam and thereby return a signal for imaging. For instance, the fluorescent components may be generated by fluorescently tagged nucleic acids that hybridize to complementary7molecules of the components or to fluorescently tagged nucleotides that are incorporated into an oligonucleotide using a polymerase. As will be appreciated by those skilled in the art, the wavelength at which the dyes of the sample are excited and the wavelength at which they fluoresce will depend upon the absorption and emission spectra of the specific dyes. The sequencing device may optionally include or be operably coupled to its own dedicated sequencer computer with its own input / output mechanisms, one or more processors, and memory. Additionally or alternatively, the sequencing device may be operably coupled to a serv er 135 or client device 105 via network 120. Client device 105 may access the raw sequencing data files from data repositories 110 and execute instructions for analyzing or communicating the sequence data to network 120. As discussed in detail below with respect to FIG. 3, the liquid biopsy platform 115 can identify nucleic acids in a sample from a subject. To achieve this, the liquid biopsy platform 115 is configured to detect nucleic acids (e.g., cell-free RNAs) in a biological sample (e.g., blood or plasma) from a subject using a sequencing method. For example, nucleic acids are isolated from the biological sample using a first set of reagents (e.g., cell lysis buffers, enzymes, alcohol precipitation, etc.). Then to detect the sequence of the nucleic acids, a second set of reagents (e.g., library construction reagents) and a sequencing instrument (e.g., sequencing device described above) are used. To determine a disease status for the subject based on the identified nucleic acids (e.g., cell-free RNAs), the liquid biopsy platform 115 leverages various bioinformatic software (e.g., alignment and variant calling tools) to process the sequencing output files and to identify nucleic acids that are present in the biological sample. Then using artificial intelligence modeling (e.g., machine learning), the liquid biopsy platform 115 leams the underlying patterns of the data (e.g., what subset of nucleic acids correspond to a particular disease status) and outputs a disease status prediction. Training and Using a Machine Learning Model for Predicting Disease Status
[0041] Artificial intelligence (Al) is a broad field of computer science focused on creating systems or machines capable of performing tasks that typically require human intelligence, such as reasoning, problem-solving, understanding language, recognizing patterns, and making decisions. Machine learning is a subset of Al that enables computers to learn from data and experience without being explicitly programmed for each task. FIG. 2 shows a block diagram of a machine learning pipeline 200 comprising several subsystems that work together to train, validate, and implement one or more machine learning models in accordance with various embodiments. The machine learning pipeline 200 may be executed as part of the liquid biopsy platform 115 of the computing environment 100 described in FIG. 1. The machine learning pipeline 200 comprises a data subsystem 205 for collecting, generating, preprocessing, and labeling of training and validation datasets 210, training and validation subsystem 215 that facilitates the training and validation of one or more machine learning algorithms 220, and inference subsystem 225 for deploying and implementing one or more trained machine learning models 230 independently or in combination with one or more other systems or services 235 for downstream processes.
[0042] As used herein, machine learning algorithms (also described herein as simply algorithm or algorithms) are procedures that are run on datasets (e.g., training and validation datasets) and perform pattern recognition on datasets, leam from the datasets, and / or are fit on the datasets. Examples of machine learning algorithms include linear and logistic regression, decision trees, artificial neural networks, k-means, and k-nearest neighbor. In contrast, machine learning models (also described herein as simply model or models) are the output of the machine learning algorithms and are comprised of model data and a prediction algorithm. In other words, the machine learning model is the program that is saved after running a machine learning algorithm on training data and represents the rules, numbers, and any other algorithm-specific data structures required to make inferences. For example, a linear regression algorithm may result in a model comprised of a vector of coefficients with specific values, a decision tree algorithm may result in a model comprised of a tree of if-then statements with specific values, or neural network, backpropagation, and gradient descent algorithms together result in a model comprised of a graph structure with vectors or matrices of weights with specific values.
[0043] Data Subsystem
[0044] Data subsystem 205 is used to collect, generate, preprocess, and label data to be used to train and validate one or more machine learning algorithms 220. The data collection can include exploring various data sources such as public datasets, private data collections, or real-time data streams, depending on a project’s needs. In some instances, a data source is a public or online repository' of information or examples pertinent to a general or target domain space. Many domains have publicly available datasets provided by governments, universities, or organizations. For example, many government and private entities offer datasets on healthcare, environmental data, and more through various portals. For proprietary needs, data might be available through partnerships or purchases from private companies that specialize in data aggregation. In other instances, a data source is a private repository of information or examples pertinent to a general or target domain space. For example, a data source can be data repositories 110 that store various sequencing data files generated by the liquid biopsy platform 115 as described in FIG. 1 or public trans criptomic databases. Once a data source is identified, data subsystem 205 can be used to collect data through appropriate methods such as downloading from online repositories, web scraping, using APIs for real-time data, creating datasets through surveys and experiments, or by running assays. The acquired raw data may be further preprocessed to generate the training and validation datasets 210.
[0045] In some instances, raw data may be generated as opposed to being collected or acquired. Data generating may comprise data synthesis and / or data augmentation. Different data synthesis and / or data augmentation techniques may’ be implemented by the data subsystem 205 to generate data to be used for the training and validation subsystem 215. Data synthesizing involves creating entirely new data points from scratch. This technique may be used when real data is insufficient, too sensitive to use, or when the cost and logistical barriers to obtaining more real data are too high. The synthesized data should be realistic enough to effectively train a machine learning model, but distinct enough to comply with regulations (e.g., copyright and data privacy), if necessary’. Techniques such as Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs) may be used to generate new data examples. These models learn the distribution of real data and attempt to produce new data examples that are statistically similar but not identical. Data augmentation, on the other hand, refers to techniques used to artificially expand the size of a dataset by creating modified versions of existing data examples. The primary goal of data augmentation is to increase variation in the data in order to make the model more robust to variations it might encounter in the real world, thereby improving its ability to generalize from the training data to unseen data. This is especially common in image and speech recognition tasks but is applicable to other data types as well. For images, data augmentation may include rotations, flipping, scaling, or altering the color / lighting conditions. For text, data augmentation may include synonyms replacement, back translation, or sentence shuffling. For audio, data augmentation may include changes made to pitch, speed, or background noise.
[0046] Preprocessing may be implemented using data subsystem 205, serving as a bridge between raw data acquisition and effective model training. The primary’ objective of preprocessing is to transform raw data into a format that is more suitable and efficient for analysis, ensuring that the data fed into machine learning algorithms is clean, consistent, and relevant. This step can be useful because raw data often comes with a variety of issues such as missing values, noise, irrelevant information, and inconsistencies that can significantly hinder the performance of a model. By standardizing and cleaning the data beforehand, preprocessing helps in enhancing the accuracy and efficiency^ of the subsequent analysis, making the data more representative of the underlying problem the model aims to solve.
[0047] Other raw data preprocessing techniques include data cleaning, normalization, feature extraction, dimensionality reduction, and the like. Data cleaning may involve removing duplicates, filling in missing values, or filtering out outliers to improve data quality. Normalization involves scaling numeric values to a common scale without distorting differences in the ranges of values, which helps prevent biases in the model due to the inherent scale of features. Feature extraction involves transforming the input data into a set of usable features, possibly reducing the dimensionality of the data in the process. For instance, in image analysis, feature reduction techniques such as Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), t-Distributed Stochastic Neighbor Embedding (t-SNE), autoencoders, and feature selection can be used for simplifying images, improving model performance, and gaining insights into the underlying structure of the images. These techniques not only help in reducing the computational load on the model but also in mitigating issues like overfitting by simplifying the data without losing critical information.
[0048] In the instance that machine learning pipeline 200 is used for supervised or semisupervised learning of machine learning models, labeling techniques can be implemented as part of the data collection. The quality and accuracy of data labeling directly influences the model's performance, as labels serve as the definitive guide that the model uses to learn the relationships between the input features and the desired output. Particularly in complex domains such as image recognition, natural language processing, or medical diagnosis, precise and consistent labeling is important because it provides the ground truth or target outcomes against which the model's predictions are compared and adjusted during training. Effective labeling ensures that the model is trained on correct and clear examples, thus enhancing its ability to generalize from the training data to real-world scenarios. In some instances, the annotation labels and ground truth values (labels) are appended or annotated within the raw data. For example, when the raw data includes text scripts, the labels may include one or more spans and corresponding named entities.
[0049] Labeling techniques can vary significantly depending on the type of data and the specific requirements of the project. Manual labeling, where human annotators label the data, is one method that can be used. This approach is useful when a detailed understanding and judgment are required, such as in labeling medical images or categorizing text data where context and subtlety are important. However, manual labeling is time-consuming and prone to inconsistency, especially with many annotators. To mitigate this, semi-automated labeling tools may be used as part of data subsystem 205 to pre-label data using algorithms, which human annotators may then review and correct as needed. Another approach is active learning, a technique where the model being developed is used to label new data iteratively. The model suggests labels for new data points, and human annotators may review and adjust certain predictions such as the most uncertain predictions. This technique optimizes the labeling effort by focusing human resources on a subset of the data, e.g., the most ambiguous cases, improving efficiency and label quality through continuous refinement. Once collected, generated, preprocessed, and / or labeled, the data may then be split into the training and validation datasets 210. The training and validation datasets 210 may comprise the raw data and / or the preprocessed data. The training and validation datasets 210 are typically split into at least three subsets of data: training, validation, and testing. The training set is used to fit the model, where the machine learning model leams to make inferences based on the training data. The validation set, on the other hand, is utilized to tune hyperparameters and prevent overfitting by providing a sandbox for model selection. Finally, the test set serves as a new and unseen dataset for the model, used to simulate real-w orld application and evaluate the final model’s performance. The process of splitting ensures that the model can perform well not just on the data it was trained on, but also on new-, unseen data, thereby validating and testing its ability to generalize.
[0050] Various techniques can be employed to split the data effectively, with each method aiming to maintain a good representation of the overall dataset in each subset. A simple random split (e.g., a 70 / 20 / 10%, 80 / 10 / 10%, or 60 / 25 / 15%) is the most straightforward approach, where examples from the data are randomly assigned to each of the three sets. However, more sophisticated methods may be necessary to preserve the underlying distribution of data. For instance, stratified sampling may be used to ensure that each split reflects the overall distribution of a specific variable, particularly useful in cases where certain categories or outcomes are underrepresented. Another technique, k-fold cross- validation, involves rotating the validation set across different subsets of the data, maximizing the use of available data for training while still holding out portions for validation. These methods help in achieving more robust and reliable model evaluation and are useful in the development of predictive models that perform consistently across varied datasets.
[0051] Data subsystem 205 is also used for collecting, generating, setting, or implementing model hyperparameters 240 for the training and validation subsystem 215. The hyperparameters control the overall behavior of the models. Unlike model parameters 245 that are learned automatically during training, hyperparameters 240 are set before training begins and have a significant impact on the performance of the model. For example, in a neural network, hyperparameters include the learning rate, number of layers, number of neurons / nodes per layer, activation functions, convolution kernel width, the number of kernels for a model, the number of graph connections to make during a lookback period, and the maximum depth of a tree in a random forest among others. These settings can determine how quickly a model leams, its capacity to generalize from training data to unseen data, and its overall complexity. Correctly setting hyperparameters is important because inappropriate values can lead to models that underfit or overfit the data. Underfitting occurs when a model is too simple to leam the underlying pattern of the data, and overfitting happens when a model is too complex, learning the noise in the training data as if it were signal.
[0052] Training, Validation, and Testing
[0053] The training and validation subsystem 215 is comprised of a combination of specialized hardware and software to efficiently handle the computational demands required for training, validating, and testing a machine learning model. On the hardware side, high- performance GPUs (Graphics Processing Units) may be used for their ability to perform parallel processing, drastically speeding up the training of complex models, especially deep learning networks. CPUs (Central Processing Units), while generally slower for this task, may also be used for less complex model training or when parallel processing is less critical. TPUs (Tensor Processing Units), designed specifically for tensor calculations, provide another level of optimization for machine learning tasks. On the software side, a variety of frameworks and libraries are utilized, including TensorFlow, PyTorch. Keras, and scikit- leam. These tools offer comprehensive libraries and functions that facilitate the design, training, validation, and testing of a wide range of machine learning models across different computing platforms, whether local machines, cloud-based systems, or hybrid setups, enabling developers to focus more on model architecture and less on underlying computational details.
[0054] Training is the initial phase of developing machine learning models 230 where the model leams to make predictions or decisions based on training data provided from the training and validation datasets 210. During this phase, the model iteratively adjusts its internal model parameters 245 to achieve a preset optimization condition. In a supervised machine learning training process, the preset optimization condition can be achieved by minimizing the difference between the model output (e.g., predictions, classifications, or decisions) and the ground truth labels in the training data. In some instances, the preset optimization condition can be achieved when the preset fixed number of iterations or epochs (full passes through the training dataset) is reached. In some instances, the preset optimization condition is achieved when the performance on the validation dataset stops improving or starts to degrade. In some instances, the preset optimization condition is achieved when a convergence criterion is met, such as when the change in the model parameters falls below a certain threshold between iterations. This process, known as fitting, is fundamental because it directly influences the accuracy and effectiveness of the model.
[0055] In an exemplary training phase performed by the training and validation subsystem 215, the training subset of data is input into the machine learning algorithms 220 to find a set of model parameters 245 (e.g., weights, coefficients, trees, feature importance, and / or biases) that minimizes or maximizes an objective function (e.g., a loss function, a cost function, a contrastive loss function, a cross-entropy loss function, an Out-of-Bag (OOB) score, etc.). To train the machine learning algorithms 220 to achieve accurate predictions, ‘'errors” (e.g., a difference between a predicted label and the ground truth label) need to be minimized. In order to minimize the errors, the model parameters can be configured to be incrementally updated by minimizing the objective function over the training phase ("optimization”). Various different techniques may be used to perform the optimization. For example, to train machine learning algorithms such as a neural network, optimization can be done using back propagation. The current error is ty pically propagated backwards to a previous layer, where it is used to modify the weights and bias in such a way that the error is minimized. The weights are modified using the optimization function. Other techniques such as random feedback, Direct Feedback Alignment (DFA), Indirect Feedback Alignment (IF A), Hebbian learning, and the like can also be used to update the model parameters 245 in a manner as to minimize or maximize an objective function. This cycle is repeated until a desired state (e.g.. a predetermined minimum value of the objective function) is reached.
[0056] The training phase is driven by three primary components: the model architecture (which defines the structure of the algorithm(s) 220), the training data (which provides the examples from which to leam), and the learning algorithm (which dictates how the model adjusts its model parameters). The goal is for the model to capture the underlying patterns of the data without memorizing specific examples, thus enabling it to perform well on new, unseen data.
[0057] The model architecture is the specific arrangement and structure of the various components and / or layers that make up a model. In the context of a neural network, the model architecture may include the configuration of layers in the neural network, such as the number of layers, the type of layers (e.g., convolutional, recurrent, fully connected), the number of neurons in each layer, and the connections between these layers. In the context of a random forest consisting of a collection of decision trees, the model architecture may include the configuration of features used by the decision trees, the voting scheme, and hyperparameters such as the number of trees in the forest, the maximum depth of each tree, the minimum number of samples required to split a node, and the maximum number of features to consider when looking for the best split. In some instances, the model architecture is configured to perform multiple tasks. For example, a first component of the model architecture may be configured to perform a feature selection function, and a second component of the model architecture may be configured to perform a feature scoring function. The different components may correspond to different algorithms or models, and the model architecture may be an ensemble of multiple components.
[0058] Model architecture also encompasses the choice and arrangement of features and algorithms used in various models, such as decision trees or linear regression. The architecture determines how input data is processed and transformed through various computational steps to produce the output. The model architecture directly influences the model's ability to leam from the data effectively and efficiently, and it impacts how well the model performs tasks such as classification, regression, or prediction, adapting to the specific complexities and nuances of the data it is designed to handle.
[0059] The model architecture can encompass a wide range of algorithms 220 suitable for different kinds of tasks and data types. Examples of algorithms 220 include, without limitation, linear regression, logistic regression, decision tree, Support Vector Machines, Naives Bayes algorithm, Bayesian classifier, linear classifier, K-Nearest Neighbors, K- Means. random forest, dimensionality reduction algorithms, grid search algorithm, genetic algorithm, AdaBoosting algorithm. Gradient Boosting Machines, and Artificial Neural Networks such as convolutional neural network (“CNN”), an inception neural network, a U- Net, a V-Net, a residual neural network (“Resnef '), a transformer neural network, a recurrent neural network, a Generative adversarial network (GAN), or other variants of Deep Neural Networks (“DNN”) (e.g., a multi-label n-binary DNN classifier or multi-class DNN classifier). These algorithms can be implemented using various machine learning libraries and frameworks such as TensorFlow, PyTorch, Keras, and scikit-leam, which provide extensive tools and features to facilitate model building, training, validation, and testing. For example, the classification algorithm described with respect to FIGS. 3 A and 3B is a set of algorithms configured as the architecture of a neural network comprised of layers of nodes, or "neurons," which are designed to mimic the way a human brain operates.
[0060] The learning algorithm is the overall method or procedure used to adjust the model parameters 245 to fit the data. It dictates how the model learns from the data provided during training. This includes the steps or rules that the algorithm follows to process input data and make adjustments to the model's internal parameters (e.g., weights in neural networks) based on the output of the objective function. Examples of learning algorithms include gradient descent, backpropagation for neural networks, and splitting criteria in decision trees.
[0061] Various techniques may be employed by training and validation subsystem 215 to train machine learning models 230 using the learning algorithm, depending on the type of model and the specific task. For supervised learning models, where the training data includes both inputs and expected outputs (e.g., ground truth labels), gradient descent is a possible method. This technique iteratively adjusts the model parameters 245 to minimize or maximize an objective function (e.g., a loss function, a cost function, a contrastive loss function, etc.). The objective function is a method to measure how well the model’s predictions match the actual labels or outcomes in the training data. It quantifies the error between predicted values and true values and presents this error as a single real number. The goal of training is to minimize this error, indicating that the model's predictions are, on average, close to the true data. Common examples of loss functions include mean squared error for regression tasks and cross-entropy loss for classification tasks.
[0062] The adjustment of the model parameters 245 is performed by the optimization function or algorithm, which refers to the specific method used to minimize (or maximize) the objective function. The optimization function is the engine behind the learning algorithm, guiding how the model parameters 245 are adjusted during training. It determines the strategy to use when searching for the best weights that minimize (or maximize) the objective function. Gradient descent is a primary example of an optimization algorithm, including its variants like stochastic gradient descent (SGD), mini-batch gradient descent, and advanced versions like Adam or RMSprop, which provide different ways to adjust learning rates or take advantage of the momentum of changes. For example, in training a neural network, backpropagation may be used with gradient descent to update the weights of the network based on the error rate obtained in the previous epoch (cycle through the full training dataset). Another technique in supervised learning is the use of decision trees, where a tree-like model of decisions is built by splitting the training dataset into subsets based on an attribute value test. This process is repeated on each derived subset in a recursive manner called recursive partitioning.
[0063] In unsupervised learning, where training data does not include labels, different techniques are used. Clustering is one method where data is grouped into clusters that maximize the similarities of data within the same cluster and maximize the differences with data in other clusters. The K-Means algorithm, for example, assigns each data point to the nearest cluster by minimizing the sum of distances between data points and their respective cluster centroids. Another technique, Principal Component Analysis (PCA), involves reducing the dimensionality7of data by transforming it into a new set of variables, the principal components, which are uncorrelated and ordered so that the first few retain most of the variation present in all of the original variables. These techniques help uncover hidden structures or patterns in the data, which can be essential for feature reduction, anomaly detection, or preparing data for further supervised learning tasks.
[0064] Validating is another phase of developing machine learning models 230 where the model is checked for deficiencies in performance and the hyperparameters 240 are optimized based on validation data provided from the training and validation datasets 210. The validation data helps to evaluate the model's performance, such as accuracy, precision, recall, or Fl -score, to gauge how well the training is ongoing, for example, by monitoring if an underfitting or overfitting is occurring. Hyperparameter optimization, on the other hand, involves adjusting the settings that govern the model's learning process (e.g., learning rate, number of layers, size of the layers in neural networks) to find the combination that yields the best performance on the validation data. One optimization technique is grid search, where a set of predefined hyperparameter values are systematically evaluated. The model is trained with each combination of these values, and the combination that produces the best performance on the validation set is chosen. Although thorough, grid search can be computationally expensive and impractical when the hyperparameter space is large. A more efficient alternative optimization technique is random search, which samples hyperparameter combinations from a defined distribution randomly. This approach can in some instances find a good combination of hyperparameter values faster than grid search. Advanced methods like Bayesian optimization, genetic algorithms, and gradient-based optimization may also be used to find optimal hyperparameters more effectively. These techniques model the hyperparameter space and use statistical methods to intelligently explore the space, seeking hyperparameters that yield improvements in model performance.
[0065] An exemplary validation process includes iterative operations of inputting the validation subset of data into the trained algorithm(s) using a validation technique such as K- Fold Cross-Validation. Leave-one-out Cross-Validation, Leave-one-group-out Cross- Validation. Nested Cross-Validation, or the like, to fine-tune the hyperparameters and ultimately find the optimal set of hyperparameters. In some instances, a 5-fold cross- validation technique may be used to avoid overfitting the trained algorithm and / or to limit the number of selected features per split to the square-root of the total number of input features. In some instances, training dataset is split into 5 equal-size cohorts (or about equal -size), and every four of the cohorts are used to train an algorithm to generate five models (e.g., cohorts #1, 2, 3, and 4 are used to train and generate model 1, cohorts #1, 2, 3, and 5 are used to train and generate model 2, cohorts #1, 2. 4, and 5 are used to train and generate model 3, cohorts #1, 3, 4, and 5 are used to train and generate model 4, and cohorts #2, 3, 4 and 5 are used to train and generate model 5). Each model is evaluated (or validated) using the unused cohort in the training (e.g., for model 5. cohort #1 is used for validation). The overall performance of the training can be evaluated by an average performance of the five models. K-fold cross- validation provides a more robust estimate of a model’s performance compared to a single training / validation split because it utilizes the entire dataset for both training and evaluation and reduces the variance in the performance estimate.
[0066] Once a machine learning model has been trained and validated, it undergoes a final evaluation using test data provided from the training and validation datasets 210, which is a separate subset of the data that has not been used during the training or validation phases. This step is crucial as it provides an unbiased assessment of the model's performance in simulating real-world operation. The test dataset serves as new, unseen data for the model, mimicking how the model would perform when deployed in actual use. During testing, the model’s predictions are compared against the true values in the test dataset using various performance metrics such as accuracy, precision, recall, Fl, AUC, and mean squared error, depending on the nature of the problem (classification or regression). This process helps to verify the generalizability of the model — its ability to perform well across different data samples and environments — highlighting potential issues like overfitting or underfitting and ensuring that the model is robust and reliable for practical applications. The machine learning models 230 are fully validated and tested once the output predictions have been deemed acceptable by user defined acceptance parameters. Acceptance parameters may be determined using correlation techniques such as Bland-Altman method and the Spearman’s rank correlation coefficients and calculating performance metrics such as the error, accuracy, precision, recall, receiver operating characteristic curve (ROC), etc. Inference Phase for Machine Learning Models
[0067] The inference subsystem 225 is comprised of various components for deploying the machine learning models 230 in a production environment (e.g., use in a liquid biopsy platform as described with respect to FIG. 1). Deploying the machine learning models 230 includes moving the models from a development environment (e.g., the training and validation subsystem 215, where it has been trained, validated, and tested), into a production environment where it can make inferences on real-world data (e.g.. input data 250). This step typically starts with the model being saved after training, including its parameters and configuration such as final architecture and hyperparameters. It is then converted, if necessary, into a format that is suitable for deployment, depending on the deployment environment. For instance, a model trained in a scientific computing environment such as Python might be converted into a Java-friendly format for integration into a larger enterprise application. Deployment can be conducted on various platforms, including on-premises serv ers, cloud environments like AWS, Azure, Google.
[0068] Once deployed, the model is ready to receive input data 250 and return outputs (e.g., inferences 255). In some instances, the model resides as a component of a larger system or service (e.g., including additional downstream applications 235). In some instances, the models 230 and / or the inferences 255 can be used by the downstream applications 235 to provide further information. For example, the inferences 255 can be used to aid qualified personnel (e.g., oncologists) to help diagnosis and / or determine whether treatment should be administered to a patient. In some instances, the inferences 255 can be used to aid qualified personnel to determine a specific type of treatment to administer to a patient based on the inference results. The downstream applications can be configured to generate an output 260. In some instances, the output 260 comprises a report including inferences 255 and information generated by the downstream applications 235.
[0069] In an exemplary inference subsystem 225, the input data 250 includes differentially expressed nucleic acids, more specifically, differentially expressed cell-free RNAs. The differentially expressed cell-free RNAs may be identified by performing one or more sequencing methods on one or more biological samples collected from subjects using the liquid biopsy platform 115 as described with respect to FIG. 1. The one or more biological samples may be a single sample or multiple samples collected from a single subject, or single samples collected from multiple subjects. To manage and maintain its performance, a deployed model may be continuously monitored to ensure it performs as expected over time. This involves tracking the model’s prediction accuracy, response times, and other operational metrics. Additionally, the model may require retraining or updates based on new data or changing conditions in the environment it is applied in. This can be useful because machine learning models can drift over time due to changes in the underlying data they are making predictions on — a phenomenon known as model drift. Therefore, maintaining a machine learning model in a production environment often involves setting up mechanisms for performance monitoring, regular evaluations against new test data, and potentially periodic updates and retraining of the model to ensure it remains effective and accurate in making predictions.
[0070] Training One or More Machine Learning Models to Predict a Disease Status of a Subject
[0071] FIG. 3 is a flowchart illustrating a process 300 for training one or more machine learning models to predict a disease status of a subject. The processing depicted in FIG. 3 may be implemented in software (e.g.. code, instructions, program) executed by one or more processing units (e.g., processors, cores) of the respective systems, hardware, or combinations thereof. The software may be stored on a non -transitory storage medium (e.g., on a memory' device). The method presented in FIG. 3 and described below is intended to be illustrative and non-limiting. Although FIG. 3 depicts the various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In certain alternative embodiments, the steps may be performed in some different order, or some steps may also be performed in parallel. In certain embodiments, such as in the embodiments depicted in FIG. 1, the processing depicted in FIG. 3 may be executed as part of the liquid biopsy platform 115 of the computing environment 100 as described in FIG. 1.
[0072] At step 305, a biological sample comprising cell-free ribonucleic acids (RNAs) is collected from a subject. The biological sample can be a cell-containing liquid cell-free RNAs. The sample can include, but is not limited to, amniotic fluid, blood, plasma, prenatal blood, blood cells, peritoneal fluid, amniotic fluid, pleural fluid, urine, saliva, semen, serum, ex vivo cell culture media, or in vitro cell culture media. Methods of obtaining a sample include but are not limited to aspirations, swabs, drawing blood or other fluids, and the like. In various embodiments, the sample is a blood sample, more preferably a plasma sample collected from the subject either having one or more diseases or without disease. To collect plasma, a whole blood sample can be collected from the subject using venipuncture of other routine methods known in the art. Plasma is separated from the whole blood sample by adding an anticoagulant to the blood sample and centrifuging the blood sample at sufficient speed to separate the plasma from the blood cells.
[0073] After a biological sample (e.g., blood or plasma) is collected, the cell-free RNAs may be isolated from the biological sample. Various methods are know n in the art for isolating cell-free RNAs from a sample, such as using a specialized cell-free RNA reagent kit (e.g., reagents 145 in the liquid biopsy platform 115) or by experimental means. A cell-free RNA reagent kit can include tubes, filters, and cell-free RNA extraction reagents specific to a manufacturer. Experimental methods for isolating / extracting cell-free RNA from a sample involve disruption and lysis of the starting material follow ed by the removal of proteins, DNA, and other contaminants and finally recovery (e.g., alcohol precipitation) of the cell-free RNAs. Following isolation, the isolated cell-free RNAs are reverse transcribed into single stranded complementary DNA (cDNA) using amplification methods known in the art. Amplification refers to production of additional copies of a nucleic acid sequence and is generally carried out using polymerase chain reaction (PCR) or other technologies for increasing concentration of a segment of a nucleic acid sequence.
[0074] As used herein, cell-free RNAs refers to polymers of single-stranded RNAs. Unless specifically limited, the RNAs encompass ribonucleic acids containing known analogues of natural nucleotides that have similar properties as the reference ribonucleic acid. Ribonucleic acids also encompass all forms of sequences including, but not limited to. singlestranded forms, double-stranded forms, hairpins, stem-and-loop structures, and the like. Cell- free RNAs (cell-free RNAs) are full-length or fragments of RNA molecules that circulate outside of cells, typically in bodily fluids such as blood plasma, serum, urine, or cerebrospinal fluid. Unlike cellular RNAs, which reside within intact cells and participate directly in gene expression and regulation, cell-free RNAs are released into the extracellular environment through various biological processes, including cell death (apoptosis or necrosis), active secretion via extracellular vesicles like exosomes, or cellular stress and injury’. A population of cell-free RNAs is highly diverse and can comprise full-length or fragments of messenger RNAs (mRNAs). which are transcribed from genes and carry protein-coding information, as well as various non-coding RNAs such as microRNAs (miRNAs), long non-coding RNAs (IncRNAs), circular RNAs (circRNAs), repetitive RNAs, unannotated RNAs, and small RNA species. Cell-free RNAs are derived from the transcriptional activity of their tissue or cellular origin, making them valuable as minimally invasive biomarkers for monitoring physiological and pathological states, such as cancer, pregnancy, organ transplantation, or other diseases.
[0075] As used herein, the terms “individual,’’ “patient,” and “subject” are used interchangeably. In certain embodiments, subjects are “patients,” i.e., living humans that are receiving medical care for a disease or condition. This includes persons with no defined illness who are being investigated for signs of a particular disease or condition or who have no other reported comorbidities. In any of the methods set forth herein, the subject can be suspected of having one or more diseases / conditions, diagnosed with one or more diseases / conditions, or receiving treatment for one or more diseases / conditions.
[0076] In some instances, the subject can have a negative disease status, a pre-disease status, or a positive disease status. A negative disease status is indicative of a subject who is negative for the disease of interest; however, the subject may still have other disease comorbidities. A pre-disease status refers to a physiological or clinical state in which an individual does not yet meet the formal diagnostic criteria for a particular disease, but exhibits certain biomarkers, risk factors, or subclinical changes that indicate an increased likelihood of developing that disease in the future. A positive disease status refers to a physiological or clinical state in which an individual does meet the formal diagnostic criteria for a particular disease. In some embodiments, the positive disease status indicates the subject having one or more diseases (or conditions) selected from a list comprising a cancerous disease, a precancerous condition, a neurodegenerative disease, an infectious disease, a genetic disease, an autoimmune disease, a metabolic disease, a cardiovascular disease, a neurological / psychiatric disease, a respiratory7disease, an endocrine disease, a digestive disease, a musculoskeletal disease, a skin disease, a congenital disease, a blood disease, or any combination thereof. In some embodiments, the positive disease status indicates the cancerous disease.
[0077] A subject with a cancerous disease status may suffer from one or more types of cancer simultaneously. For example, the subject can have pancreatic, kidney, renal, pelvic, colorectal, stomach, thymic, head and neck, mesothelial, prostate, cervical, thyroid, adrenal, testicular, breast, uterine, bone, esophageal, lung, liver, bile duct, ovarian, bladder, nervous system, and / or skin cancer. In some embodiments, the cancerous disease is esophageal cancer. Accordingly, the pre-disease status indicates the precancerous condition, and in a more preferred embodiment, the precancerous condition is Barrett’s Esophagus. The negative disease status indicates the subject is negative for precancerous Barrett’s Esophagus and esophageal cancer; however, the subject may still have other disease comorbidities. In some instances, subjects who are at risk of or have just been diagnosed with one or more diseases / conditions (e.g., cancer or precancer) are likely to have not yet received treatment(s). In other instances, a subject having been diagnosed with one or more diseases / conditions is receiving treatment(s).
[0078] At step 310, the cell-free RNAs in the biological sample are detected using a sequencing method, wherein the sequencing method generates sequencing read data for cell- free RNA in the biological sample. The sequencing method is performed using a sequencing device (e.g., instrumentation 140 in the liquid biopsy platform 115). In some instances, the sequencing method identifies the full sequence, a substantially full sequence, and / or a partial sequence of the cell-free RNAs in the biological sample. In a preferred embodiment, the sequencing method identifies the full length of the cell-free RNAs in the biological sample. The sequencing method comprises a nanopore sequencing method, a sequencing by expansion method, a long-read sequencing method, a single-molecule sequencing method, a sequencing by synthesis method, or a short read sequencing method. In some embodiments, the sequencing method is the long-read sequencing method. In some embodiments, detecting the cell-free RNAs may optionally be performed using RT-PCR, a probe-based target capture method with or without subsequent sequencing, a primer-based target enrichment method with or without subsequent sequencing, or a CRISPR-based detection method.
[0079] The sequencing method described above generate a large number of reads. As used herein, '‘reads” (e g., “a read,” “a sequence read”) are short nucleotide sequences produced by any sequencing process described herein or known in the art. Reads can be generated from one end of nucleic acid fragments (“single-end reads”), or they can be generated from both ends of nucleic acid fragments (e.g.. paired-end reads, double-end reads). The length of a sequence read is often associated with the particular sequencing technology. High-throughput methods, for example, provide sequence reads that can vary in size from tens to thousands of base pairs (bp). Sequencing reads may have a mean, median, average, or absolute length of about 15bp to about 2000bp. For example, sequencing reads may be about 15bp, 16bp. 17bp, 18bp, 19bp, 20bp, 25bp. 50bp. lOObp. 150bp. 200bp, 300bp, 400bp, 500bp, 600bp, 700bp. 800bp, 900bp, lOOObp, or about 2000bp or about any integer value between 15bp and 2000bp. The raw7sequencing read data and their associated metadata are stored in FASTQ, FASTA or POD5 files. In addition to sequencing read data, the sequencing method can generate other types of sequencing data including alignment files (e.g..BAM or SAM files) which store the results of mapping sequencing reads to a reference genome, detailing the position, alignment quality, and count of each read. Moreover, sequencing data can include variant files, such as VCF files, which contain information about detected genetic variants, including their ty pe, location, and annotations. To obtain alignment files, VCF files, and other processed sequencing data files, various bioinformatic software (e.g., software 150 in the liquid biopsy platform 115) tools may be used.
[0080] At step 315, a subset of cell-free RNAs from the cell-free RNAs in the biological sample are identified by annotating the sequencing read data for the cell-free RNAs to sequencing read data for a standard reference trans criptome. Optionally, the identification of the subset of cell-free RNAs from the cell-free RNAs in the biological sample are not found in the standard reference transcriptome. Thus, for example, a subset of cell-free RNAs from the cell-free RNAs in the biological sample can be identified by annotating the sequencing read data for the cell-free RNAs that are not present in sequencing read data from a standard reference transcriptome. Read annotation is performed by various bioinformatics software (e.g., software 150 in the liquid biopsy platform 115) such as IsoQuant as a nonlimiting example. The subset of cell-free RNAs can comprise those cell-free RNAs having no annotation data in the standard reference transcriptome. The standard reference transcriptome is a rigorously curated and comprehensive collection of RNA transcripts — including all known coding and non-coding RNAs — for a given organism, characterized by high accuracy and completeness. It serves as the definitive reference for gene expression analysis, transcript annotation, and RNA sequencing validation, and may be sourced from publicly available databases (such as GENCODE, RefSeq. Ensembl. or RepeatMasker). proprietary in-house curation, or tailored to a particular disease or condition. The standard reference transcriptome encompasses detailed genomics data and annotations for known transcripts, including transcript isoforms, gene coordinates, functional and biotype classifications (e.g., proteincoding. IncRNA. pseudogenes), sequence information, and regulatory’ elements for the organism of interest (such as human, mouse, drosophila, yeast, or zebrafish).
[0081] At step 320, the subset of cell-free RNAs are combined with the standard reference transcriptome to generate a custom reference transcriptome. The combination is based on the annotated sequencing read data for the identified subset of cell-free RNAs. Optionally, the standard reference transcriptome (e.g.. GENCODE) is concatenated with known repeat element sequences (e.g., RepeatMasker) prior to combining the cell-free RNAs to generate the custom reference transcriptome. In various embodiments, the custom reference transcriptome comprises transcripts and their corresponding metadata (e.g., annotations, transcript isoforms, gene coordinates, functional and biotype classifications, sequence information, etc.), wherein the transcripts comprise the cell-free RNAs, the standard reference transcriptome transcripts, and the repeat-derived transcripts.
[0082] At step 325, transcript abundance for the custom set of cell-free RNAs in the custom reference transcriptome are quantified. In some embodiments, the quantification is done using another sequencing method. In some embodiments, the detecting sequencing method and the quantification sequencing method are the same. In some embodiments, the detecting sequencing method and the quantification sequencing method are long-read sequencing methods. In some embodiments, the detecting sequencing method and the quantification sequencing method are different. In some embodiments, the detecting sequencing method is a long-read sequencing method and the quantification sequencing method is a short-read sequencing method.
[0083] To quantify the transcript abundance for the custom set of cell-free RNAs in the custom reference transcriptome, the sequencing read data are analyzed using specialized bioinformatics tools (e.g., Salmon). These tools compare segments (e.g., predefined k-mer length of the transcript reads) against a genomic reference genome (e.g., the human reference genome HG38) to estimate the likelihood that each read originates from specific transcripts. Transcript abundance is calculated using statistical methods such as expectationmaximization (EM) algorithms, bias correction, bootstrapping, the like, or any combination thereof.
[0084] At step 330, a list of differentially expressed cell-free RNAs are identified by comparing the transcript abundance of the custom set of cell-free RNAs in the custom reference transcriptome to cell-free RNA transcript abundance values from healthy subjects. Differentially expressed cell-free RNAs are transcripts showing significant differences in transcript abundance between conditions (e.g., healthy vs. disease), biological states (e.g., pre-disease vs. disease), or groups. As described herein, the term ■■significant’’ refers to "statistical significance," indicating that an observed effect, relationship, or difference in data is unlikely to have occurred by chance alone, according to a predetermined threshold (e.g., a p-value, adjusted p-value, q-value less than or equal to 0.05, or 0.01). In some embodiments, differentially expressed cell-free RNAs indicate the disease status of the subject. In some embodiments, the list of differentially expressed cell-free RNAs distinguish healthy (i.e.. negative disease status) versus precancerous Barretf s Esophagus, heathy versus esophageal cancer (i.e., positive disease status), and healthy versus disease (e.g., precancerous Barrett’s Esophagus and esophageal cancer). In various embodiments, the list of differentially expressed cell-free RNAs comprise all or a subset thereof: *chr22_KI270877vl_alt_2, ENSG00000274383, MT-ND1, MT-ATP6, *chr5_23822. MT-RNR1, MT-ND4L, MT-C02, MTATP6P1. MT-C03, MT-ND4. MT-C01, MT-ND3. *chrl_28278, MT-ND2. MT-RNR2, MT-ND5, MT-CYB, MT-ATP8, RPL23AP7, *chrl5_l 1653, *chrl5_l 1661, MTND2P28, MT-ND6, MTND1P23, HHEX, ENSG00000280064, *chrl5_11709, *chrl7_7889, FRAT2, ITGB3BP, S100A16, PPBP, DYRK4, TREML1, *chrl_37700. *chr6_GL000253v2_alt_80, H2BC4, PSMD10. DERA. MAP3K7CL, SCOC, *chrX 8889, *chrl7 23634, DEFA4, *chr3_27456, GP1BB, F13A1, RAB24, RAB27B, CDK1 IB, PLEKHF2, 0RMDL1, PGRMC1, SMCO4, MS4A1, ABO, SPARC, IRF8, PIKA API, AT0X1, OST4, PCNA, BTG2, IGLC2, GNG11, MRPL27, COA3, SMC6, PRKAR2B, H2AC6, VPREB3, IGLC3, NDUFV2, ZRANB2, EIF6, NGDN, MRPL47, *chrl0_23323, MRPL13, GOLPH3, *chr8_13328, DPM2, SAE1, IGKV1-5, CAPG, IGLC1, FRA10AC1, IGKC, ENSG00000288869, CAVIN2, ENSG00000288796, UTP18, CEBPD, *chrl0_36630, GCH1, RAB5IF, HSPA1A, MFNG, AP1B1, FABP4, UBE2Q1, ISCU, RGS2, MNDA, SELENOP, TUBB4B, AL0X5, FCER2, ESD, CA2, FAM89B, ZNF217, USO1, GCA, PYM1. ZCCHC17, MRPL51, TUBA1B, PF4. DEF A3, HSPB1, GAPT, DHX29. NCF4, MYL9, TAF7, MT1X, ZFAND2B, ENSG00000251487, ID3, LTA4H, IFIT2, *chrl 7_17297, CALHM6, ARMCX3, ANXA3, MFSD1, NADK, UBE2Q2, SRP9, *chrl_2906, STARD3NL, TERF2IP, MRPS25. TXN, RAB1A, RNF20, ID2, SNX3, CTNNBL1, NMI, C4orf3, NRDC, HNRNPAB, GABARAPL2, UBE2J1, SEC11C, NBPF14, EIF4G2, FLII, LSM6, D0K2, PRKACB, PNRC2, POLR2J, PDCD10, UBE2N, MTCH1, SRA1, RALB, SLC9A3R2, LAMT0R3, RPA2, DEFA1, CLU, GIMAP6, CTDSP1, HMGN3, ATP6V1B2, LYRM1, CASP1, Clorfl62, YPEL5, COX17, NDUFA4, ZNF438, MTCH2, DNAJC15, PTRHD1, IGHM, CAPZA1. S100A12, IRAG2, SLA. JCHAIN, HSPE1, RAB2A, TALDO1, SF3B6. GTF2F2, E2F4, NIFK, PSMB8-AS1. SIPA1. PLAC8, VPS26A, PRDX5, PSME2. MRPL20, LINC01857, TPI1, MRPL33, SNX6, TMEM126B, WARSI, IFI16, RBBP8, ENSG00000233968, UBE2A, ARHGDIB, ARPC1A, BLOC1S1, TMSB4X, RPS28, NAA10, MYD88, SERPINB1, CLC, *chr2_46125, FBXL5, TMEM258. DDT, LCP2, CASP4, SMIM30, PICALM, *chr2 36264, PSMA7. UBE2F, MYL12A, TAF10, LYL1, UCP2, SNRPD3, SDHD, TMEM219, PTGS1, SAP18, FTL, VAMP3, RHOC, COP1, CARD16, ATP6V1F. COX6A1, HSBP1. RMND5B. TPRKB, CDC26, MMD, DYNLL1. CHCHD1, DNAJA1, UBXN1, SEC61B, NCF2, SLC25A5, CSTB, TBCB, *chr8_26652, BRK1, TRAPPCI, PEA15, NDUFB3, CTDNEP1, PSMA1, ECHI, PRDX1, BLOC1S2, PLPBP, RPL9, KANSL1-AS1, *chr22_2098, GIMAP4, SNRPD2, MGP, SCP2, C2orf49, COX20, ARPC2, SNRPG, NAA38. DCK, GLIPR2, RBX1, *chr3_27583, SRP14, TCL1A, RNF11, ATP5PO, C0X5A, *chr5_20581, FAM117A. ACTR2, PDCD5, NDUFB4, TXNIP, EIF2S2, SEPTIN11, ANP32A, CEP20, COX6C, NAP1L4, SERP1, PTMA, PAK1, HSPA8, ANAPC11, ARPC3, DPYSL2, SRSF9, RBM17, ASXL2, H4C3, H2AZ1, CALM2, Clorfl l5, MICOSIO, TCP1, RABI 1 A, NDUFC1, GDI2, YWHAQ, SUMO3, *chr5_33000, SH3GLB1. PFN1, GNG5, USB1, ATP5MK, SNRPF. *chr20 11440, ACYP2, MTMR12, AIF1L, AFTPH, C22orf39, TBCA, FABP5, ANXA2, PXN, SDHB, RNH1, ATP6V1E1, BECN1, GRB2, PSME3IP1, YY1, RAB7A, TRMT112, TOMM5, HLA-DPB1, WIPI1, RAB I 3. UBL5, HSP90AA1, H2AZ2, CHMP3, PGK1, PLCG2, MS4A7, *chr6_GL000255v2_alt_118, MTURN, TUBB, HNRNPA2B1. UQCRQ, MBD2, ENSG00000228021, APIP, SPIDR, COMMD6, APAF1, ACTG1, TFG, CAPZB, LBH, RNF114, *chr22_8633, EPAS1, CNBP, PPIA, PPP1CA, *chr3_29499, SNRPB, MRPL3, FDPS, CRLF3, NME2, IGSF6, ENSG00000226471, OSGEP, DYNC1LI1, *chr7_3815, SLIRP. PURA, ZCRB1, CRPPA, TPM3, YWHAB, *chrl9_792, *chrl_40462, GPR174, UQCRB. LDHA, NFATC2IP. DNAJC2, ARPC5. *chrX 10000, ENSG00000286896.
[0085] SUMO1 , NUP214, MTSS1 , VCL, RFLNB, *chr7_8482, *chr2_l 1817, ENSG00000279440, NRBP1, HNRNPC, CARD8, OXAIL, SLC25A3, WNK1, *chrl_61255, *chr3_13866, UBXN8, *chr4_16544, *chrl7_25614, *chr4_22561, *chr20_14359, *chrl2_238, *chr4_33244, ENSG00000271109, *chrl6_9155, LINC00491, RASA4B, *chrl_34706, RASEF, *chr20_6072, *chr3_44232, *chrl l_15696, *chr3_6460, *chr9_27845, ENSG00000284607, DNM1P47, ENSG00000288531, *chrl7_1739, *chrl_4734, CASC9, *chr4_22712, CCN5, DMXL1, *chrl9_14862, *chrl9_4729, and CXCL8. (*) indicates novel cell-free RNAs.
[0086] At step 335, at least the list of differentially expressed cell-free RNAs are used as features to train one or more machine learning models to predict the disease status of the subject (e.g., artificial intelligence models 155 in the liquid biopsy platform 115). As previously described, machine learning is a subset of artificial intelligence that uses algorithms and statistical models to make predictions based on input data features. In various instances, machine learning occurs in a computer environment (see computer environment 100 as discussed with respect to FIG. 1) using any of the training, validation, and inference approaches described with respect to FIG. 2. The one or more machine learning models leverage machine learning algorithms (e.g., algorithms 220 described with respect to FIG.2) that can have either (i) high interpretability and low accuracy to model linear relationships, (ii) low interpretability and high accuracy to model nonlinear relationships, or (iii) high interpretability and high accuracy to model both linear relationships and non-linear relationships. Examples of machine learning algorithms with high interpretability and low accuracy include, without limitation, k-nearest neighbors, decision trees, linear regression, logistic regression, and classification rules. Examples of machine learning algorithms with low interpretability and high accuracy include, without limitation, deep neural networks (DNN), graph neural network (GNN), and support vector machine (SVM). Examples of machine learning algorithms with high interpretability and high accuracy include Energy - Based Models (EBMs). In various embodiments, the one or more machine learning models comprise a regression model, a classification model, a clustering model, a statistical model, a decision tree, an ensemble model, a Bayesian model, a deep learning model, or any combination thereof. In some embodiments, at least one of the one or more machine learning models is a logistic regression model.
[0087] In the context of machine learning models, interpretability refers to the ability to understand and explain how a model generates its predictions or decisions, while accuracy measures the model's performance in correctly predicting outcomes or classifying data. High accuracy is important for reliable predictions, but it often comes at the expense of interpretability, especially in complex black box models like deep neural networks, which operate with intricate layers and parameters that are not easily understandable by humans. As such, black box models may optionally implement explainable Al (XAI) techniques to provide insights into their decision-making processes. Examples of XAI techniques include, without limitation, SHapley Additive exPlanations (SHAP), gradient based approaches such as integrated gradients, back propagation approaches such as DeepLIFT, model agnostic techniques such as Local Interpretable Model-agnostic Explanations (LIME), neural network and attention weight approaches such as Attention-Based Neural Network Models, and Deep Taylor Decomposition approaches such as Layer-wise Relevance Propagation (LRP). On the other hand, glass box models, such as EBMs, are inherently interpretable, allowing users to inspect the model's parameters and understand the relationships between inputs and outputs directly. While glass box models may sometimes sacrifice accuracy for simplicity', their transparency makes them highly valuable in applications requiring both trust and accountability.
[0088] To train the one or more machine learning models, iterative operations to adjust a set of model parameters to minimize a loss or error function of the one or more machine learning models are performed. The loss or error function is configured to measure the difference between the output predictions of the one or more machine learning models (e.g., a negative disease status, a pre-disease status, or a positive disease status for a subject) and the ground truth dataset. More specifically, during training, the one or more machine learning models leam to associate patterns of transcript abundance with disease status based on the input training dataset (i.e., the list of differentially expressed cell-free RNAs). The performance of the one or more models is assessed using cross-validation or a separate test dataset, with metrics such as accuracy, sensitivity, specificity, ROC-AUC, and Fl -score being evaluated. In some instances, a subset of the most important features (e.g., cell-free RNAs) for influencing disease status prediction are determined using feature importance scores, coefficients, SHAP / LIME values, etc.
[0089] By way of example, two or more of the differentially expressed cell-free RNAs listed in Table 1 can be used to distinguish healthy (i.e., negative disease status) versus precancerous Barrett’s Esophagus, heathy versus esophageal cancer (i.e., positive disease status), and healthy versus disease (e.g.. precancerous Barrett’s Esophagus and esophageal cancer). In some cases, the expression level of a given RNA is elevated in the precancerous stage, the esophageal cancer stage, or both as compared to control. In some cases, the expression level of a given RNA is reduced in the precancerous stage, the esophageal cancer stage, or both as compared to control. Optionally, all the differentially expressed cell-free RNAs or various subsets of the differentially expressed cell-free RNAs are analyzed. For example, the whole transcriptome of a blood or plasma sample could be sequenced in which all or substantially all of the differentially expressed cell-free RNAs that indicate the precancerous status (i.e., the pre-disease status), the esophageal cancer status (i.e., positive disease status), or distinguishes between precancerous status and esophageal cancer status are detected and quantified. As used herein, all or substantially all of the differentially expressed cell-free RNAs include for example, at least 80%, 85%, 90%, 95%, or 99% of the identified differentially expressed cell-free RNAs. One of skill in the art will appreciate that the pattern of expression of a certain subset of the differentially expressed cell-free RNAs in an identified subset correlates more closely with a disease status than other subsets, and, depending upon the desired accuracy, a critical subset of differentially expressed cell-free RNAs can be analyzed using, for example, a targeted gene panel. Limiting the analysis to the critical subset of the identified differentially expressed cell-free RNAs may be particularly useful in quick screening assays, whereas the full set or substantially all of the differentially expressed cell-free RNAs in the subset of transcripts would be used for a robust analysis.
[0090] In some instances, the precancerous Barrett’s Esophagus status is predicted by the one or more machine learning models based on the list of differentially expressed cell-free RNAs comprising at least two, three, four, five, six, seven, eight, nine, ten or more of the following cell-free RNA transcripts: *chr22_KI27087, MT-ND1, MT-ATP6. MT-RNR1, MT-C02, MT-C01, MT-C03, MTATP6P1, MT-ND4L, MT-RNR2. MT-ND4, MT-ND3. PVALB, MT-ND5, MT-ND2, *chrl2_8540, MT-CYB, MT-ATP8, *chr5_23822, *chrl_28278, ENSG00000274383, MTND2P28, MT-ND6, *chrl5_l 1639, MTNDIP23, ENSG00000266401, *chr6_9458, AIM2, CXCL8, *chr3_0052, *chr7_24810, ZCH312A, *chrl l_2126, *chr21_7097, *chr!7_23338, *chr!3_18256. *chrl_48852, *chr5_14282, *chrl2_19620, *chr6_26573, RPSAP4, *chrl 1_29228, *chr3_41901, *chr4_4793, DLXL ENSG00000288980, *chrl 1_29228, *chr4_42463, *chr3_564, *chrl0_28310, *chrl7_14620, *chrl7_10248, *chr3_15268, AIRN, *chrl4_6471, ANKRD11P1, or any combination thereof. See Table 1 and (*) indicates novel cell-free RNAs. Optionally, the precancerous Barrett’s Esophagus status is predicted based on a certain subset of the differentially expressed cell-free RNAs, wherein the certain subset comprises: L3MBTL1 , DLX1, *chr6_29343, *chr!2_12317, *chr5_31003, *chr20_8731, NADK, RPSAP4, AL0X5, *chr20_12853, *chr2_41140, *chrl3_19893, and *chr21_3060.
[0091] In some instances, the one or more machine learning models predict the esophageal cancer status based on the list of differentially expressed cell-free RNAs comprising at least two, three, four, five, six, seven, eight, nine, ten or more of the following RNAs: *chrl2_8540, *chr5_14206, ENSG00000274383, *chr5_23822, *chrl_28278, MT-C03, MT-C02, MT-ATP6, MT-CYB. *chr!0_36912, MT-ATP8, MT-ND4, MT-ND3, MT-ND1, *chrl5_11661, *chrl6_5531, MT-ND4L. MT-ND2, MT-ND6. MT-TQ, MT-C01. MT-ND5, *chrl5_l 1639, GNG8, MT-RNR2, MTATP6P1, *chr6_18454, HHEX, MTND2P41, *chr4_36834, *chr4_37169, *chr6_15475, *chr8_23641, *chr6_16530, *chr21_3639, *chrl_32054, *chr3_42038, *chr!8_6774, *chr!9_15033, LINC01487, *chr!2_15574, *chrl5 21741. EIF4EP1, *chrl 38898, *chrlO 35990. *chrl2 7067, *chrl2 25273, *chrl2_26920, *chrl5_6605, *chrl0_32620, *chr9_1146, RPL7AP10, *chrl2_10146, *chrl2_1619, *chr2_38235, *chr5_131. or any combination thereof. See Table 1 and (*) indicates novel cell-free RNAs. Optionally, the esophageal cancer status is predicted based on a certain subset of the differentially expressed cell-free RNAs, wherein the certain subset comprises: *chrl_34706, *chr7_7352, *chr9_29702, *chrl6_5531, *chr20_2254, *chr6_11543, ENTR1P1, *chr20_12853, *chr3_14400, PPP2CA, MRPL20-DT, *chr8_33134, ENSG00000286896, CASC9. *chr9_6537. DYNC1LI1, *chr8_24909, *chrl_12702, ENSG00000268047, *chr7_452, *chrl_49351, BHLHE40, *chrl6_943, *chr2_38235, *chr4_40238, *chr4_6748, *chr2_34817, and *chr4_29780.
[0092] In some instances, the one or more machine learning models prediction distinguish disease (e.g.. precancerous status and esophageal cancer status) from the negative disease status (i.e., healthy) based on the list of differentially expressed cell-free RNAs comprising at least two, three, four, five, six, seven, eight, nine, ten or more of the following cell-free RNAs: *chr22_KI27087, ENSG00000274383, MT-ND1, MT-ATP6, *chr5_23822, MT- RNR1, MT-ND4L, MT-C02, MTATP6P1, MT-C03. MT-ND4, MT-C01. MT-ND3, *chrl_28278, MT-ND2, MT-RNR2, MT-ND5, MT-CYB, MT-ATP8, RPL23AP7, *chrl5_11653, *chrl5_l 1661, MTND2P28, MT-ND6, MTND1P23, HHEX, ENSG00000280064, *chrl5_11709, CXCL8, *chrl9_4729, *chrl9_14862, DMXL1, CCN5, *chr4_22712, CASC9, *chrl_4734, *chrl7_1739, ENSG00000288531, DNM1P47, ENSG00000284607, *chr9_27845, *chr3 6460, *chrl 1 15696, *chr3_44232. *chr20_6072, RASEF, *chrl_34706, RASA4B, LINC00491, *chrl 6_9155, ENSG00000271 109, *chr4_33244, *chr!2_238, *chr20_14359, *chr4_22561, *chrl7_25614, or any combination thereof. See Table 1 and (*) indicates novel cell-free RNAs. Optionally, distinguishing between disease and healthy subjects is predicted based on a certain subset of the differentially expressed cell-free RNAs, wherein the certain subset comprises: PXN, *chr3_13866, *chrl l_15696, IGHM, ENSG00000284607, CXCL8, VAMP3, *chr4_22561, *chrl2_238, *chr20_14359, OSGEP, ENSG00000286896, DYNC1LI1, *chr20_6072, and *chr9_27845.
[0093] Table 1:
[0094] * Indicates novel gene.
[0095] Surprisingly, mitochondrial RNAs are also significantly differentially expressed in the precancerous Barrett’s Esophagus status, the esophageal cancer status, or both. By way of example, the differentially expressed mitochondrial RNAs listed in Table 2 can be used to distinguish healthy (i.e., negative disease status) versus precancerous Barret’s Esophagus, heathy versus esophageal cancer (i.e., positive disease status), and healthy versus disease (e.g., precancerous Barret’s Esophagus and esophageal cancer).
[0096] In some instances, the precancerous Barrett's Esophagus status is predicted by the one or more machine learning models based on the list of differentially expressed mitochondrial RNAs comprising at least two, three, four, five, six, seven, eight, nine, ten or more of the following mitochondrial RNAs: ENSG00000280064, MT-ND1, MT-RNR1, MT- ATP6, MT-C02. MT-C01, MT-C03, MTATP6P1,MT-RNR2, MT-ND4, MT-ND4L, MENDS. MT-ND2, MT-ND5, MT-CYB, MT-ATP8, ENSG00000274383, MTND2P28. MT- ND6, ENSG00000266401, MT-TQ, MTND1P23, TOMM34, MT-TS 1, AIM2, ITGB3BP, H2BC4, PR0SER2, CXCL8, ENSG00000277991, H4C15, ZC3H12A, ENSG00000267784. TAF1C, RPSAP4, ENSG00000288980, B3GAT2, ANKRD11P1, ENSG00000289410, ENSG00000287322, HSP90AA5P, CXCR4, ZFP3-DT, RPL21P37, RPL7P60, UBE2V1P2, ENSG00000283025, ENSG00000259199, ENSG00000287754, CCR7, LINC00463, DGAT2, INAFM2, GAPDHP42, TNFAIP3, ENSG00000286301, or any combination thereof. See Table 2.
[0097] In some instances, the esophageal cancer status is predicted by the one or more machine learning models based on the list of differentially expressed mitochondrial RNAs comprising at least two, three, four, five, six, seven, eight, nine, ten or more of the following mitochondrial RNAs: ENSG00000274383, MT-ND1, MT-C02, MT-ATP6, MT-CYB, MT- ATP8, MT-ND4, RFX5-AS1, MT-ND3, FRAT1, MT-ND2, MT-ND4L, RAB6B, MT-ND6. MT-CO1, MT-ND5, MT-TQ, MT-RNR2, MT-CO3, MTATP6P1, GNG8, HHEX, MTND1P23, MT-RNR1, MT-TS1, FRAT2, PSMD10, CLC, ENSG00000277991, LINC01487, EIF3EP1, APOM, LINC01747, MTCO3P29, PHB2P1, ENSG00000240174, ENSG00000235713. ENSG00000237378, NANOGP6, STIP1P3, MTND6P24, LINC02151. ENSG00000229853, TTC30A, ENSG00000283286, ST6GALNAC2, RGS 1, RPL7L1P3, ENSG00000248911, FAM71B, ENSG0000025I010, MRPS18BP2, THAP12P6, ZNF695, QRICH2, DU0XA1, or any combination thereof. See Table 2.
[0098] In some instances, the one or more machine learning models prediction distinguish disease (e.g.. precancerous status and esophageal cancer status) from the negative disease status (i.e., healthy) based on the list of differentially expressed mitochondrial RNAs comprising at least two, three, four, five, six, seven, eight, nine, ten or more of the following mitochondrial RNAs: xxx, or any combination thereof. See Table 2.
[0099] Table 2:
[0100]
[0101] At step 340, the one or more trained machine learning models that predict the disease status of the subject based differentially expressed cell-free RNAs are output for deployment into a production environment. In some embodiments, one trained machine learning model (e.g.. a trained logistic regression model) is selected, based on its performance during training, to be implemented into a production environment to predict whether a subject has a negative disease status, a precancerous Barrett’s Esophagus status, or an esophageal cancer status. One of skill in the art can appreciate how the type of machine learning model selected for production environment implementation depends on the disease status(es) being investigated and the number / complexity of the cell-free RNA data being input as features to distinguish disease status. Accordingly, machine learning models other than logistic regression can be selected for production environment deployment based on its performance during training.
[0102] Once deployed, the one or more trained machine learning models are integrated as a component of a larger system or service, wherein one or more of the methods described in steps 305-335 are also performed as part of the larger system or service. By way of example, the trained model(s) may be integrated with clinical laboratory information systems or electronic health records if deployed into a clinical laboratory or hospital IT infrastructure. In another example, the trained model(s) may be deployed on secure cloud infrastructure (e.g., AWS, Azure, Google Cloud) and accessed through web-based portals or APIs by healthcare providers or laboratories. As another nonlimiting example, the trained model(s) may be deployed within academic medical centers or research hospitals for translational research and technology development.
[0103] The one or more trained machine learning models, once deployed, are continuously monitored to ensure it performs as expected. This involves tracking the model’s prediction accuracy, response times, and other operational metrics. Additionally, the trained models may require retraining or updates based on new data or changing conditions in the environment it is applied in. This can be useful because machine learning models can drift over time due to changes in the underlying data the models are making predictions on, a phenomenon known as model drift.
[0104] In another aspect, the one or more trained machine learning models are adapted from a first task (e.g., predicting esophageal cancer statuses based on differentially expressed cell-free RNAs) to a second task (e.g., predicting the disease statuses for other non- esophageal cancer conditions or diseases based on differentially expressed cell-free RNAs) using transfer learning. Examples of other non-esophageal cancer conditions or diseases can include non-esophageal cancers and precancerous conditions, neurodegenerative diseases, infectious diseases, genetic diseases, autoimmune diseases, metabolic diseases, cardiovascular diseases, neurological / psychiatric diseases, respiratory' diseases, endocrine diseases, digestive diseases, musculoskeletal diseases, skin diseases, congenital diseases, blood diseases, or any combination thereof.
[0105] Transfer learning is a powerful technique in machine learning that leverages a pretrained model developed for one task and adapts it to perform a related but different task. In the present aspect, the one or more machine learning models have been trained to predict esophageal cancer disease status in a subject based on their differentially expressed cell-free RNA profile in a blood or plasma sample. These one or more trained models have already learned to identify various features and patterns (e.g., RNA sequence motifs, expression patterns) related to distinguishing between a negative disease status from a precancerous Barretf s Esophagus status, a negative disease status from an esophageal cancer status, and a negative disease status from a disease status (i.e., precancerous and cancerous). To adapt these models for predicting other non-esophageal cancer conditions or diseases based on differentially expressed cell-free RNA profiles, transfer learning is employed. The initial layers of the pre-trained models, which capture general features (e.g.. convolutional layers for sequence motifs, embedding layers for expression data) are retained. These features are often generic enough to be useful across different ty pes of conditions and diseases. By keeping these layers, the one or more models benefit from the knowledge already gained during the initial training phase.
[0106] To specialize the one or more models for the second task, the later layers, which are more task-specific, can be fine-tuned and / or retrained on a smaller, specific dataset of differentially expressed cell-free RNAs (the "target domain") relevant to the new task. This process involves adjusting the weights of these layers to better capture the specific characteristics and patterns relevant to the new task. In addition to, or in alternative of. the later layers may be replaced with new output layers tailored to the new task (e.g.. binary classification for disease vs. healthy, multi-class for disease subtypes). Transfer learning is highly efficient because it requires significantly less training data and computational resources compared to training a model from scratch. Moreover, transfer learning can lead to faster convergence and improved performance, as the one or more models start with a solid base of learned features, allowing it to adapt more quickly to the new task.
[0107] Also provided herein are methods of treating a subject having precancerous Barretf s Esophagus or esophageal cancer with an effective amount of one or more therapeutic agents based on the predicted precancerous status or the cancer status. The treatment method includes performing the method described herein and selecting and administering treatment based on the results of method. In various embodiments, the subject having the precancerous status or the cancer status is treated with an effective amount of one or more therapeutic agents based on the precancerous status or the cancer status. In various embodiments, the subject having the precancerous status or the cancer status is treated with an effective amount of one or more therapeutic agents based on one or more detected differentially expressed cell- free RNAs. Optionally, the method is repeated after treatment to track progression or improvement based on therapeutic intervention.
[0108] In some embodiments, the method of treating the subject further comprises repeating the method of determining the disease status in the subject after one or more treatments with the one or more therapeutic agents to determine the presence or recurrence of minimal residual disease or molecular residual disease. In some embodiments, the method of treating the subj ect further comprises repeating the method of determining the disease status in the subject after one or more surgeries to determine the presence or recurrence of minimal residual disease or molecular residual disease. In some embodiments, the method of treating the subject further comprises repeating the method of determining the disease status in the subject after one or more treatments with one or more therapeutic agents to determine one or more mechanisms of drug resistance. Minimal residual disease (MRD) refers to the small number of cancer cells that can remain in a patient’s body after treatment, such as chemotherapy or radiation, and are undetectable using traditional diagnostic methods like microscopy. These residual cells are clinically significant because they have the potential to survive initial therapy and eventually lead to relapse or recurrence of the disease. The detection and monitoring of MRD are important because they provide a more sensitive and precise assessment of a patient’s response to therapy. Molecular residual disease, by contrast, is a more specific term that refers to the detection of these remaining cancer cells using molecular techniques that analyze genetic or epigenetic alterations unique to the cancer, such as mutations, rearrangements, or other DNA or RNA changes. This approach is especially relevant in both hematologic and solid tumors, where molecular assays such as quantitative PCR, digital droplet PCR, or NGS of circulating tumor DNA (ctDNA) in the blood (‘‘liquid biopsy”) are employed to provide a highly sensitive and specific measure of residual disease.
[0109] Treatment refers to improving or slowing progression of one or more symptoms of either precancerous Barrett’s Esophagus or esophageal cancer in the subject being treated. Treatment can include providing to the subject an effective amount of a therapeutic agent. Therapeutic agents used to treat precancerous Barrett's Esophagus or esophageal cancer include, without limitation, traditional chemotherapies, targeted therapies (such as HER2 or PI3K inhibitors), and immune checkpoint inhibitors (such as ipilimumab, tremelimumab, pembrolizumab, nivolumab, tislelizumab, atezolizumab, and durvalumab). For early or precancerous stages, local ablative or endoscopic therapies and acid suppression are standard. Treatment selection is highly individualized and may involve multimodal approaches. The term effective amount, as used throughout, is defined as any amount necessary to produce a desired physiologic response, for example, reducing or delaying one or more effects or symptoms of a disease or disorder. Effective amounts and schedules for administering the therapeutic agent can be determined empirically, making such determinations within the skill of one in the art. The dosage ranges for administration are those large enough to produce the desired effect in which one or more symptoms of the disease or disorder are affected (e.g., reduced or delayed). The dosage should not be so large as to cause substantial adverse side effects, such as unwanted cross-reactions, unwanted cell death, and the like. Generally, the dosage will vary with the species, age. body weight, general health, sex and diet of the subject, the mode and time of administration and severity of the particular condition and can be determined by one of skill in the art. The dosage can be adjusted by the individual phy sician in the event of any contraindications. Dosages can vary and can be administered in one or more doses.
[0110] The therapeutic agent described herein are administered in a number of ways depending on whether local or systemic treatment is desired. The compositions are administered via any of several routes of administration, including intraparenchymal injection, intravenously, intrathecally, intramuscularly, intracistemally, transdermally, or a combination thereof. Effective doses for any of the administration methods described herein can be extrapolated from dose-response curves derived from in vitro or animal model test systems.
[0111] Also provided herein is a method for determining one or more drug targets to treat a disease in a subject. The method includes (a) collecting a blood sample from the subject comprising cell-free ribonucleic acids (RNAs); (b) detecting, using a sequencing method, the cell-free RNAs in the blood sample, wherein the sequencing method generates sequencing read data for the cell-free RNAs; (c) quantifying, using the sequencing read data, transcript abundance for the cell-free RNAs in the blood sample; (d) identifying one or more differentially expressed cell-free RNA by comparing the transcript abundance of the cell-free RNAs in the blood sample to cell-free RNA transcript abundance values from healthy subjects; and (e) providing, based on the one or more differentially expressed cell-free RNAs, one or more drug targets to treat the disease in a subject. Optionally, the one or more drug targets is targeted by small molecule drugs, antibodies, antibody drug conjugates, RNA medicines, RNA vaccines, gene therapies, or cell therapies to treat the disease in a subject.
[0112] By analyzing the cfRNA profiles of patients with a particular disease and comparing them to those of healthy individuals, researchers can identify genes and regulatory RNAs that are differentially expressed in the disease state. These disease-associated cfRNAs can reveal which cellular pathways are abnormally activated or repressed, highlighting potential molecular drivers of the condition. Additionally, cfRNA analysis can uncover mutations, gene fusions, or alternative splicing events unique to the disease, which may produce novel proteins or regulatory elements that can be specifically targeted by drugs. Through these insights, cfRNAs help pinpoint the genes, proteins, or non-coding RNAs that play critical roles in disease progression, guiding the selection and development of targeted therapies. This non-invasive, dynamic approach not only facilitates the identification of new drug targets but also supports personalized medicine by enabling ongoing monitoring of how patients’ molecular profiles respond to treatment.
[0113] EXAMPLES
[0114] The following examples are provided by way of illustration only and not by way of limitation. Those of skill in the art will readily recognize a variety of non-critical parameters that could be changed or modified to yield essentially the same or similar results.
[0115] RNA liquid biopsies detect cancer noninvasively by profiling cell-free RNAs that are secreted from cancer cells into the circulation. However, the utility of RNA liquid biopsies for precancer detection remains unclear. As shown below, full-length nanopore sequencing of blood plasma from healthy individuals, precancerous Barrett’s esophagus patients with highgrade dysplasia, or patients with early -stage esophageal adenocarcinoma revealed a diverse and dynamic cell-free RNA transcriptome that can be leveraged for disease detection. Using the method described herein 270,679 novel, unannotated cell-free RNAs, were discovered and used as features for training machine learning models to accurately classify precancerous dysplasia or early-stage cancer. Moreover, mitochondrial cell-free RNAs were also found to be highly enriched in precancerous and cancer patient plasma. These findings highlight the utility of full-length cell-free RNA sequencing and novel transcript discovery to enable both precancer and cancer early detection.
[0116] EXAMPLE 1: Methods
[0117] Cell-free RNA isolation: Blood plasma was collected from de-identified healthy controls, precancerous Barret’s esophagus patients with high-grade dysplasia, and patients with early-stage esophageal adenocarcinoma. Healthy donors were identified as those without disease by BioIVT sample metadata; however, healthy donors may have other unreported comorbidities. Samples were initially filtered through a 0.8 pm filter to remove any contaminants, such as cellular debris. Filtered plasma was then processed using the Norgen Biotek Plasma / Serum Circulating and Exosomal RNA purification slurry format kit to isolate cell-free RNAs. The cell-free RNAs were eluted into 50 pL and stored at -80°C prior to cDNA preparation. cDNA preparation: The Takara SMART-Seq HT kit plus kit was used to synthesize cDNA from the isolated cell-free RNA samples from healthy controls, precancerous Barret’s esophagus patients with high-grade dysplasia, and patients with early-stage esophageal adenocarcinoma. cDNA was prepared as specified by the manufacturer’s protocol and amplified for 24 cycles. Quantification of the resulting cDNA was performed using a Nanodrop spectrophotometer and a Qubit Fluorometer. The size distribution of the cDNA was verified using an Agilent Bioanalyzer prior to sequencing library preparation.
[0118] Long-read sequencing: Barcoded nanopore sequencing libraries were created using a custom low-input protocol for the Oxford Nanopore Technologies SQK-NBD-114.24 and LSK-114 kits. Briefly, cDNA was end-repaired and A-tailed as specified, but incubated for 30 minutes, followed by a 30-minute deactivation. Barcodes from the SQK-NBD-114.24 kit were incubated at 20°C for 4.5 hours. Ligation was terminated by the addition of 2 pL of EDTA to each reaction. Barcoded samples were pooled and subject to a 1.8x Ampure Bead cleanup and then used as input for the LSK-114 protocol. All bead cleanups were modified to use a 1.8x ratio and all bead washes were performed at 37°C. Adapter ligation was performed as specified but incubated for 20 minutes to increase ligation efficiency. Short fragment buffer was used to capture the full gamut of cDNA lengths. The final library was eluted into 25 pL of elution buffer and loaded onto an R10.4 Promethion flow cell. POD5s were base called using and demultiplexed Dorado version 0.5.3 using the SUPv4.3.0 base calling model. Reads that failed demultiplexing were re-demultiplexed using a custom python script that exactly matches barcodes, and later combined with their respective FASTQs.
[0119] Short-read sequencing: Illumina libraries were prepared using full-length cDNA generated from the isolated plasma RNA of healthy, Barrett’s esophagus, and esophageal cancer samples using the SMART-Seq HT kit from Takara Bio USA. Final libraries were prepared with the Nextera XT Prep Kit from Illumina and sequenced on an Illumina Novaseq using a 2x151 sequencing chemistry.
[0120] Novel transcript discovery: Isoquant version 3.3.1 with the Human reference genome build HG38 and the GENCODE version 39 comprehensive transcript annotation set was used to identify novel transcripts against all FASTQs in one batch with the following parameters: -d nanopore, -stranded none, -report_novel_unspliced true, -matching_strategy loose, — gene quantification all, — transcript quantification all
[0121] Novel transcripts were merged with the GENCODE annotations and with repeat elements from Repeatmasker for HG38 to generate a reference transcriptome that was used as input for COMPLETE-seq analysis as previously described. Reference transcriptome assembly quality was assessed using SQANTI version 3 with default parameters.
[0122] RNA-seq quality and quantification: RNA-seq short reads (FASTQ) were trimmed using FastP (vO.23.4), assessed for quality using FastQC (vO. 12.1), and quantified using Salmon (vl.10. 1). Quantification was run with the below parameters to reduce sequences biases, enable selective alignments, rescue reads with an unmapped pair, and improve quantification accuracy.
[0123] — gcBias, — seqBias, — validateMappings. -recoverOrphans, — rangeFactorizationBins 4
[0124] The Salmon mappings were performed against a concatenation of:
[0125] 1) The GENCODE consortium Hg38 reference annotation (v.39), the RepeatMasker track from the University of California Santa Cruz genome browser, and
[0126] 2) The annotation of novel genes identified via IsoQuant (3.3.1) This assembled a transcriptome of ~5.5 million transcripts, annotated with the corresponding ENSEMBL transcript IDs. UCSC genome browser repeat instance names, and novel gene IDs from IsoQuant. Transcript counts were aggregated to the gene level for GENCODE genes, and to the repeat ‘family’ level for repeat elements as previously described.
[0127] Differential Expression analysis with the short read: Salmon quantifications were loaded into R (v4.3.3) using tximport (vl.22.0) and converted to DESeq2 (vl.42.0) objects. Biotypes were assigned to genes using their respective GENCODE biotypes, or RepeatMasker 'subclass’ annotation for gene and repeat transcripts respectively. IsoQuant annotated novel genes were assigned the ‘Novel’ biotype. Count normalization and differential expression analyses were performed using DESeq2 (vl.42.0). Significant differential expression was limited to genes with a Benjamini-Hochberg adjusted p-value < 0.05, and an absolute value log2 fold-change > 1.
[0128] Modeling’. Logistic regression models were trained in R using the glmnet package (v4.1), with the identified differentially expressed genes as input features. Models were trained with respect to the sample age and sex (~condition+age+sex). One sample was missing the age information in the metadata, and the average age of samples with the same histology and sex was used. Optimal lambda values for the logistic regression model were identified using 5-fold cross validation, and the value 1 standard error away from the minimum cross-validated value of lambda w as selected for final model training. The final model was trained utilizing the entire dataset, and further feature reduction was performed using LASSO regression.
[0129] EXAMPLE 2: LOCATE-seq RNA liquid biopsy platform
[0130] An RNA liquid biopsy platform called LOCATE-seq was developed that leverages the power of long-read nanopore sequencing and machine learning to discover novel cell-free RNAs and to precisely detect whether a patient has precancer or cancer in their body (FIG. 4A). Cell-free RNA was isolated from the plasma of 16 healthy individuals, 12 patients diagnosed with precancerous high-grade dysplasia, and 19 patients diagnosed with early - stage esophageal adenocarcinoma. Full-length cDNA synthesis was performed, and full- length cDNA libraries were generated for nanopore sequencing on the PromethlON 48 sequencer. For novel transcript discovery, all 47 nanopore sequencing libraries (40,709,643 total reads), which revealed 276.346 previously unannotated transcripts that were not found in the GENCODE annotation of the human transcriptome. Of these novel cell-free RNA transcripts, 270,679 occurred in intergenic regions. Examples of novel cell-free RNA transcripts are provided in FIGs 4B-4E.
[0131] A custom transcriptome reference was generated by combining the 276,346 novel cell-free RNA transcripts with all of the known transcripts in GENCODE, substantially expanding the annotated human cell-free RNA transcriptome. To more deeply profile the healthy and patient plasma samples, high-depth short-read sequencing was performed on the full-length cell-free RNA libraries using the NovaSeq X sequencing (1.26 billion total reads). The reads were mapped to the custom cell-free RNA transcriptome reference for quantification of well-annotated mRNAs (FIG. 4F) and IncRNAs (FIG. 4G), as well as the novel cell-free RNAs (FIG. 4H) discovered using nanopore sequencing. Novel cell-free RNAs were distributed throughout the genome (FIG. 41), and the majority of novel cell-free RNAs were less than 1,000 nucleotides in length (FIG. 4J).
[0132] EXAMPLE 3: Mitochondrial RNA enrichment in dysplasia and cancer
[0133] Quantification of cell-free RNA using only the GENCODE annotation and subsequent differential expression (DE) analysis revealed significant enrichment of protein-coding and mitochondrial RNA in precancerous high-grade dysplasia when compared to healthy (FIG. 5 A), early-stage esophageal cancer when compared to healthy (FIG. 5B), and when comparing healthy to both dysplasia and cancer cell-free RNA (FIG. 5C). The list of differentially expressed mitochondrial RNA transcripts for FIGs. 5A-5C are provided in Table 2 above. Protein-coding RNAs were the most commonly enriched across each of the DE analyses (FIG. 5D), while IncRNAs w ere the most commonly depleted across each of the comparisons (FIG. 5E). Significantly more mitochondrial genes w ere represented in the cell- free RNA of patients with precancerous dysplasia or early-stage cancer (FIG. 5F), suggesting that mitochondrial RNAs may serve as highly abundant biomarkers for both precancer and cancer early detection. Notably, a subset of patients exhibited substantially higher levels of mitochondrial RNA in their cell-free RNA transcriptomes, while other patients show ed more moderate levels of mitochondrial RNA enrichment (FIG. 5G).
[0134] EXAMPLE 4: Differential abundance of novel cell-free RNA in dysplasia and cancer
[0135] Next, the custom transcriptome reference containing all of the novel cell-free RNA transcripts was used for quantification and DE analysis in precancerous high-grade dysplasia when compared to healthy (FIG. 6A), early-stage esophageal cancer when compared to healthy (FIG. 6B), and when comparing healthy to both dysplasia and cancer cell-free RNA (FIG. 6C). Protein-coding RNAs were again the most commonly enriched in disease across each of the comparisons (FIG. 6D), while novel cell-free RNAs were significantly depleted in dysplasia or early-stage cancer cell-free RNA (FIG. 6E). There was also a cluster of 20 novel cell-free RNAs that exhibited higher overall expression across both dysplasia and cancer patients (FIG. 6F). EXAMPLE 5: Novel cell-free RNA features enable accurate disease classification
[0136] When the novel cell-free RNAs that were unique to each healthy individual or patient were examined, there was a significant increase in unique novel cell-free RNAs in both dysplasia and cancer (FIG. 7A), with a few patients expressing over 1,000 unique novel cell- free RNAs (FIG. 7B). Using an expanded feature set that contained both well-annotated GENCODE transcripts and the novel cell-free RNA transcripts, machine learning models (logistic regression) were trained with LI regularization using the significantly DE cell-free RNAs as input features (p-adjusted < 0.05). The models segregated both high-grade dysplasia (FIG. 7C) and early-stage esophageal cancer (FIG. 7D) from healthy samples with perfect sensitivity7(100%) and specificity7(100%) using a limited number of features, highlighting the potential of our approach for both precancer and cancer early detection using a noninvasive RNA liquid biopsy.
[0137] EXAMPLE 6: Dysregulated pathways as potential therapeutic targets in dysplasia and cancer
[0138] Common gene sets and pathways that were significantly dysregulated in both highgrade dysplasia and early-stage esophageal cancer patients were characterized to identify potential therapeutic targets. Gene seat enrichment analysis (GSEA) was performed and it was found that the mitochondrial process of “oxidative phosphorylation” was among the top two enriched gene sets across dysplasia and early-stage cancer (FIG. 8A). “PI3K AKT MTOR signaling” was also in the top 10 enriched gene sets in both dysplasia and cancer (FIG. 8A), suggesting that small molecule inhibitors of oxidative phosphorylation or the PI3K pathway may serve as effective treatments for patients with high-grade dysplasia and / or early-stage cancer. We also identified several more potential therapeutic targets, including immune checkpoint genes, that were upregulated in dysplasia and / or cancer (FIG. 8B-G), suggesting that HER2 inhibition or immune checkpoint inhibitors of CTLA-4 and / or PD- 1 / PD-L1 may also serve as effective treatments for dysplasia and / or cancer patients.
[0139] DISCUSSION
[0140] The method provided herein shows that comprehensive profiling of cell-free RNA using full-length nanopore sequencing enables the discovery of over 270.000 unannotated. novel cell-free RNA transcripts in healthy individuals and patients with precancerous highgrade dysplasia or early-stage esophageal adenocarcinoma. The study substantially increases the breadth of the human cell-free RNA transcriptome and provides proof-of-concept that building and leveraging a custom transcriptome reference that incorporates known and novel cell-free RNAs provides a diverse set of novel features with w hich to train highly accurate machine learning models for disease classification and the early detection of both cancer and precancerous conditions.
[0141] Additional Considerations
[0142] Although specific embodiments have been described, various modifications, alterations, alternative constructions, and equivalents are also encompassed within the scope of the disclosure. Embodiments are not restricted to operation within certain specific data processing environments, but are free to operate within a plurality of data processing environments. Additionally, although embodiments have been described using a particular series of transactions and steps, it should be apparent to those skilled in the art that the scope of the present disclosure is not limited to the described series of transactions and steps. Various features and aspects of the above-described embodiments may be used individually or jointly.
[0143] Further, while embodiments have been described using a particular combination of hardw are and softw are, it should be recognized that other combinations of hardware and software are also within the scope of the present disclosure. Embodiments may be implemented only in hardware, or only in software, or using combinations thereof. The various processes described herein can be implemented on the same processor or different processors in any combination. Accordingly, where components or services are described as being configured to perform certain operations, such configuration can be accomplished, e g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Processes can communicate using a variety of techniques including but not limited to conventional techniques for inter process communication, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0144] The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. It will, however, be evident that additions, subtractions, deletions, and other modifications and changes may be made thereunto without departing from the broader spirit and scope as set forth in the claims. Thus, although specific disclosure embodiments have been described, these are not intended to be limiting. Various modifications and equivalents are within the scope of the following claims.
[0145] The use of the terms '‘a” and “an” and “the” and similar referents in the context of describing the disclosed embodiments (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms “comprising,” “having,” “including,” and “containing” are to be construed as open-ended terms (i.e., meaning “including, but not limited to,”) unless otherwise noted. The term “connected” is to be construed as partly or wholly contained within, attached to, or joined together, even if there is something intervening. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate embodiments and does not pose a limitation on the scope of the disclosure unless otherw ise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0146] Disjunctive language such as the phrase “at least one of X, Y. or Z,” unless specifically stated otherwise, is intended to be understood within the context as used in general to present that an item, term, etc., may be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y. or at least one of Z to each be present.
[0147] The term “about” is used to provide flexibility to a numerical range endpoint by providing that a given value may be “slightly above” or “slightly below” the endpoint without affecting the desired result.
[0148] As used herein, the transitional phrase "consisting essentially of (and grammatical variants) is to be interpreted as encompassing the recited materials or steps "and those that do not materially affect the basic and novel characteristic(s)" of the claimed invention. See, In re Herz, 537 F.2d 549, 551-52, 190 U.S.P.Q. 461, 463 (CCPA 1976) (emphasis in the original); see also MPEP §2111.03. Thus, the term "consisting essentially of as used herein should not be interpreted as equivalent to "comprising."
[0149] Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise-indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. For example, if a concentration range is stated as 1% to 50%, it is intended that values such as 2% to 40%, 10% to 30%, or 1% to 3%, etc., are expressly enumerated in this specification. These are only examples of what is specifically intended, and all possible combinations of numerical values between and including the lowest value and the highest value enumerated are to be considered to be expressly stated in this disclosure.
[0150] Preferred embodiments of this disclosure are described herein, including the best mode known for carrying out the disclosure. Variations of those preferred embodiments may become apparent to those of ordinary skill in the art upon reading the foregoing description. Those of ordinary skill should be able to employ such variations as appropriate and the disclosure may be practiced otherwise than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein.
[0151] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0152] In the foregoing specification, aspects of the disclosure are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the disclosure is not limited thereto. Various features and aspects of the above-described disclosure may be used individually or jointly. Further, embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive.
Claims
1. WHAT IS CLAIMED IS:
1. A method for determining a disease status of a subject comprising:(a) collecting a biological sample from the subject comprising cell-free ribonucleic acids (RNAs);(b) detecting, using a sequencing method, the cell-free RNAs in the biological sample, wherein the sequencing method generates sequencing read data for the cell-free RNAs;(c) identifying a subset of cell-free RNAs from the cell-free RNAs in the biological sample by annotating the sequencing read data for the cell-free RNAs to sequencing read data for a standard reference transcriptome;(d) combining, based on the annotated sequencing read data for the identified subset of cell-free RNAs, the subset of cell-free RNAs with the standard reference transcriptome to generate a custom reference transcriptome comprising a custom set of cell-free RNAs;(e) quantifying transcript abundance for the custom set of cell-free RNAs in the custom reference transcriptome;(f) identifying a list of differentially expressed cell-free RNAs by comparing the transcript abundance of the custom set of cell-free RNAs in the custom reference transcriptome to cell-free RNA transcript abundance values from healthy subjects;(g) training, using at least the list of differentially expressed cell-free RNAs, one or more machine learning models to predict the disease status of the subject; and(h) outputting the one or more trained machine learning models that predict the disease status of the subject into a production environment.
2. The method of claim 1, wherein the disease status is a negative disease status, a pre-disease status, or a positive disease status.
3. The method of claim 1. wherein the positive disease status indicates the subject having one or more diseases selected from a list comprising a cancerous disease, a precancerous condition, a neurodegenerative disease, an infectious disease, a genetic disease, an autoimmune disease, a metabolic disease, a cardiovascular disease, a neurological / psychiatric disease, a respiratory disease, an endocrine disease, a digestive disease, a musculoskeletal disease, a skin disease, a congenital disease, a blood disease, or any combination thereof.
4. The method of claim 3, wherein the positive disease status indicates the cancerous disease.
5. The method of claim 4. wherein the cancerous disease is esophageal cancer.
6. The method of claim 2, wherein the pre-disease status indicates the precancerous condition.
7. The method of claim 6, wherein the precancerous condition is Barret’s Esophagus.
8. The method of claim 1, wherein the biological sample is blood.
9. The method of claim 1. wherein the biological sample is plasma.
10. The method of claim 1. wherein the sequencing method identifies the full length of the cell-free RNAs in the biological sample.
11. The method of claim 10, wherein the sequencing method comprises a nanopore sequencing method, a sequencing by expansion method, a long-read sequencing method, a single-molecule sequencing method, a sequencing by synthesis method, or a short read sequencing method.
12. The method of claim 10, wherein the sequencing method is the long-read sequencing method.
13. The method of claim 1. wherein the detecting is done using RT-PCR, a probebased target capture method with or without subsequent sequencing, a primer-based target enrichment method with or without subsequent sequencing, or a CRISPR-based detection method.
14. The method of claim 1. wherein the quantification is done using a sequencing method.
15. The method of claim 14, wherein the detecting sequencing method and the quantification sequencing method are the same.
16. The method of claim 15, wherein the detecting sequencing method and the quantification sequencing method are long-read sequencing methods.
17. The method of claim 1, wherein the detecting sequencing method and the quantification sequencing method are different.
18. The method of claim 17, wherein the detecting sequencing method is a long- read sequencing method and the quantification sequencing method is a short-read sequencing method.
19. The method of claim 1, wherein the one or more machine learning models comprise a regression model, a classification model, a clustering model, a statistical model, a decision tree, an ensemble model, a Bayesian model, a deep learning model, or any combination thereof.
20. The method of claim 19, wherein at least one of the one or more machine learning models is a logistic regression model.
21. The method of claim 1, wherein the training comprises: performing iterative operations to adjust a set of model parameters to minimize a loss or error function of the one or more machine learning models, wherein the loss or error function is configured to measure the difference between the output predictions of the one or more machine learning models and the ground truth dataset.
22. The method of claim 1 , wherein the one or more machine learning models predicts, based at least on the list of differentially expressed cell-free RNAs, that the disease status of the subject is a negative disease status, a pre-disease status, or a positive disease status.
23. The method of claim 22, wherein the one or more machine learning model predicts, based at least on the list of differentially expressed cell-free RNAs, the pre-disease status, wherein the pre-disease status indicates precancerous Barrett’s Esophagus.
24. The method of claim 23, wherein the list of expressed cell-free RNAs indicating precancerous Barrett’s Esophagus comprises two or more of the cell-free RNAs comprising *chr22_KI27087, MT-ND1, MT-ATP6, MT-RNR1, MT-CO2, MT-CO1, MT- CO3, MTATP6P1, MT-ND4L, MT-RNR2, MT-ND4, MT-ND3, PVALB, MT-ND5, MT- ND2. *chrl2_8540, MT-CYB, MT-ATP8, *chr5_23822, *chrl_28278, ENSG00000274383, MTND2P28. MT-ND6, *chrl5 11639, MTNDIP23, ENSG00000266401, *chr6 9458, AIM2, CXCL8, *chr3_0052, *chr7_24810, ZCH312A, *chrl 1 2126, *chr21_7097, *chrl7_23338, *chrl3_18256, *chrl_48852, *chr5_14282, *chrl2_19620, *chr6_26573, RPSAP4, *chrl 1_29228, *chr3_41901, *chr4_4793, DLX1, ENSG00000288980,*chrl 1_29228. *chr4_42463, *chr3_564, *chrl0_28310, *chrl7_14620, *chrl7_10248, *chr3_15268, AIRN, *chrl4_6471, ANKRD11P1, or any combination thereof.
25. The method of claim 24, wherein the list of differentially expressed cell-free RNAs indicating precancerous Barrett’s Esophagus comprises two or more of the cell -free RNAs comprising L3MBTL1, DLXL *chr6_29343, *chrl2_12317, *chr5_31003, *chr20_8731, NADK, RPSAP4, ALOX5, *chr20_12853, *chr2_41140, *chrl3_19893, *chr21_3060, or any combination thereof.
26. The method of claim 1, wherein the one or more machine learning models predict, based at least on the list of differentially expressed cell-free RNAs. the positive disease status, wherein the positive disease status indicates esophageal cancer.
27. The method of claim 26, wherein the list of differentially expressed cell-free RNAs indicating esophageal cancer comprises two or more of the cell-free RNAs comprising *chrl2_8540, *chr5_14206, ENSG00000274383, *chr5_23822, *chrl_28278, MT-C03, MT-C02, MT-ATP6, MT-CYB. *chrl0_36912, MT-ATP8, MT-ND4. MT-ND3, MT-ND1. *chrl5_11661, *chrl6_5531, MT-ND4L, MT-ND2, MT-ND6, MT-TQ, MT-C01, MT-ND5, *chrl5_l 1639, GNG8, MT-RNR2, MTATP6P1, *chr6_18454, HHEX, MTND2P41, *chr4_36834, *chr4_37169, *chr6_15475, *chr8_23641, *chr6_16530, *chr21_3639, *chrl 32054, *chr3 42038, *chrl8 6774, *chrl9 15033, LINC01487, *chrl2 15574, *chrl5_21741, EIF4EP1, *chrl_38898, *chrl0_35990, *chrl2_7067, *chrl2_25273, *chrl2_26920, *chrl5_6605, *chrl0_32620, *chr9_1146, RPL7AP10, *chrl2_10146, *chrl2_1619, *chr2_38235, *chr5_131, or any combination thereof.
28. The method of claim 27, wherein the list of differentially expressed cell-free RNAs indicating esophageal cancer comprises two or more of the cell-free RNAs comprising *chrl_34706, *chr7_7352, *chr9_29702, *chrl6_5531, *chr20_2254, *chr6_11543, ENTR1P1, *chr20_12853. *chr3_14400, PPP2CA, MRPL20-DT, *chr8_33134, ENSG00000286896. CASC9, *chr9_6537, DYNC1LI1, *chr8_24909. *chrl_12702, ENSG00000268047, *chr7_452, *chrl_49351, BHLHE40, *chrl6_943, *chr2_38235, *chr4_40238, *chr4_6748, *chr2_34817, *chr4_29780, or any combination thereof.
29. The method of claim 1. wherein the list of differentially expressed cell-free RNAs comprise mitochondrial RNAs.
30. The method of claim 29, wherein the mitochondrial RNAs comprise two or more of the cell-free RNAs listed in Table 2.
31. A method for determining a disease status in a subj ect comprising:(a) collecting a biological sample from the subject comprising cell-free ribonucleic acids (RNAs);(b) detecting, using a sequencing method, the cell-free RNAs in the biological sample, wherein the sequencing method generates sequencing read data for the cell-free RNAs;(c) identifying a subset of cell-free RNAs from the cell-free RNAs in the biological sample by annotating the sequencing read data for the cell-free RNAs to sequencing read data for a standard reference transcriptome;(d) combining, based on the annotated sequencing read data for the identified subset of cell-free RNAs, the subset of cell-free RNAs with the standard reference transcriptome to generate a custom reference transcriptome comprising a custom set of cell-free RNAs;(e) quantifying transcript abundance for the custom set of cell-free RNAs in the custom reference transcriptome;(f) identifying a list of differentially expressed cell-free RNAs by comparing the transcript abundance of the custom set of cell-free RNAs in the custom reference transcriptome to cell-free RNA transcript abundance values from healthy subjects;(g) predicting, using one or more trained machine learning models, the disease status of the subject based at least on the list of differentially expressed cell-free RNAs, wherein the disease status is a negative disease status, a pre-disease status, or a positive disease status; and(h) outputting the predicted disease status of the subject in a production environment.
32. The method of claim 31, wherein the positive disease status indicates a cancerous disease.
33. The method of claim 32, wherein the cancerous disease is esophageal cancer.
34. The method of claim 31, wherein the pre-disease status indicates the precancerous condition.
35. The method of claim 34, wherein the precancerous condition is Barrett's Esophagus.
36. The method of claim 31 further comprising treating the subject having the precancerous status or the cancer status with an effective amount of one or more therapeutic agents based on the precancerous status or the cancer status.
37. The method of claim 31 further comprising treating the subject having the precancerous condition status or the positive cancer status with an effective amount of one or more therapeutic agents based on one or more detected cell-free RNAs.
38. The method of claim 36 or 37, further comprising repeating the method of claim 31 after one or more treatments with the one or more therapeutic agents to determine the efficacy of the one or more therapeutic agents.39 . The method of claim 36 or 37, further comprising repeating the method of claim 31 after one or more treatments with the one or more therapeutic agents to determine the presence or recurrence of minimal residual disease or molecular residual disease.
40. The method of claim 36 or 37, further comprising repeating the method of claim 31 after one or more surgeries to determine the presence or recurrence of minimal residual disease or molecular residual disease.
41. The method of claim 36 or 37, further comprising repeating the method of claim 31 after one or more treatments with one or more therapeutic agents to determine one or more mechanisms of drug resistance.
42. A method for determining one or more drug targets to treat a disease in a subject comprising:(a) collecting a blood sample from the subject comprising cell-free ribonucleic acids (RNAs);(b) detecting, using a sequencing method, the cell-free RNAs in the blood sample, wherein the sequencing method generates sequencing read data for the cell-free RNAs;(c) quantifying, using the sequencing read data, transcript abundance for the cell-free RNAs in the blood sample;(d) identifying one or more differentially expressed cell-free RNA by comparing the transcript abundance of the cell-free RNAs in the blood sample to cell-free RNA transcript abundance values from healthy subjects; and(e) providing, based on the one or more differentially expressed cell-free RNAs, one or more drug targets to treat the disease in a subject.
43. The method of claim 42, wherein the one or more drug targets is targeted by small molecule drugs, antibodies, antibody drug conjugates, RNA medicines, RNA vaccines, gene therapies, or cell therapies to treat the disease in a subject.
Citation Information
Patent Citations
Detecting esophageal disorders
US20180037958A1
Methods for detecting disease using analysis of RNA
US20230072924A1
Methods for polynucleotide sequencing
WO2023196983A2
Treatment and detection of cancers having a neural-like progenitor, squamoid / basaloid / mesenchymal, or classical phenotype
WO2023230632A2
Repeat-aware profiling of cell-free RNA
WO2024010875A1