Machine learning techniques for using DNA methylation in clinical diagnosis and prognosis of acute myeloid leukemia
Biologically-informed machine learning models using DNA methylation data effectively address the challenges in AML diagnosis and prognosis, providing accurate predictions for improved clinical management.
Patent Information
- Application Number
- PCT/US2024/058595
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-12-05
- Publication Date
- 2025-06-12
AI Technical Summary
Acute myeloid leukemia (AML) diagnosis and prognosis are challenging due to the heterogeneity and complexity of leukemic stem cells, and existing technologies struggle to effectively translate genomic and epigenomic data into clinically relevant information.
The development of biologically-informed machine learning models that utilize DNA methylation data from peripheral or bone marrow blood samples to predict clinical diagnoses and prognoses for AML patients, providing accurate classification and treatment insights.
These machine learning models efficiently and accurately provide clinical diagnosis and prognosis predictions for AML patients, enabling informed treatment decisions and improving patient outcomes.
Smart Images

Figure IMGF000016_0001 
Figure 00000031_0000 
Figure 00000032_0000
Abstract
Description
MACHINE LEARNING TECHNIQUES FOR USING DNA METHYLATION INCLINICAL DIAGNOSIS AND PROGNOSIS OF ACUTE MYELOID LEUKEMIASUPPORT STATEMENT
[0001] This invention was made with government support under R01 CA132946 awarded by the National Institutes of Health. The government has certain rights in the invention.CROSS-REFERENCE TO RELATED APPLICATION(S)
[0002] This application claims priority to U.S. Application No. 63 / 606,254, filed December 5, 2023, the content of which is incorporated herein by reference in its entirety.TECHNOLOGICAL FIELD
[0003] The present disclosure generally relates to the technical field of clinical diagnosis and prognosis of acute myeloid leukemia. In particular, the present disclosure includes embodiments directed to technical improvements in predicting clinical diagnoses and prognoses related to acute myeloid leukemia using DNA methylation. The present disclosure additionally relates to the technical field of machine learning models and techniques.BACKGROUND
[0004] Acute myeloid leukemia (AML) is a devastating disease associated with high morbidity and mortality. Decades of fundamental research have posed leukemic stem cells are highly heterogenous, elusive, and ever evolving. Translating that many layers of complexity into clinical practice has been a monumental challenge. Through applied effort, ingenuity, and innovation, the problems identified herein have been solved by developing solutions that are included in embodiments of the present disclosure, many examples of which are described in detail herein.BRIEF SUMMARY
[0005] Various embodiments of the present disclosure provide predictive clinical tools that includes respective biologically-informed machine learning-trained models that are configured to determine diagnostic classification and / or prognostic classification with respect to acute myeloid leukemia (AML) for a subject (e g., patient, study participant, and / or the like). In variousembodiments, at least a portion of the input data provided to the machine learning-trained model is DNA methylation data. In certain embodiments, the DNA methylation data is determined via analysis and / or sequencing of a peripheral and / or bone marrow blood sample.
[0006] According to an aspect of the present disclosure, a method for determining a clinical diagnosis prediction and / or prognosis prediction with respect to acute myeloid leukemia (AML) is provided. In an example embodiment, the method includes obtaining, by a prediction model operating on a system computing entity, subject data comprising DNA methylation data corresponding to a subject; providing, by the prediction model, the subject data to an input layer of a machine learning-trained model configured to generate at least one of a clinical diagnosis prediction or a prognosis prediction for the subject based on a learned transformation of the subject data; extracting, by the prediction model, at least one prediction from an output layer of the machine learning-trained model; and causing, by the prediction model, the at least one prediction to be provided for human review.
[0007] According to another aspect, an apparatus configured to determine and provide a clinical diagnosis prediction and / or prognosis prediction with respect to acute myeloid leukemia (AML) is provided. In an example embodiment, the apparatus includes at least one processor and a memory storing computer program code. The memory and the computer program code are configured, when executed by the at least one processor, to cause the apparatus to perform obtaining subject data comprising DNA methylation data corresponding to a subject; providing the subject data to an input layer of a machine learning-trained model configured to generate at least one of an acute myeloid leukemia (AML) clinical diagnosis prediction or a prognosis prediction for the subject based on a learned transformation of the subject data; extracting at least one prediction from an output layer of the machine learning-trained model; and causing the at least one prediction to be provided for human review.
[0008] According to another aspect, a computer program product configured to cause an apparatus to determine and provide a clinical diagnosis prediction and / or prognosis prediction with respect to acute myeloid leukemia (AML) is provided. In an example embodiment, the computer program product includes at least one non-transitory computer-readable memory media storing computer-executable instructions. The computer executable instructions are configured to, when executed by a processor of an apparatus, cause the apparatus to perform obtaining subject data comprising DNA methylation data corresponding to a subject; providing the subject data to aninput layer of a machine learning-trained model configured to generate at least one of an acute myeloid leukemia (AML) clinical diagnosis prediction or a prognosis prediction for the subject based on a learned transformation of the subject data; extracting at least one prediction from an output layer of the machine learning-trained model; and causing the at least one prediction to be provided for human review.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Having thus described the present disclosure in general terms, reference will now be made to the accompanying drawings, which are not necessarily drawn to scale.
[0010] Figure 1 is a diagram of a system architecture that can be used in conjunction with various embodiments of the present disclosure;
[0011] Figure 2 is a schematic of a computing entity that may be used in conjunction with various embodiments of the present disclosure;
[0012] Figure 3 provides a flowchart illustrating an example process for training a prediction model and using the prediction model to provide a diagnosis and / or prognosis prediction, in accordance with various embodiments of the present disclosure;
[0013] Figure 4 illustrates an example machine learning model that may be configured, trained, and used as a prediction model, in accordance with various embodiments described herein;
[0014] Figures 5A, 5B, and 5C illustrate example views of an example interactive user interface configured for providing at least one prediction for human review; and
[0015] Figure 6 provides a flowchart illustrating an example process for pre-processing subject data for use in determining at least one prediction.DETAILED DESCRIPTION
[0016] Various embodiments of the present disclosure now will be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all embodiments of the disclosure are shown. Indeed, the disclosure may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. The term “or” (also designated as “ / ”) is used herein in both the alternative and conjunctive sense, unless otherwiseindicated. The terms “illustrative” and “exemplary” are used to be examples with no indication of quality level. Like numbers refer to like elements throughout.I. General Overview and Exemplary Technical Advantages
[0017] Various embodiments of the present disclosure provide a machine learning-based framework for predicting clinical diagnoses and / or prognoses of subjects with respect to acute myeloid leukemia (AML) for a subject (e.g., patient, study participant, and / or the like). For example, a prediction model is configured to receive input data. The input data include DNA methylation data for a subject. The prediction model includes at least one machine learning-trained model that is trained to generate and / or provide a clinical diagnosis prediction and / or a prognosis prediction for a subject based on the input data provided thereto. The prediction model may then provide the clinical diagnosis prediction and / or prognosis prediction for review by a medical professional.
[0018] Acute myeloid leukemia (AML) is a devastating disease associated with high morbidity and mortality. Decade of fundamental research have posed leukemic stem cells are highly heterogenous, elusive, and ever evolving. Translating that many layers of complexity into clinical practice has been a monumental challenge. Genomic and epigenomic analyses are plagued by the overwhelming amount of data they generate, which greatly obfuscates clinically relevant information. For example, determining the type of AML a subject is experiencing can be quite difficult. However, determining an appropriate course of treatment for the subject may depend on an accurate determination of the type of AML the subject is experiencing. Thus, various technical limitations and challenges exist regarding effective prediction of clinical diagnosis and prognosis relating to AML.
[0019] Various embodiments provide technical solutions to these technical limitations and challenges. In various embodiments, a prediction model includes at least one biologically- informed machine learning-trained model trained to determine a clinical diagnosis prediction and / or a prognosis prediction for a subject based on input data provided thereto. The input data includes DNA methylation data corresponding to the subject. In certain embodiments, the DNA methylation data is determined via sequencing of a peripheral blood sample, which is a significantly less intrusive tissue sample to collect compared to a bone marrow sample. The prediction model efficiently and accurately provides the clinical diagnosis prediction and / orprognosis prediction for the subject. Thus, the prediction model enables determination of treatment results for different types of AML and / or important information for use by a clinical team in determining a treatment plan for a subject. For example, certain embodiments support the integration of epigenomic profiling into the hematology / oncology unit setting, offering a pathway toward more personalized and effective trial designs and treatment modalities for patients with AML. Various embodiments therefore provide technical improvements and / or advantages to the relevant technical fields.II. Exemplary Technical Implementation of Various Embodiments
[0020] Embodiments of the present disclosure may be implemented in various ways, including as computer program products that comprise articles of manufacture. Such computer program products may include one or more software components including, for example, software objects, methods, data structures, and / or the like. A software component may be coded in any of a variety of programming languages. An illustrative programming language may be a lower-level programming language such as an assembly language associated with a particular hardware architecture and / or operating system platform. A software component comprising assembly language instructions may require conversion into executable machine code by an assembler prior to execution by the hardware architecture and / or platform. Another example programming language may be a higher-level programming language that may be portable across multiple architectures. A software component comprising higher-level programming language instructions may require conversion to an intermediate representation by an interpreter or a compiler prior to execution.
[0021] Other examples of programming languages include, but are not limited to, a macro language, a shell or command language, a job control language, a script language, a database query, or search language, and / or a report writing language. In one or more example embodiments, a software component comprising instructions in one of the foregoing examples of programming languages may be executed directly by an operating system or other software component without having to be first transformed into another form. A software component may be stored as a file or other data storage construct. Software components of a similar type or functionally related may be stored together such as, for example, in a particular directory, folder, or library. Softwarecomponents may be static (e.g., pre-established or fixed) or dynamic (e.g., created or modified at the time of execution).
[0022] A computer program product may include a non-transitory computer-readable storage medium storing applications, programs, program modules, scripts, source code, program code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like (also referred to herein as executable instructions, instructions for execution, computer program products, program code, and / or similar terms used herein interchangeably). Such non-transitory computer-readable storage media include all computer-readable media (including volatile and non-volatile media).
[0023] In one embodiment, a non-volatile computer-readable storage medium may include a floppy disk, flexible disk, hard disk, solid-state storage (SSS) (e.g., a solid-state drive (SSD), solid state card (SSC), solid state module (SSM)), enterprise flash drive, magnetic tape, or any other non-transitory magnetic medium, and / or the like. A non-volatile computer-readable storage medium may also include a punch card, paper tape, optical mark sheet (or any other physical medium with patterns of holes or other optically recognizable indicia), compact disc read only memory (CD-ROM), compact disc-rewritable (CD-RW), digital versatile disc (DVD), Blu-ray disc (BD), any other non-transitory optical medium, and / or the like. Such a non-volatile computer- readable storage medium may also include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory (e.g., Serial, NAND, NOR, and / or the like), multimedia memory cards (MMC), secure digital (SD) memory cards, SmartMedia cards, CompactFlash (CF) cards, Memory Sticks, and / or the like. Further, a non-volatile computer- readable storage medium may also include conductive-bridging random access memory (CBRAM), phase-change random access memory (PRAM), ferroelectric random-access memory (FeRAM), non-volatile random-access memory (NVRAM), magnetoresistive random-access memory (MRAM), resistive random-access memory (RRAM), Silicon-Oxide-Nitride-Oxide- Silicon memory (SONOS), floating junction gate random access memory (FJG RAM), Millipede memory, racetrack memory, and / or the like.
[0024] In one embodiment, a volatile computer-readable storage medium may include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), fast page mode dynamic random access memory (FPM DRAM), extended data-outdynamic random access memory (EDO DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), double data rate type two synchronous dynamic random access memory (DDR2 SDRAM), double data rate type three synchronous dynamic random access memory (DDR3 SDRAM), Rambus dynamic random access memory (RDRAM), Twin Transistor RAM (TTRAM), Thyristor RAM (T-RAM), Zero-capacitor (Z-RAM), Rambus in-line memory module (RIMM), dual in-line memory module (DIMM), single in-line memory module (SIMM), video random access memory (VRAM), cache memory (including various levels), flash memory, register memory, and / or the like. It will be appreciated that where embodiments are described to use a computer-readable storage medium, other types of computer-readable storage media may be substituted for or used in addition to the computer-readable storage media described above.
[0025] As should be appreciated, various embodiments of the present disclosure may also be implemented as methods, apparatus, systems, computing devices, computing entities, and / or the like. As such, embodiments of the present disclosure may take the form of a data structure, apparatus, system, computing device, computing entity, and / or the like executing instructions stored on a computer-readable storage medium to perform certain steps or operations. Thus, embodiments of the present disclosure may also take the form of an entirely hardware embodiment, an entirely computer program product embodiment, and / or an embodiment that comprises a combination of computer program products and hardware performing certain steps or operations.
[0026] Embodiments of the present disclosure are described with reference to example operations, steps, processes, blocks, and / or the like. Thus, it should be understood that each operation, step, process, block, and / or the like may be implemented in the form of a computer program product, an entirely hardware embodiment, a combination of hardware and computer program products, and / or apparatus, systems, computing devices, computing entities, and / or the like carrying out instructions, operations, steps, and similar words used interchangeably (e.g., the executable instructions, instructions for execution, program code, and / or the like) on a computer- readable storage medium for execution. For example, retrieval, loading, and execution of code may be performed sequentially such that one instruction is retrieved, loaded, and executed at a time. In some exemplary embodiments, retrieval, loading, and / or execution may be performed in parallel such that multiple instructions are retrieved, loaded, and / or executed together. Thus, such embodiments can produce specifically configured machines performing the steps or operationsspecified in the block diagrams and flowchart illustrations. Accordingly, the block diagrams and flowchart illustrations support various combinations of embodiments for performing the specified instructions, operations, or steps.
[0027] Figure 1 provides an illustration of a system architecture 100 that may be used in accordance with various embodiments of the disclosure. Here, the architecture 100 includes various components involved in training one or more machine learning-trained models of a prediction model and / or using a prediction model to provide clinical diagnosis predictions and / or prognosis predictions with respect to AML for one or more subjects. As illustrated, the architecture 100 includes one or more user computing entities 120, one or more system computing entities 110, and one or more networks 105.
[0028] In various embodiments, a user computing entity 120 is configured for user interaction. For example, a clinical diagnosis prediction and / or prognosis prediction may be provided (e.g., displayed) via a user computing entity 120. In another example, a user may operate a user computing entity 120 to provide (e.g., via user input) at least a portion of input data corresponding to a subject (e.g., patient, study participant, and / or the like) and / or to provide user input initiating the generation of a clinical diagnosis prediction and / or prognosis prediction for a subject. In some embodiments, a user computing entity 120 is configured to operate and / or execute a prediction model including one or more machine learning-trained models for determining and AML-related diagnostic and / or prognostic prediction for a subject.
[0029] In various embodiments, the system computing entity 110 is configured for training one or more machine learning-trained models of a prediction model and / or operating / executing the prediction model to generate and / or provide an AML-related clinical diagnosis prediction and / or a prognosis prediction for a subject. For example, in various embodiments, the system computing entity 110 may be configured for performing various operations related to training, implementation, and use of a prediction model comprising one or more machine learning-trained models. In various embodiments, the system computing entity 110 is used for storing the machine learning-trained models of the prediction model, and / or the like. In various embodiments, the system computing entity 110 may be a cloud computing system, a distributed computing system, an edge computing system, and / or the like that may perform computational operations as a service.
[0030] As noted, the user computing entity 120 and the system computing entity 110 may communicate with one another over one or more networks 105. Depending on the embodiment,these networks 105 may comprise any type of known network such as a land area network (LAN), wireless land area network (WLAN), wide area network (WAN), metropolitan area network (MAN), wireless communication network, cellular communication networks, the Internet, and / or combinations thereof. In addition, these networks 105 may comprise any combination of standard communication technologies and protocols. For example, communications may be carried over the networks 105 by link technologies such as Ethernet, 802.11, CDMA, 3G, 4G, 5G or digital subscriber line (DSL). Further, the networks 105 may support a plurality of networking protocols, including the hypertext transfer protocol (HTTP), the transmission control protocol / internet protocol (TCP / IP), or the file transfer protocol (FTP), and the data transferred over the networks 105 may be encrypted using technologies such as, for example, transport layer security (TLS), secure sockets layer (SSL), and internet protocol security (IPsec). Those skilled in the field of the present disclosure will recognize Figure 1 represents but one possible configuration of a system architecture 100, and that variations are possible with respect to the protocols, facilities, components, technologies, and equipment used.
[0031] Figure 2 provides a schematic of an exemplary apparatus 200 that may be used in accordance with various embodiments of the present disclosure. In particular, the apparatus 200 may be configured to perform various example operations described herein to relating to the training of one or more machine learning-trained models of a prediction model, operation / execution of a prediction model to generate AML-related clinical diagnosis predictions and / or prognosis predictions, and / or provision of clinical diagnosis predictions and / or prognosis predictions for review by clinical personnel. In some example embodiments, the apparatus 200 may be embodied by the user computing entity 120. In some other example embodiments, the apparatus 200 may be embodied by the system computing entity 110.
[0032] In general, the terms computing entity, entity, device, and / or similar words used herein interchangeably may refer to, for example, one or more computers, computing entities, desktop computers, mobile phones, tablets, phablets, notebooks, laptops, distributed systems, items / devices, terminals, servers or server networks, blades, gateways, switches, processing devices, processing entities, set-top boxes, relays, routers, network access points, base stations, the like, and / or any combination of devices or entities adapted to perform the functions, operations, and / or processes described herein. Such functions, operations, and / or processes may include, for example, transmitting, receiving, operating on, processing, displaying, storing, determining,creating / generating, monitoring, evaluating, comparing, and / or similar terms used herein interchangeably. In one embodiment, these functions, operations, and / or processes can be performed on data, content, information, and / or similar terms used herein interchangeably. In an example embodiment, a user computing entity 120 is a tablet, laptop, or desktop computer and the system computing entity 110 is a server or part of a Cloud-computing resource.
[0033] Although illustrated as a single computing entity, those of ordinary skill in the field should appreciate that the apparatus 200 shown in Figure 2 may be embodied as a plurality of computing entities, tools, and / or the like operating collectively to perform one or more processes, methods, and / or steps. As just one non-limiting example, the apparatus 200 may comprise a plurality of individual data tools, each of which may perform specified tasks and / or processes.
[0034] Depending on the embodiment, the apparatus 200 may include one or more network and / or communications interfaces 220 for communicating with various computing entities, such as by communicating data, content, information, and / or similar terms used herein interchangeably that can be transmitted, received, operated on, processed, displayed, stored, and / or the like. Thus, in certain embodiments, the apparatus 200 may be configured to receive data from one or more data sources and / or devices as well as receive data indicative of input, for example, from a device.
[0035] The networks used for communicating may include, but are not limited to, any one or a combination of different types of suitable communications networks such as, for example, cable networks, public networks (e.g., the Internet), private networks (e.g., frame-relay networks), wireless networks, cellular networks, telephone networks (e.g., a public switched telephone network), or any other suitable private and / or public networks. Further, the networks may have any suitable communication range associated therewith and may include, for example, global networks (e.g., the Internet), MANs, WANs, LANs, or PANs. In addition, the networks may include any type of medium over which network traffic may be carried including, but not limited to, coaxial cable, twisted-pair wire, optical fiber, a hybrid fiber coaxial (HFC) medium, microwave terrestrial transceivers, radio frequency communication mediums, satellite communication mediums, or any combination thereof, as well as a variety of network devices and computing platforms provided by network providers or other entities.
[0036] Accordingly, such communication may be executed using a wired data transmission protocol, such as fiber distributed data interface (FDDI), digital subscriber line (DSL), Ethernet, asynchronous transfer mode (ATM), frame relay, data over cable service interface specification(DOCSIS), or any other wired transmission protocol. Similarly, the apparatus 200 may be configured to communicate via wireless external communication networks using any of a variety of protocols, such as general packet radio service (GPRS), Universal Mobile Telecommunications System (UMTS), Code Division Multiple Access 2000 (CDMA2000), CDMA2000 IX (IxRTT), Wideband Code Division Multiple Access (WCDMA), Global System for Mobile Communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), Time Division- Synchronous Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), 5G New Radio (5G NR), Evolved Universal Terrestrial Radio Access Network (E-UTRAN), Evolution-Data Optimized (EVDO), High Speed Packet Access (HSPA), High-Speed Downlink Packet Access (HSDPA), IEEE 802.11 (Wi-Fi), Wi-Fi Direct, 802.16 (WiMAX), ultra-wideband (UWB), infrared (IR) protocols, near field communication (NFC) protocols, Wibree, Bluetooth protocols, wireless universal serial bus (USB) protocols, and / or any other wireless protocol. The apparatus 200 may use such protocols and standards to communicate using Border Gateway Protocol (BGP), Dynamic Host Configuration Protocol (DHCP), Domain Name System (DNS), File Transfer Protocol (FTP), Hypertext Transfer Protocol (HTTP), HTTP over TLS / SSL / S ecure, Internet Message Access Protocol (IMAP), Network Time Protocol (NTP), Simple Mail Transfer Protocol (SMTP), Telnet, Transport Layer Security (TLS), Secure Sockets Layer (SSL), Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Datagram Congestion Control Protocol (DCCP), Stream Control Transmission Protocol (SCTP), HyperText Markup Language (HTML), and / or the like.
[0037] In addition, in various embodiments, the apparatus 200 includes or is in communication with one or more processing elements 205 (also referred to as processors, processing circuitry, and / or similar terms used herein interchangeably) that communicate with other elements within the apparatus 200 via a bus, for example, or network connection. As will be understood, the processing element 205 may be embodied in several different ways. For example, the processing element 205 may be embodied as one or more complex programmable logic devices (CPLDs), microprocessors, multi-core processors, coprocessing entities, application-specific instruction-set processors (ASIPs), and / or controllers. Further, the processing element 205 may be embodied as one or more other processing devices or circuitry. The term circuitry may refer to an entirely hardware embodiment or a combination of hardware and computer program products. Thus, the processing element 205 may be embodied as integrated circuits, application specific integratedcircuits (ASICs), field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), hardware accelerators, other circuitry, and / or the like.
[0038] As will therefore be understood, the processing element 205 may be configured for a particular use or configured to execute instructions stored in volatile or non-volatile media or otherwise accessible to the processing element 205. As such, whether configured by hardware, computer program products, or a combination thereof, the processing element 205 may be capable of performing steps or operations according to embodiments of the present disclosure when configured accordingly.
[0039] In various embodiments, the apparatus 200 may include or be in communication with non-volatile media (also referred to as non-volatile storage, memory, memory storage, memory circuitry and / or similar terms used herein interchangeably). For instance, the non-volatile storage or memory may include one or more non-volatile storage or non-volatile memory media 210 such as hard disks, ROM, PROM, EPROM, EEPROM, flash memory, MMCs, SD memory cards, Memory Sticks, CBRAM, PRAM, FeRAM, RRAM, SONOS, racetrack memory, and / or the like. As will be recognized, the non-volatile storage or non-volatile memory media 210 may store files, databases, database instances, database management system entities, images, data, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like. The term database, database instance, database management system entity, and / or similar terms used herein interchangeably and in a general sense to refer to a structured or unstructured collection of information / data that is stored in a computer-readable storage medium.
[0040] In particular embodiments, the non-volatile memory media 210 may also be embodied as a data storage device or devices, as a separate database server or servers, or as a combination of data storage devices and separate database servers. Further, in some embodiments, the non-volatile memory media 210 may be embodied as a distributed repository such that some of the stored information / data is stored centrally in a location within the system and other information / data is stored in one or more remote locations. Alternatively, in some embodiments, the distributed repository may be distributed over a plurality of remote storage locations only. As already discussed, various embodiments contemplated herein use data storage in which some or all the information / data required for various embodiments of the disclosure may be stored.
[0041] In various embodiments, the apparatus 200 may further include or be in communication with volatile media (also referred to as volatile storage, memory, memory storage, memory circuitry and / or similar terms used herein interchangeably). For instance, the volatile storage or memory may also include one or more volatile storage or volatile memory media 215 as described above, such as RAM, DRAM, SRAM, FPM DRAM, EDO DRAM, SDRAM, DDR SDRAM, DDR2 SDRAM, DDR3 SDRAM, RDRAM, RIMM, DIMM, SIMM, VRAM, cache memory, register memory, and / or the like.
[0042] As will be recognized, the volatile storage or volatile memory media 215 may be used to store at least portions of the databases, database instances, database management system entities, data, images, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like being executed by, for example, the processing element 205. Thus, the databases, database instances, database management system entities, data, images, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like may be used to control certain aspects of the operation of the apparatus 200 with the assistance of the processing element 205 and operating system.
[0043] In various embodiments, the apparatus 200 includes one or more user interfaces 225. For example, the user interfaces 225 may include one or more output devices such as one or more displays, one or more speakers, and / or the like. For example, the user interfaces 225 may include one or more input devices such as a hard or soft keyboard or keypad, mouse, touchscreen, touchpad, microphone, and / or the like.
[0044] As will be appreciated, one or more of the computing entity’s components may be located remotely from other computing entity components, such as in a distributed system. Furthermore, one or more of the components may be aggregated, and additional components performing functions described herein may be included in the apparatus 200. Thus, the apparatus 200 can be adapted to accommodate a variety of needs and circumstances.III. Exemplary Framework and Operations
[0045] Various embodiments described herein address technical limitations relating to the difficulties in diagnosing AML and various subtypes of AML and accurate prognosis prediction. Moreover, various embodiments described herein address technical limitations regardingdetermining effective strategies for harnessing the overwhelming amount of data generated by genomic and epigenomic analyses.
[0046] Figure 3 provides a flowchart illustrating various processes, procedures, operations, and / or the like performed by a system computing entity 110, for example, to train one or more biologically-informed machine learning-trained models of a prediction model and to use the prediction model to generate and provide a clinical diagnosis prediction and / or prognosis prediction for a subject (e.g., a patient, a study participant, and / or the like).
[0047] In an example embodiment, the prediction model comprises a diagnosis model that includes a machine-learning trained model configured to receive subject data as input and provide a clinical diagnostic prediction as output. In an example embodiment, the clinical diagnostic prediction may identify whether or not the subject has AML or not and / or a probability that the subject has AML. In another example embodiment, the clinical diagnostic prediction identifies a sub-type of AML the subject is predicted to have. In another example embodiment, the clinical diagnostic prediction indicates the likelihood that the subject has a respective AML sub-type for one or more AML sub-types.
[0048] In certain embodiments, the prediction model includes a prognosis model that includes a machine-learning trained model configured to receive subject data as input and provide a prognosis prediction as output. In various embodiments, the prognosis prediction may indicate the likelihood of the subject being alive (due to AML-related issues) after one or more time periods (e g., six months, a year, two years, five years, ten years, and / or the like). In certain embodiments, the prognosis prediction may indicate the likelihood of the subject being alive (due to AML-related issues) after one or more time periods (e.g., six months, a year, two years, five years, ten years, and / or the like) depending on whether the subject receives one or more treatments. For example, in an example scenario the prognosis prediction may indicate that the subject is expected to be alive in five years if the subject receives a first treatment and to be alive in two years but not alive in five years if the subject does not receive the second treatment. In various embodiments, a prognosis prediction may include one or more of predicted or probability of one or more outcome endpoints such as survival, overall survival, event free survival or minimal residual disease after induction therapy corresponding to the subject.
[0049] In certain embodiments, the prediction model is configured to execute the diagnosis model and, based on the output thereof (e.g., a clinical diagnosis prediction) determine whether totrigger execution of the prognosis model. For example, the prediction model may trigger execution of prognosis model responsive to the diagnosis model indicating that there is a likelihood greater than a threshold likelihood that the subject has AML. In some embodiments, the clinical diagnosis prediction (e.g., likelihood that the subject has one or more AML sub-types) is provided as input to the prognosis model in addition to the subject data.
[0050] In various embodiments, the prediction model is configured to generate and / or provide a clinical diagnosis prediction. In an example such embodiment, the prediction uses an unsupervised machine learning model, such as Pairwise Controlled Manifold Approximation (PaCMAP), for example, to reduce dimensions from 331,556 variables selected according to biology, into a handful of coordinates (e.g., 2 or 5 coordinates or dimensions). In various embodiments, the dimension reduction is performed using a PaCMAP model that is modified by adding hyperparameter tuning with human feedback based on clinical annotations. This novel technique generates a novel set of parameters: n_neighbors=10, MN_ ratio =0.4, FP ratio=20.0, random state=42, lr=0.1, num iter s=4500. Notably, the prediction model describes, for the first time, a complete picture of the landscape of acute leukemias according to diagnostic guidelines. Moreover, visualization thereof is enabled based on the reduced dimensionality of the variables. Additionally, this example embodiment is probably one of the first studies to ever use PaCMAP for the purpose of creating diagnostic tools for blood cancers.
[0051] Furthermore, the resulting coordinates of the model (PaCMAP dimensions 1 to 5) are linked with a supervised multi-class classifier grid search. In an example embodiment, supervised training of the multi-class classifier grid search resulted in the optimized parameters — Best parameters diagnostic model: {'class weight': 'balanced', 'reg alpha': 0.1, 'reg lambda': 4}, Best parameters prognostic model: {'class_weigh : 'balanced', 'reg_alpha': 8, 'reg_lambda': 4}. Both models were trained using GridSearchCV(lgbm, param_grid, cv=5, njobs=-l, scoring='roc_auc_ovr_weighted'). See full information at C
[0052] In various embodiments, the prediction model is configured to generate and / or provide a prognosis prediction. In an example such embodiment, a custom format of the conventional Epigenome-Wide-Association-Study (EWAS) is generated by replacing logistic regression with Cox Proportional Hazards Regression and adding risk group categories as covariates. This model is then used to map the risk adjusted associations of CpGs with clinical outcome- overall survival.In an example embodiment, a threshold is used to select the top results from the EWAS pipeline above with a Cox-PH-Lasso model that generates weighted methylation levels in 38 CpGs, which for the first time is being described as a robust predictor of clinical outcomes in pediatric AML patients. Importantly, this is the first application of Cox-PH-EWAS in pediatric AML. For example, the prediction model defines a 38-CpG signature that robustly predicts overall survival of AML patients, in an example embodiment. In certain embodiments, the prognosis prediction for a subject is determined and / or calculated by multiplying the average coefficient from 1000 Lasso Models for the 38 CpGs followed by its association with multiple clinical endpoints.
[0053] Starting with step 302 of Figure 3, the system computing entity 110 obtains training data. For example, the system computing entity 110 may receive training data via network 105, read training data from memory 210, 215, and / or the like. For example, the training data may be accessed from one or more databases of data including DNA methylation data for one or more subjects. In various embodiments, the training data includes data from one or more clinical studies. In some embodiments, the DNA methylation data of an instance of training data is associated with one or more demographic identifiers (e.g., age, race, sex, geographic location, and / or the like). In certain embodiments, an instance of training data may include clinical outcome information for the subject. For example, the instance of training data my include information regarding one or more treatments received by the subject and the clinical outcomes of each of the one or more treatments.
[0054] At step 304, the system computing entity 110 filters and / or pre-processes the training data. For example, the system computing entity 110 may remove instances of data from the training data that are incomplete or do not include one or more necessary components. For example, the system computing entity 110 may pre-process the training data to harmonize the training data, place instances of training data in a pre-defined format, and / or the like. In an example embodiment, the system computing entity 110 pre-processes the training data to reduce the dimensionality thereof. For example, in an example embodiment, the system computing entity 110 uses an unsupervised learning algorithm such as Pairwise Controlled Manifold Approximation (PaCMAP), t-distributed stochastic neighbor embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), or Large-scale Dimensionality Reduction Using Triplets (TriMAP), for example, to reduce the dimensionality of the training data.
[0055] At step 306, the system computing entity 110 trains one or more machine learning- trained models of the prediction model using at least a portion of the training data. For example, in an example embodiment, the prediction model includes a diagnosis model configured for generating and / or providing a clinical diagnosis prediction and / or a prognosis model configured for generating and / or providing a prognosis prediction. The diagnosis model and / or the prognosis model are respective machine learning-trained models. In an example embodiment, the prediction model includes a diagnosis and prognosis model which is a machine learning-trained model configured to generate and / or provide a clinical diagnosis prediction and a prognosis prediction. In various embodiments, the diagnosis model, prognosis model, and / or diagnosis and prognosis model are classifiers. In certain embodiments, the diagnosis model, prognosis model, and / or diagnosis and prognosis model are multi-class classifiers. In various embodiments, the diagnosis model, prognosis model, and / or diagnosis and prognosis model are configured to receive subject data including reduced dimensionality DNA methylation data via respective input layers.
[0056] In various embodiments, training data may be associated with demographic indicators. For example, an instance of training data may be associated with one or more indicators indicating the subjects age, sex, race, geographic location, and / or other demographic information. In an example embodiment, the prediction model is trained to determine which demographic information is relevant for the determination of an accurate diagnosis prediction and / or prognosis prediction. For example, the prediction model may include one or more diagnosis models that are each trained and / or have their parameters tuned based on one or more demographic categories (e.g., age, sex, race, geographic location, etc.) and / or combinations of demographic categories, one or more prognosis models that are each trained and / or have their parameters tuned based on one or more demographic categories (e.g., age, sex, race, geographic location, etc.) and / or combinations of demographic categories, and / or one or more diagnosis and prognosis models that are each trained and / or have their parameters tuned based on one or more demographic categories (e.g., age, sex, race, geographic location, etc.) and / or combinations of demographic categories.
[0057] In various embodiments, the diagnosis model and / or a diagnosis and prognosis model comprise a LightGBM, Gaussian Process, Random Forrest, Logistic Regression, and K-Nearest Neighbors that is trained via a supervised learning algorithm with, for example, hyperparameter tuning and / or 10-fold cross validation.
[0058] In various embodiments, the prognosis model and / or a diagnosis and prognosis model comprise Cox Proportional Hazards (Cox-PH) model and / or a LASSO-based multivariate Cox- PH model configured to generate an epigenetic signature for a subject based on data provided thereto and a classifier configured to classify the epigenetic signature into a risk score category. In various embodiments, the epigenetic signature includes a plurality of CpG values. In certain embodiments, the plurality of CpG values are standardized (e.g., value standardized =,valueunstandarized uiean value) / value standard deviation-) for a distribution of the value established based on, for example, the training data) and the standardized CpG values are provided to the classifier for classification into a risk score category.
[0059] In various embodiments, the one or more machine learning-trained models of the prediction model are trained using appropriate training algorithms and / or loss functions for the respective models. Once the prediction model and / or the one or more machine learning-trained models of the prediction model are trained, the prediction model may be used to generate and / or provide predictions for one or more subjects (e.g., patients, study participants, and / or the like).
[0060] At step 312, the system computing entity 110 obtained subject data. In various embodiments, the subject data comprises data corresponding to a subject. In various embodiments, the subject data comprises DNA methylation data corresponding to the subject. The subject data may also include demographic information and / or one or more demographic indicators corresponding to the subject. In various embodiments, the subject data is obtained by receiving the subject data via network 105 (e.g., provided by a user computing entity 120, and / or the like), accessing the subject data from memory, and / or the like.
[0061] At step 314, the system computing entity 110 pre-processes the subject data. For example, the subject data may be formatted in accordance with a pre-defined format. For example, the dimensionality of the subject data may be reduced. In various embodiments, the pre-processing of the subject data prepares the subject data to be input to the input layer of a respective machinelearning trained model (e.g., a diagnosis model, a prognosis model, a diagnosis and prognosis model) of the prediction model. In various embodiments, pre-processing the subject data includes performance of one or more steps shown in Figure 6 and described in detail elsewhere herein.
[0062] At step 316, the system computing entity 110 inputs the subject data to the input layer of a respective machine-learning trained model (e.g., a diagnosis model, a prognosis model, a diagnosis and prognosis model) of the prediction model. For example, the system computing entity110 causes one or more machine learning-trained models (e.g., a diagnosis model, a prognosis model, a diagnosis and prognosis model) of the prediction model to process the subject data (e.g., the pre-processed subject data) and transform the subject data (based on the learned weights of the respective machine learning-trained model) into one or more predictions. In various embodiments, the prediction model is configured to select a particular diagnosis model from one or more diagnosis models, a particular prognosis model from one or more prognosis models, and / or a particular diagnosis and prognosis model from one or more diagnosis and prognosis models based at least in part on demographic information and / or demographic identifiers included in and / or associated with the subject data.
[0063] In certain embodiments, the prediction model is configured to execute the diagnosis model and, based on the output thereof (e.g., a clinical diagnosis prediction) determine whether to trigger execution of the prognosis model. For example, the prediction model may trigger execution of the prognosis model responsive to the diagnosis model indicating that there is a likelihood greater than a threshold likelihood that the subject has AML (e.g., the sum of likelihoods determined for a plurality of AML sub-types is greater than the threshold likelihood). In various embodiments, the prediction model triggers the execution of the prognosis model by provided the (pre-processed) subject data as input to the prognosis model. In some embodiments, the clinical diagnosis prediction (e.g., likelihood that the subject has one or more AML sub-types) is provided as input to the prognosis model in addition to the subject data.
[0064] At step 318, the system computing entity 110 receives the one or more predictions from the prediction model. For example, the transformation of the subject data by the machine learning- trained model of the prediction model causes the output layer of the machine learning-trained model to be populated with one or more predictions.
[0065] For example, the prediction model may receive a clinical diagnosis prediction from an output layer of a diagnostic model. In an example embodiment, the clinical diagnostic prediction may identify whether or not the subject has AML or not and / or a probability that the subject has AML. In another example embodiment, the clinical diagnostic prediction identifies a sub-type of AML the subject is predicted to have. In another example embodiment, the clinical diagnostic prediction indicates the likelihood that the subject has a respective AML sub-type for one or more AML sub-types.
[0066] For example, the prediction model may receive a prognosis prediction of an output layer of a prognosis model. In various embodiments, the prognosis prediction may indicate the likelihood of the subject being alive (due to AML-related issues) after one or more time periods (e.g., six months, a year, two years, five years, ten years, and / or the like). In certain embodiments, the prognosis prediction may indicate the likelihood of the subject being alive (due to AML-related issues) after one or more time periods (e.g., six months, a year, two years, five years, ten years, and / or the like) depending on whether the subject receives one or more treatments. For example, in an example scenario the prognosis prediction may indicate that the subject is expected to be alive in five years if the subject receives a first treatment and to be alive in two years but not alive in five years if the subject does not receive the second treatment.
[0067] The prediction model may process the one or more predictions and provide them as output of the prediction model. The system computing entity 110 (e.g., one or more processing elements thereof) receives the one or more predictions.
[0068] At step 320, the system computing entity 110 provides the one or more predictions for review by clinical personnel. For example, the system computing entity 110 may transmit the one or more predictions to cause the one or more predictions to be provided (e.g., displayed or audibly provided) via a user interface of a user computing entity 120. For example, the system computing entity 110 may cause an electronic health record (EHR) corresponding to the subject to be updated to include the one or more predictions. For example, the system computing entity 110 may cause a clinical study results database to be updated to include the one or more predictions.
[0069] Thus, as described above, various machine learning models are configured to be trained as part of a prediction model configured to provide clinical diagnosis predictions and / or prognosis predictions with respect to AML using DNA methylation data as at least a portion of the input to the prediction model. In various embodiments, the DNA methylation data is generated and / or determined via analysis of a peripheral and / or bone marrow blood sample from the subject using nanopore epigenomic sequencing.
[0070] According to various embodiments described herein, various configurations and architectures of machine learning models may be used to implement the prediction model. As previously identified, these configurations and architectures of machine learning models may include convolutional neural network (CNN) or deep neural network (DNN) machine learning models, generative adversarial network (GAN) machine learning models, encoder-decoder orautoencoder machine learning models, dictionary learning machine learning models, and / or the like, depending on the embodiment. In an example embodiment, the prediction model includes a light gradient-boosting machine (LightGBM) classifier configured to generate and / or provide clinical diagnosis predictions and / or a Lasso-based Cox Proportional Hazards (Cox-PH) model configured to generate and / or provide prognosis predictions.
[0071] Figure 4 provides a diagram illustrating a general configuration of a DNN machine learning model 400, and in various embodiments, the DNN machine learning model 400 may embody a machine learning-trained model of a prediction model. For example, the DNN machine learning model 400 enables output of a clinical diagnosis prediction and / or prognosis prediction for a subject (e.g., patient, study participant) with respect to AML in response to receiving input data that includes DNA methylation data corresponding to the subject. DNN machine learning models 400 may include or be embodied by convolutional neural network (CNN) machine learning models, recurrent neural network (RNN) machine learning models, graph neural network (GNN) machine learning models, and / or the like. Generally, the DNN machine learning model 400 includes an input layer 410, one or more hidden layers 420, and an output layer 430, with each layer comprising a plurality of neuronal units 402. The neuronal units 402 of a given layer are connected and feed-forward to neuronal units 402 of the succeeding layer via trainable weights 404, in various embodiments. The neuronal units 402 of the output layer 430 then provide a weighted and learned representation of the input and its portions and / or features, for example.
[0072] To embody a machine learning-trained model of a prediction model, a DNN machine learning model 400 may be trained (e.g., via the trainable weights 404) to perform classification processes based on input data. In various embodiments, the output layer 430 of the DNN machine learning model 400 may be coupled with activation functions and / or attention mechanisms that select portions and / or features of the input data for transformation.
[0073] In various embodiments, a DNN machine learning model 400 may be trained to recognize patterns within input data via supervised or unsupervised learning. For example, the DNN machine learning model 400 may learn which features or elements and / or combinations thereof of the input data are most predictive for clinical diagnosis and / or prognosis with respect to AML when the input data includes DNA methylation data. In one example, a historical dataset that includes DNA methylation data for a plurality of subjects and diagnosis and / or outcome data for the plurality of subjects may be used to train the parameters of the DNN machine learningmodel 400. A learned classification (e g., clinical diagnosis prediction and / or prognosis prediction) of a subject based on the input data including DNA methylation data is generated by the DNN machine learning model 400 and then compared with corresponding diagnosis and / or outcome data, and a loss measure resulting from the comparison may be backpropagated through the DNN machine learning model 400 to configure the trainable weights 404 and / or other parameters.
[0074] Similarly, dictionary learning may be used to train DNN machine learning models 400 to embody a machine learning-trained model of a prediction model. In dictionary learning, mappings between input data including DNA methylation data and corresponding clinical diagnoses and / or prognoses are defined, and these mappings are learned by the DNN machine learning model 400 such that, when provided with input data for a subject, the DNN machine learning model 400 outputs a mapped clinical diagnosis prediction and / or prognosis prediction.
[0075] In various embodiments, the system computing entity 110 is configured to cause an interactive user interface (IUI) 550 to be displayed via a display 228 of a user interface 225 of a user computing entity 120, as shown in Figures 5A, 5B, and 5C. In various embodiments, the prediction model may be configured to generate diagnosis and / or prognosis predictions using reduced dimensionality DNA methylation data. For example, the dimensionality of the DNA methylation data of the subject data may be reduced from thousands of dimensions to reduced number of dimensions in a range of three to twenty dimensions (e.g., as part of the pre-processing of the subject data at step 314). The prediction model may be configured to provide the diagnostic and / or prognosis prediction using a further reduced dimensionality. For example, the diagnostic and / or prognosis prediction may be provided by projecting the reduced dimensionality DNA methylation data into a two-dimensional plane or a dimension reduction model (e.g., PaCMAP and / or the like) may be used to reduce the dimensionality of the DNA methylation data to enable display thereof in two-dimensions.
[0076] Figure 5 A shows an IUI displaying a first plot 510 illustrating a plurality of clusters of data corresponding to anonymized DNA methylation data of a plurality of individuals with each data point color-coded based on a World Health Organization (WHO) 2022 diagnosis for the corresponding individual. Second plot 512 shows a zoomed in view of a particular cluster of data points and the star indicates the location of a data point corresponding to the subject in the two- dimensional DNA methylation data. Figure 5B shows an IUI 550 displaying a first plot 520 illustrating a plurality of clusters of data corresponding to anonymized DNA methylation data ofa plurality of individuals with each data point color-coded based on diagnosis prediction for the corresponding individual determined by the prediction model. Second plot 522 shows a zoomed in view of a particular cluster of data points and the star indicates the location of a data point corresponding to the subject in the two-dimensional DNA methylation data. Figure 5C shows an IUI 550 displayed via display 228 displaying a first plot 530 illustrating a plurality of clusters of data corresponding to anonymized DNA methylation data of a plurality of individuals with each data point color-coded based on a prognosis prediction (whether the subject will be alive or dead in five years) for the corresponding individual determined by the prediction model. Second plot 532 shows a zoomed in view of a particular cluster of data points and the star indicates the location of a data point corresponding to the subject in the two-dimensional DNA methylation data. Clinical personnel may review these plots and use the information provided thereby to determine treatment strategies for the subject.
[0077] Figure 6 provides a flowchart illustrating various processes, procedures, operations, and / or the like that may be performed by a system computing entity 110, for example, to pre- process subject data (e.g., at step 314 of Figure 3).
[0078] Starting at step 602, the system computing entity 110 defines a subject data format for the subject data. In various embodiments, the prediction model is configured to process subject data having a set or predefined data format. For example, the prediction model may be configured to read a subject data object comprising one or more strings and / or arrays of values where each value of the one or more strings and / or arrays of values is associated with a particular semantic meaning. A first array of values may provide CpG values for a respective set of CpGs, where the order of the values within the first array of values indicates which CpG the particular value corresponds. A second array of values may provide demographic indicators. For example, the second array may be sequences of binary values that indicate whether or not the subject is a member of a group corresponding to the sequence position of a respective binary value or not. In another example, the second array may be a sequence of values where the value indicates of which group(s) the subject is a member. Notably, the position of a value in an array or string may be used to map that value to a semantic meaning based on the set of predefined data format.
[0079] In various embodiments, the system computing entity 110 may determine whether the subject data obtained at step 312 is in the set or predefined subject data format. For example, the obtained subject data may be obtained as a subject data object including metadata and / or a headerthat indicates that the subject data is provided by the subject data object in the set or predefined subject data format. In another example, the obtained subject data may be determined to not be in the set or predefined subject data format and the system computing entity 110 may generate a subject data object that provides the subject data in the set or predefined subject data format by using the subject data format to map values from the obtained subject data into corresponding sequence positions in strings and / or arrays of subject data object using the set or predefined subject data format.
[0080] At step 604, the system computing entity 110 filters the CpG probes of the DNA methylation data of the subject data. In various embodiments, step 604 may be performed in parallel with step 602. For example, the filtering of the CpG probes may be performed during the generation of the subject data object. In various embodiments, the CpG probes of the DNA methylation data may be filtered to remove CpG probe values associated with low levels of confidence (e.g., where the DNA methylation data is inconclusive regarding a particular CpG probe value). In certain embodiments, the CpG probes of the DNA methylation data of the subject data may be filtered to remove CpG probe values for CpG probes that are determined (e.g., via the training process of steps 302-306, for example) to not be of interest and / or that are not significant for the operation of the prediction model. For example, if a subset of the CpG probes of the DNA methylation data are determined (e.g., via the training process of steps 302-306, for example) to not affect the predictions generated by the prediction model, the subset of CpG probes may be filtered from the DNA methylation data included in the subject data object to reduce memory and computation expenses.
[0081] At step 606, the system computing entity 110 may define one or more demographic indicators of the subject data. In various embodiments, step 606 may be performed in parallel with step 602. For example, the defining one or more demographic indicators may be performed during the generation of the subject data object. For example, if the obtained subject data does not include demographic indicators but includes demographic information, the system computing entity 110 may convert the demographic information into demographic indicators. In another example, if the subject data does not include demographic information, the system computing entity 110 may request and receive (e.g., from an electronic health record (EHR) corresponding to the subject and / or the like) demographic information corresponding to the subject and then convert the demographic information into the demographic indicators of the subject data object.
[0082] At step 608, the system computing entity 110 may transform and / or standardize one or more values associated with CpG probes. In various embodiments, step 608 may be performed in parallel with step 602. For example, the transforming and / or standardizing one or more values associated with CpG probes may be performed during the generation of the subject data object. In various embodiments, transforming and / or standardizing the one or more values associated with CpG probes may include converting or mapping a value associated with a CpG probe to a standardized value based at least in part on how the CpG probe value was determined. For example, similar to how measuring a subject’s temperature by sampling the subject’s oral temperature, ear temperature, or forehead temperature may affect the determined temperature of the subject, the analysis and / or sequencing process to determine the value of one or more CpG probes may affect the determined value. To remove the effects of the analysis and / or sequencing process, the values of one or more CpG probes may be transformed, converted, and / or mapped to corresponding standardized values.
[0083] At step 610, the system computing entity 110 reduces the dimensionality of the DNA methylation data of the subject data. For example, a dimension reduction of the DNA methylation data may be performed using an unsupervised learning algorithm such as PaCMAP, t-SNE, UMAP, or TriMAP. In various embodiments, the dimensionality of the DNA methylation data of the subject data may be reduced from thousands of dimensions to two to twenty dimensions. In an example embodiment, the reduced dimension DNA methylation data is stored in the subject data object in accordance with the set or predefined subject data format.
[0084] The system computing entity 110 may store the subject data object (e.g., in memory 210, 215) and / or use at least a portion of the subject data object (based at least in part on the set of predefined subject data format) to execute the prediction model to cause generation of the one or more predictions (e.g., diagnosis prediction and / or prognosis prediction) for the subject.IV. Conclusion
[0085] Many modifications and other embodiments of the present disclosure set forth herein will come to mind to one skilled in the art to which the present disclosures pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the present disclosure is not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be includedwithin the scope of the appended claim concepts. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
CLAIMSThat which is claimed:
1. A method for determining a clinical diagnosis prediction and / or prognosis prediction with respect to acute myeloid leukemia (AML), the method comprising: obtaining, by a prediction model operating on a system computing entity, subject data comprising DNA methylation data corresponding to a subject; providing, by the prediction model, the subject data to an input layer of a machine learning- trained model configured to generate at least one of a clinical diagnosis prediction or a prognosis prediction for the subject based on a learned transformation of the subject data; extracting, by the prediction model, at least one prediction from an output layer of the machine learning-trained model; and causing, by the prediction model, the at least one prediction to be provided for human review.
2. The method of claim 1, further comprising pre-processing, by the prediction model, the subj ect data to prepare the subj ect data to be provided to the input layer of the machine learning- trained model of the prediction model.
3. The method of claim 2, wherein the pre-processing of the subject data comprises at least one of formatting the subject data, filtering the subject data, defining one or more demographics indicators associated with the subject data, standardizing one or more values of the subject data, or reducing dimensionality of the subject data.
4. The method of claim 1, wherein at least a portion of the subject data provided to the input layer includes one or more data elements of the DNA methylation data.
5. The method of claim 1, wherein the at least one prediction comprises at least one of an AML sub-type prediction or an AML-related prognosis, outcome endpoints such as survival, overall survival, event free survival or minimal residual disease after induction therapy corresponding to the subject.
6. The method of claim 1 , wherein the machine learning-trained model comprises a classifier.
7. The method of claim 1 , wherein one or more data elements of the DNA methylation data are determined based analysis of a peripheral or bone marrow blood sample corresponding to the subject.
8. The method of claim 7, wherein the one or more data elements of the DNA methylation data are determined by nanopore epigenomic sequencing of the peripheral blood sample.
9. The method of claim 1, wherein the subject data provided to in the input layer includes at least one demographic indicator configured to indicate a demographic category of the subject and the machine learning-trained model is configured to generate the at least one of the clinical diagnosis prediction or the prognosis prediction based at least in part on the at least one demographic indicator.
10. An apparatus comprising at least one processor and a memory storing computer program code, the memory and the computer program code configured, when executed by the at least one processor, to cause the apparatus to perform: obtaining subject data comprising DNA methylation data corresponding to a subject; providing the subject data to an input layer of a machine learning-trained model configured to generate at least one of an acute myeloid leukemia (AML) clinical diagnosis prediction or a prognosis prediction for the subject based on a learned transformation of the subject data; extracting at least one prediction from an output layer of the machine learning-trained model; and causing the at least one prediction to be provided for human review.
11. The apparatus of claim 10, wherein the memory and the computer program code are further configured, when executed by the at least one processor, to cause the apparatus to perform pre-processing the subject data to prepare the subject data to be provided to the input layer of the machine learning-trained model.
12. The apparatus of claim 11, wherein the pre-processing of the subject data comprises at least one of formatting the subject data, filtering the subject data, defining one or more demographics indicators associated with the subject data, standardizing one or more values of the subject data, or reducing dimensionality of the subject data.
13. The apparatus of claim 10, wherein at least a portion of the subject data provided to the input layer includes one or more data elements of the DNA methylation data.
14. The apparatus of claim 10, wherein the at least one prediction comprises at least one of an AML sub-type prediction or an AML-related prognosis corresponding to the subject.
15. The apparatus of claim 10, wherein the machine learning-trained model comprises a classifier.
16. The apparatus of claim 10, wherein one or more data elements of the DNA methylation data are determined via analysis of a peripheral or bone marrow blood sample corresponding to the subject.
17. The apparatus of claim 16, wherein the one or more data elements of the DNA methylation data are determined by nanopore epigenomic sequencing of the peripheral or bone marrow blood sample.
18. The apparatus of claim 10, wherein the subject data provided to in the input layer includes at least one demographic indicator configured to indicate a demographic category of the subject and the machine learning-trained model is configured to generate the at least one of the clinical diagnosis prediction or the prognosis prediction based at least in part on the at least one demographic indicator.
19. A computer program product comprising at least one non-transitory computer- readable memory media storing computer-executable instructions, the computer executable instructions configured to, when executed by a processor of an apparatus, cause the apparatus to perform:obtaining subject data comprising DNA methylation data corresponding to a subject; providing the subject data to an input layer of a machine learning-trained model configured to generate at least one of an acute myeloid leukemia (AML) clinical diagnosis prediction or a prognosis prediction for the subject based on a learned transformation of the subject data; extracting at least one prediction from an output layer of the machine learning-trained model; and causing the at least one prediction to be provided for human review.
20. The computer program product of claim 19, wherein one or more data elements of the DNA methylation data are determined via analysis of a peripheral or bone marrow blood sample corresponding to the subject.
Citation Information
Patent Citations
Cancer detection, classification, prognostication, therapy prediction and therapy monitoring using methylome analysis
US20210156863A1