Machine learning-enabled analysis of computed tomography and positron emission tomography scans for origin cell prediction.

A machine learning-based model for PET and CT scans addresses the limitations of invasive assays by accurately determining cancer cell origin, enhancing treatment selection and prognosis through non-invasive molecular subtype analysis.

JP2026516465APending Publication Date: 2026-05-25GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GENENTECH INC
Filing Date
2024-05-07
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Conventional techniques for determining the molecular subtype of heterogeneous diseases like diffuse large B-cell lymphoma require invasive assays and suffer from limited availability, performance, reproducibility, and resource constraints, making accurate and precise origin cell identification cumbersome.

Method used

A machine learning-based origin cell classification model is applied to non-invasive medical imaging modalities such as PET and CT scans to determine the origin cells of cancer, using models like recurrent neural networks and radiomic features to analyze lesions and molecular subtypes.

Benefits of technology

Enables accurate and precise non-invasive determination of molecular subtypes, facilitating effective treatment selection by identifying the origin cells of cancer cells, improving patient prognosis and treatment response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516465000001_ABST
    Figure 2026516465000001_ABST
Patent Text Reader

Abstract

The method may include receiving a positron emission tomography (PET) scan depicting multiple cancer cells. One or more lesions depicted in the PET scan may be identified. An origin cell classification model may be applied to determine the origin cells of each lesion depicted in the PET scan. A molecular subtype profile of multiple cancer cells depicted in the PET scan may be determined based on the origin cells of at least the individual lesions depicted in the PET scan. The molecular subtype profile may include the overall origin cells of the multiple cancer cells and / or the proportion of lesions having each possible origin cell. Related systems and computer program products are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Application No. 63 / 500,717, filed on 8 May 2023, entitled “MACHINE LEARNING ENABLED ANALYSIS OF COMPUTED TOMOGRAPHY AND POSITRON EMISSION TOMOGRAPHY SCANS FOR CELL-OF-ORIGIN PREDICTION,” the disclosure thereof is incorporated herein by reference in its entirety.

[0002] Technical field The subject matter described herein generally relates to machine learning, and more specifically to machine learning-based techniques for determining the origin cell (COO) based on positron emission tomography (PET) and computed tomography (CT) scans. [Background technology]

[0003] introduction Medical imaging refers to the techniques and processes for obtaining data that characterizes the internal anatomical form and pathophysiology of a subject, including images created by detecting radiation passing through the body (e.g., X-rays) or radiation emitted by administered radiopharmaceuticals (e.g., gamma rays from intravenously administered radiotracers). By revealing internal anatomical structures obscured by other tissues such as skin, subcutaneous fat, and bone, medical imaging is essential for numerous medical diagnoses and / or treatments. Examples of medical imaging modalities include two-dimensional imaging such as X-ray plain film, bone scintigraphy, and thermography. Examples of three-dimensional imaging modalities include magnetic resonance imaging (MRI), computed tomography (CT), cardiac sestamivi scanning, and positron emission tomography (PET). [Overview of the project]

[0004] overview Systems, methods, and products including computer program products are provided for machine learning-enabled analysis of positron emission tomography (PET) and computed tomography (CT) scans for determining the origin cell (COO) of cancer cells. Embodiments of this subject matter include, but are not limited to, methods according to the descriptions provided herein, and articles comprising tangibly embodied machine-readable media capable of operating one or more machines (e.g., computers) to cause operations that implement one or more of the features described herein. Similarly, computer systems that may include one or more processors and one or more memories coupled to the one or more processors are also described. Memories that may include non-temporary computer-readable or machine-readable storage media may encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implementations of one or more embodiments of this subject matter may be implemented by one or more data processors in a single computing system or in multiple computing systems. Such multiple computing systems can be connected via one or more connections, including, for example, connections through a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), or via direct connections between one or more of the multiple computing systems, and can exchange data and / or instructions or other commands, etc.

[0005] In one embodiment, a method is provided for machine learning-enabled analysis of positron emission tomography (PET) and computed tomography (CT) scans to determine the origin cell (COO) of cancer cells. The method may include: receiving a first positron emission tomography (PET) scan depicting multiple cancer cells; identifying a first lesion depicted in the first PET scan; applying an origin cell classification model to determine the first origin cell associated with the first lesion based on the first lesion depicted in the first PET scan; and determining the molecular subtype profiles of the multiple cancer cells depicted in the first PET scan based on the first origin cell of at least the first lesion.

[0006] In another embodiment, a system is provided for machine learning-enabled analysis of positron emission tomography (PET) and computed tomography (CT) scans for determining the origin cells (COO) of cancer cells. The system may include at least one processor and at least one memory. The at least one memory may contain program code that, when executed by at least one processor, brings about an operation. The operation may include receiving a first positron emission tomography (PET) scan depicting multiple cancer cells; identifying a first lesion depicted in the first PET scan; applying an origin cell classification model to determine the first origin cells associated with the first lesion based on the first lesion depicted in at least the first PET scan; and determining molecular subtype profiles of the multiple cancer cells depicted in the first PET scan based on the first origin cells of at least the first lesion.

[0007] In another embodiment, a computer program product is provided for machine learning-enabled analysis of positron emission tomography (PET) and computed tomography (CT) scans for determining the origin cells (COO) of cancer cells. The computer program product may include a non-temporary computer-readable medium that stores instructions that cause an action when executed by at least one data processor. The action may include receiving a first positron emission tomography (PET) scan depicting multiple cancer cells; identifying a first lesion depicted in the first PET scan; applying an origin cell classification model to determine the first origin cells associated with the first lesion based on the first lesion depicted in at least the first PET scan; and determining molecular subtype profiles of the multiple cancer cells depicted in the first PET scan based on the first origin cells of at least the first lesion.

[0008] In some variations of the method, system, and non-temporary computer-readable medium, one or more of the following features may be included in any feasible combination:

[0009] In some variations, the method may receive a first positron emission tomography (PET) scan depicting multiple cancer cells, identify a first lesion depicted in the first PET scan, apply an origin cell classification model to determine the first origin cell associated with the first lesion based on at least the first lesion depicted in the first PET scan, and determine the molecular subtype profile of the multiple cancer cells depicted in the first PET scan based on at least the first origin cell of the first lesion.

[0010] In some variations, the first origin cell of the first lesion may include, for each possible origin cell, the probability that one or more cancerous cells forming the first lesion originate from that origin cell.

[0011] In some variations, the method may identify a second lesion depicted in a first PET scan, apply an origin cell classification model to determine the second origin cells associated with the second lesion based on at least the second lesion depicted in the first PET scan, and further determine the molecular subtype profiles of multiple cancerous cells depicted in the first PET scan based on the second origin cells of the second lesion.

[0012] In some variations, the molecular subtype profiles of multiple cancer cells depicted in the first PET scan may include, for each possible origin cell, the probability that the overall origin cell of the multiple cancer cells originated from that origin cell.

[0013] In some variations, the probability that the overall origin of multiple cancerous cells is a specific origin cell may be the maximum, minimum, mean, median, and / or mode of the respective probabilities that each of the first and second lesions has that specific origin cell.

[0014] In some variations, the molecular subtype profiles of multiple cancerous cells depicted in the first PET scan may include, for each possible origin cell, the corresponding proportion of lesions containing that origin cell.

[0015] In some variations, the molecular subtype profiles of multiple cancer cells depicted in the first PET scan can be determined by generating an embedding sequence that includes at least the first origin cells of the first lesion and the second origin cells of the second lesion, and by applying a machine learning model to determine the overall origin cells of the multiple cancer cells in the first PET scan, based at least on the embedding sequence.

[0016] In some variations, the machine learning model can be a recurrent neural network.

[0017] In some variations, the method may extract a volume including a first lesion identified in a first PET scan, and apply an origin cell classification model to determine a first origin cell associated with the first lesion based at least on the volume extracted from the first PET scan.

[0018] In some variations, the volume including the first lesion may be extracted by at least determining the centroid of the first lesion within the first PET scan and extracting the volume based at least on the centroid of the first lesion.

[0019] In some variations, the volume may be a three-dimensional volume including a plurality of two-dimensional patches centered on the centroid of the first lesion.

[0020] In some variations, the plurality of two-dimensional patches may include a plurality of axial patches or a plurality of coronal patches.

[0021] In some variations, the first lesion may be identified by at least applying a segmentation model to identify a plurality of pixels corresponding to the first lesion within the first PET scan

[0022] In some variations, the method may extract a plurality of features associated with a first lesion identified in a first PET scan from the first PET scan, and apply an origin cell classification model to determine a first origin cell associated with the first lesion based at least on the plurality of features extracted from the first PET scan.

[0023] In some variations, the plurality of features may include the size of the first lesion, the shape of the first lesion, and / or the texture of the first lesion.

[0024] In some variations, the plurality of features may include one or more first-order statistics associated with one or more pixels depicting the first lesion in the first PET scan.

[0025] In some variations, the plurality of features can include a gray level co-occurrence matrix, a gray level size zone matrix, and / or a gray level run length matrix of one or more pixels depicting the first lesion in the first PET scan.

[0026] In some variations, the method can receive a computed tomography (CT) scan from the same time point as the first PET scan, identify the first lesion depicted in the CT scan, and apply an origin cell classification model to determine a first origin cell associated with the first lesion based on the first lesion depicted in the first PET scan and the CT scan.

[0027] In some variations, each pixel included in the CT scan can be associated with a value of intensity corresponding to the density of the tissue or the attenuation of the X-ray.

[0028] In some variations, each pixel of the first PET scan can be associated with an intensity value corresponding to the level of metabolic activity.

[0029] In some variations, the method can determine a tumor mask corresponding to the first lesion based on at least one of the CT scan and the first PET scan, and apply an origin cell classification model to determine the origin cell of the first lesion based on at least the tumor mask.

[0030] In some variations, the method can determine an organ mask corresponding to one or more organs depicted in the CT scan and the first PET scan based on at least one of the CT scan and the first PET scan, and apply an origin cell classification model to determine the origin cell of the first lesion based on at least the organ mask.

[0031] Details of one or more variations of the subject matter described herein are described in the accompanying drawings and the following description. Other features and advantages of the subject matter described herein will become apparent from the description and drawings, as well as from the claims. Certain features of the subject matter of this disclosure are described for illustrative purposes in relation to fluorodeoxyglucose-dependent (FDG-dependent) cancers such as non-Hodgkin lymphoma (NHL), but it should be readily understood that such features are not limiting. The claims following this disclosure define the scope of the subject matter protected. [Brief explanation of the drawing]

[0032] The accompanying drawings incorporated herein and constituting part of this specification illustrate specific aspects of the subject matter disclosed herein and, together with the description, help to illustrate some of the principles associated with the disclosed embodiments. In the drawings,

[0033] [Figure 1] A diagram of a system illustrating an example of a machine learning-based medical imaging analysis system, based on several exemplary embodiments, is drawn.

[0034] [Figure 2A] A schematic diagram illustrates an example of a process for predicting machine learning-enabled origin cells (COOs) using several exemplary embodiments.

[0035] [Figure 2B] A schematic diagram illustrates another example of the process for predicting machine learning-enabled origin cells (COOs) using several exemplary embodiments.

[0036] [Figure 3] A flowchart illustrating an example of the process for predicting machine learning-enabled origin cells is drawn using several exemplary embodiments.

[0037] [Figure 4A]A flowchart illustrating another example of the process for predicting machine learning-enabled origin cells is drawn using several exemplary embodiments.

[0038] [Figure 4B] A flowchart illustrating another example of the process for predicting machine learning-enabled origin cells is drawn using several exemplary embodiments.

[0039] [Figure 5] The graph illustrates a comparison of various performance metrics for different embodiments of the origin cell classification model across different clinical datasets, using several exemplary embodiments.

[0040] [Figure 6] A graph is drawn showing a comparison of receiver operating characteristic (ROC) curves of different molecular subtype classifications created by origin cell classification models across different test sets, using several exemplary embodiments.

[0041] [Figure 7] A graph illustrating the importance of various radiomic features for different embodiments of the origin cell classification model, using several exemplary embodiments, is drawn.

[0042] [Figure 8] A block diagram illustrating an example of a computing system in several exemplary embodiments is drawn.

[0043] In practical terms, similar reference numbers indicate similar structures, features, or elements. [Modes for carrying out the invention]

[0044] Detailed explanation Heterogeneous diseases, including cancers such as diffuse large B-cell lymphoma (DLBCL), are associated with complex etiologies and diverse molecular and cellular dysfunctions. Biological insights into disease heterogeneity can play a crucial role in patient prognosis and treatment selection. For example, in some cases, molecular subtypes derived based on the origin cells of cancer cells can serve as important biomarkers for predicting a patient's response to treatment and survival. In the case of diffuse large B-cell lymphoma (DLBCL), patients with the germinal center B-cell (GCB) molecular subtype tend to have a better prognosis than patients with the non-germinal center B-cell (non-GCB) or activated B-cell (ABC) molecular subtypes. Therefore, identifying a patient's molecular subtype may be essential for identifying effective treatment options.

[0045] Despite genetic and phenotypic variations, heterogeneous diseases such as diffuse large B-cell lymphoma frequently present with similar clinical manifestations. Therefore, determining the molecular subtype of heterogeneous diseases, such as the origin cells of cancer (e.g., diffuse large B-cell lymphoma), generally requires expensive procedures that are not common in standard clinical practice. Conventional techniques for determining origin cells (COO), such as ribonucleic acid sequencing (RNASeq) and immunohistochemistry (IHC)-based Hans algorithms, rely on invasive assays to extract tumor tissue. Subsequent analytical techniques suffer from limited availability (e.g., gene expression profiling (GEP)), limited performance, limited reproducibility, and large inter-leader variability (IHC). As a result, accurate and precise molecular subtype determination remains a cumbersome task beyond the resources of a typical cancer patient, even with conventional techniques for determining origin cells. To overcome the aforementioned limitations of conventional molecular subtyping techniques, the analysis controller may apply a machine learning-based origin cell classification model to determine the origin cell of multiple cancer cells based on at least one or more medical images depicting multiple cancer cells. For example, in some cases, a machine learning-based origin cell classification model may be applied to positron emission tomography (PET) scans and / or computed tomography (CT) scans depicting multiple cancer cells. Medical images such as positron emission tomography (PET) scans and computed tomography (CT) scans can be acquired non-invasively and are routinely acquired from cancer patients. Therefore, unlike conventional techniques, the origin cell classification models described herein may enable accurate and precise origin cell-based molecular subtype determination non-invasively and with minimal additional resources.

[0046] Various medical imaging modalities can be applied to obtain data characterizing the anatomical structures and pathophysiology of a subject's body. Computed tomography (CT) is an example of a three-dimensional imaging modality that captures a series of X-rays to create cross-sectional images (e.g., patches, slices, etc.) of bone, blood vessels, and soft tissues within the body. A computed tomography scan can be a three-dimensional volume formed by a series of two-dimensional images, where each pixel is associated with an intensity value indicating tissue density or X-ray attenuation at a corresponding location in the subject's body. Another example of a three-dimensional imaging modality is positron emission tomography (PET), which captures radioactive signals indicating cellular metabolic activity within a subject's body. A positron emission tomography scan can be a three-dimensional volume formed by a series of two-dimensional images, where each pixel is associated with an intensity value indicating the level of cellular metabolic activity (e.g., glucose uptake) at a corresponding location in the subject's body. In some cases, a single gantry incorporating both a positron emission tomography (PET) scanner and a computed tomography (CT) scanner may be able to acquire both PET and CT scans during the same session. The acquired PET and CT scans can be combined into a single superimposed (e.g., jointly aligned) image (e.g., a PET-CT scan) in which the spatial distribution of metabolic activity depicted in the PET scan is aligned with the anatomical structures depicted in the CT scan.

[0047] In some exemplary embodiments, the analysis controller may determine the origin cells of multiple cancer cells depicted in a positron emission tomography (PET) scan, based on at least the PET scan. Multiple cancer cells may correspond to one or more lesions depicted in the PET scan. Therefore, in some cases, the analysis controller may apply an origin cell classification model to determine the primary origin cell of a first lesion, based on at least the primary lesion depicted in the PET scan. Furthermore, in some cases, the origin cell classification model may be applied to determine the secondary origin cell of a second lesion, based on at least the secondary lesion depicted in the PET scan. The molecular subtype profiles of multiple cancer cells depicted in the PET scan may be determined based on the primary origin cell of the first lesion and the secondary origin cell of the second lesion. For example, in some cases, the molecular subtype profiles of multiple cancer cells may include the overall origin cells of the multiple cancer cells determined based on the primary origin cell of the first lesion and the secondary origin cell of the second lesion. Alternatively and / or additionally, the molecular subtype profile of multiple cancer cells may include the proportion of different origin cells present in the multiple cancer cells. This proportion of different origin cells can be determined based on at least the primary origin cells of the primary lesion and the secondary origin cells of the secondary lesion.

[0048] In some exemplary embodiments, the origin cell classification model may determine the probability that, for each of the first and second lesions, the constituent cancer cells have each of several different possible origin cells. For example, in diffuse large B-cell lymphoma (DLBCL), the origin cell classification model may determine, for each of the first and second lesions, the first probability that the constituent cancer cells have a first origin cell (e.g., germinal center B cells (GCB)) and the second probability that the constituent cancer cells have a second origin cell. The overall origin cells of multiple cancer cells may include the maximum, minimum, mean, median, and / or mode of the first probability for each lesion having a first origin cell. In some cases, the overall origin cells of multiple cancer cells may also include the maximum, minimum, mean, median, and / or mode of the second probability for each lesion having a second origin cell. Furthermore, in some cases, the overall origin cell of multiple cancerous cells may be determined based on one or more other characteristics of the first and second lesions, including, for example, dimensions (e.g., length, width, volume) and location (e.g., spatial coordinates, distance to other lesions). In some cases, the probability of each lesion having a specific origin cell may be weighted based on one or more additional characteristics, for example, when determining the probability of the overall origin cell being that specific origin cell. Furthermore, in some cases, the overall origin cell of multiple cancerous cells may be identified as a specific origin cell (e.g., germinal center B cell (GCB) or activated B cell (ABC) (or non-germinal center B cell (non-GCB)) if the probability of a threshold amount (e.g., percentage, ratio, etc.) of the lesion associated with that specific origin cell satisfies one or more thresholds.

[0049] In some exemplary embodiments, the overall origin of multiple cancer cells depicted in a positron emission tomography (PET) scan may be determined for each lesion present in the PET scan by at least applying a machine learning model (e.g., a neural network such as a recurrent neural network (RNN)) that manipulates an embedding sequence that includes the probability that each constituent cancer cell has each possible origin. In some cases, the embedding sequence may also include one or more other characteristics of each lesion, such as one or more dimensions of each lesion (e.g., length, width, volume) and the location of each lesion (e.g., spatial coordinates, distance to other lesions).

[0050] Alternatively and / or additionally, molecular subtype profiles of multiple cancer cells depicted in positron emission tomography (PET) scans may include one or more metrics indicating heterogeneity and / or homogeneity among different origin cells contained within the multiple cancer cells of the primary and secondary origin cells of the secondary lesions. For example, in some cases, the molecular subtype profile of multiple cancer cells may include the proportion (e.g., percentage, ratio, etc.) of lesions identified as having each possible origin cell. For example, in the case of diffuse large B-cell lymphoma (DLBCL), the molecular subtype profile of multiple cancer cells may include a primary proportion of lesions identified as germinal center B cells (GCBs) and a secondary proportion of lesions identified as activated B cells (ABCs) (or non-germinal center B cells (non-GCBs)).

[0051] In some exemplary embodiments, the origin cells of each lesion depicted by a positron emission tomography (PET) scan may be determined based on the corresponding volume extracted from the PET scan. For example, in some cases, the volume extracted from the PET scan may be a three-dimensional volume formed from a series of two-dimensional patches (e.g., axial patches, coronal patches, etc.). Thus, the analysis controller may extract from the PET scan a first volume centered on the first centroid of a first lesion, and a second volume centered on the second centroid of a second lesion. In some cases, the analysis controller may apply an origin cell classification model to determine the first origin cells of the first lesion based on at least the first volume extracted from the PET scan. Furthermore, in some cases, the analysis controller may apply an origin cell classification model to determine the second origin cells of the second lesion based on the second volume last extracted from the PET scan.

[0052] In some exemplary embodiments, the origin cells of each lesion depicted in a positron emission tomography (PET) scan may be determined based on one or more corresponding radiomic features extracted from the PET scan. For example, in some cases, the analysis controller may extract the size, shape, and / or texture of the lesion for each lesion depicted in the PET scan. Alternatively and / or additionally, the analysis controller may extract one or more primary statistics, a gray-level co-occurrence matrix, a gray-level size zone matrix, and / or a gray-level run-length matrix for each lesion depicted in the PET scan. In some cases, the analysis controller may apply an origin cell classification model to determine the first origin cells of a first lesion, based on at least a plurality of first features of the first lesion extracted from the PET scan. Furthermore, in some cases, the analysis controller may apply an origin cell classification model to determine the second origin cell of the second lesion, based at least on a second set of features of the second lesion extracted from the second positron emission tomography (PET) scan.

[0053] In some exemplary embodiments, the analysis controller may determine the molecular subtypes (or overall origin cells) of multiple cancer cells based on multiple modalities of medical images from the same time point. For example, in some cases, the molecular subtypes (or overall origin cells) of multiple cancer cells may be determined based on positron emission tomography (PET) scans depicting multiple cancer cells and computed tomography (CT) scans from the same time point. That is, in some cases, the first origin cells of a first lesion and the second origin cells of a second lesion may be determined based on positron emission tomography (PET) scans and corresponding computed tomography (CT) scans. Furthermore, in some cases, the first origin cells of a first lesion and the second origin cells of a second lesion may be determined by applying an origin cell classification model to tumor masks of the first and second lesions determined based on at least positron emission tomography (PET) scans and corresponding computed tomography (CT) scans. Alternatively and / or additionally, the primary origin cells of the primary lesion and the secondary origin cells of the secondary lesion can be determined by applying an origin cell classification model to an organ mask of at least one organ determined based on positron emission tomography (PET) scans and corresponding computed tomography (CT) scans.

[0054] In some exemplary embodiments, the analysis controller may determine molecular subtype profiles for multiple cancer cells based on positron emission tomography (PET) scans from multiple time points. For example, the analysis controller may, in some cases, apply an origin cell classification model to determine the first origin cell of a first lesion based on the first lesion depicted in a first positron emission tomography (PET) scan from at least a first time point and a second positron emission tomography (PET) scan from a second time point. Furthermore, in some cases, the analysis controller may,

[0055] An origin cell classification model may be applied to determine the primary origin cells of a first lesion based on the first computed tomography (CT) scan depicted in the first computed tomography (CT) scan from the same time point as the first positron emission tomography (PET) scan, and the second computed tomography (CT) scan from the same time point as the second positron emission tomography (PET) scan. In some cases, the molecular subtype profile of multiple cancer cells may include multiple overall origin cells determined based at least on the primary origin cells of the first lesion and the secondary origin cells of the second lesion. Alternatively or additionally, the molecular subtype profile of multiple cancer cells may include determining the proportion of different origin cells present in the positron emission tomography (PET) scans based at least on the primary origin cells of the first lesion and the secondary origin cells of the second lesion.

[0056] In some exemplary embodiments, the analysis controller may further determine the diagnosis of the patient's disease associated with the positron emission tomography (PET) scan, the prognosis of the disease, the progression of the disease, the treatment, and / or the response to treatment, based on the molecular subtype profiles associated with multiple cancer cells depicted in the positron emission tomography (PET) scan. For example, in some cases, multiple cancer cells may be associated with diffuse large B-cell lymphoma (DLBCL). Thus, the origin cell classification model may be trained to determine, for each lesion present in the positron emission tomography (PET) scan, the first probability that the constituent cancer cells are germinal center B cells (GCBs) and the second probability that the constituent cancer cells are activated B cells (ABCs) (or non-germinal center B cells (non-GCBs)). The molecular subtype profiles of the multiple cancer cells may be determined based on the first probability that each lesion is germinal center B cells (GCBs) and the second probability that the constituent cancer cells are activated B cells (ABCs) (or non-germinal center B cells (non-GCBs)). For example, in some cases, the molecular subtype profiles of multiple cancer cells may include the proportion of the overall origin cells and / or different origin cells present in the multiple cancer cells depicted in the positron emission tomography (PET) scan.

[0057] In some exemplary embodiments, the origin cell classification model may include one or more of the following: artificial neural networks (ANNs) (e.g., visual transformer models such as visual transformer models with shifted patch tokenization and localized autoattention), tree-based classifiers (e.g., gradient boosted decision trees, random forests, extreme gradient boosted decision trees (XGBoost)), ridge classifiers, etc. The origin cell classification model may be trained on a training set that includes one or more annotated training samples. For example, in some cases, the training set may be generated to include a first training sample having a first volume containing lesions depicted by positron emission tomography (PET) scans and a first ground truth annotation of the origin cells of the lesions. Furthermore, in some cases, the training set may be generated to include a second training sample having a second volume containing the same lesions depicted by positron emission tomography (PET) scans and a second ground truth annotation of the origin cells of the lesions. The second volume may be generated by modifying the first volume, for example, by one or more of the following: normalization, rotation, inversion, and zoom modification of the first volume. In some cases, modifications to the primary volume may be limited to one or more slices of the primary volume that are within a threshold distance of the centroid of the lesion contained within the primary volume.

[0058] Figure 1 shows a diagram of a system illustrating an example of a machine learning-based medical imaging analysis system 100 according to several exemplary embodiments. Referring to Figure 1, the machine learning-based medical imaging analysis system 100 may include an analysis controller 110, one or more imaging devices 120, and a client device 130. As shown in Figure 1, the analysis controller 110, one or more imaging devices 120, and the client device 130 may be communicatively coupled via a network 140. One or more imaging devices 120 may include, for example, a computed tomography (CT) scanner 121 and a positron emission tomography (PET) scanner 123. The client device 130 may be a processor-based device, including, for example, a smartphone, a tablet computer, a wearable device, a virtual assistant, or an Internet of Things (IoT) device. The network 140 may be a wired network and / or a wireless network, including, for example, a wide area network (WAN), a local area network (LAN), a virtual local area network (VLAN), a public land mobile network (PLMN), or the internet.

[0059] In the example shown in Figure 1, the analysis controller 110 may include an extraction engine 111 containing a segmentation model 112, an origin cell classification model 113, and an analysis engine 115. In some exemplary embodiments, the analysis controller 110 may determine the molecular subtype profile of one or more lesions formed by multiple cancer cells based on at least a positron emission tomography (PET) scan depicting multiple cancer cells. In some cases, the analysis controller 110 may further determine the molecular subtype profile of one or more lesions based on computed tomography (CT) scans from the same time point. The positron emission tomography (PET) scans and computed tomography (CT) scans may be generated by one or more imaging devices 120 (e.g., computed tomography scanner 121, positron emission tomography (PET) scanner 123, etc.).

[0060] In some cases, the analysis controller 110 may apply an origin cell classification model 113 that can be trained to determine the origin cell of each lesion. For example, the origin cell classification model 113 may be based on positron emission tomography (PET) scans and, optionally, computed tomography (CT) scans from the same time point.

[0061] It can be applied to determine the primary origin cells of the primary lesion. Furthermore, the origin cell classification model 113 can be applied to determine the secondary origin cells of the secondary lesion based on positron emission tomography (PET) scans and, optionally, computed tomography (CT) scans from the same time point. As will be described in more detail, in some cases, the primary origin cells of the primary lesion and the secondary origin cells of the secondary lesion can be determined based on the volume containing the primary and secondary lesions extracted by the extraction engine 111. Alternatively, the primary origin cells of the primary lesion and the secondary origin cells of the secondary lesion can be determined based on the features extracted by the extraction engine 111. Furthermore, the evaluation engine 115 can determine molecular subtype profiles based on the primary origin cells of the primary lesion and the secondary origin cells of the secondary lesion.

[0062] Figure 2A shows a schematic diagram illustrating an example of a process 200 for predicting machine learning-enabled origin cells, according to several exemplary embodiments. Referring to Figure 2A, the analysis controller 110 may receive a positron emission tomography (PET) scan 210 and, optionally, a computed tomography (CT) scan 220 from one or more imaging devices 120. For example, the analysis controller 110 may, depending on the case, determine molecular subtype profiles 180 of multiple cancer cells depicted in the positron emission tomography (PET) scan 210 based on the PET scan 210. Alternatively, the analysis controller 110 may determine the molecular subtype profiles 180 based on both the positron emission tomography (PET) scan 210 and the computed tomography (CT) scan 220. Depending on the case, the positron emission tomography (PET) scan 210 and the computed tomography (CT) scan 220 may be from the same point in time. Therefore, the positron emission tomography (PET) scan 210 can be superimposed (or aligned together) with the computed tomography (CT) scan 220 such that each pixel of the positron emission tomography (PET) scan 210 is mapped to the corresponding pixel of the computed tomography (CT) scan 220.

[0063] In some exemplary embodiments, the analysis controller 110 may determine molecular subtype profiles 180 of multiple cancer cells depicted in the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220 based on the origin cells of each individual lesion contained in one or more volumes extracted from the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220. For example, multiple cancer cells depicted in the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220 may form one or more lesions.

[0064] Therefore, each volume extracted from the positron emission tomography (PET) scan 210, and optionally a jointly aligned computed tomography (CT) scan 220, may be a three-dimensional volume having multiple two-dimensional patches (e.g., 96 pixels × 96 pixels or different sizes) centered on the centroid of the corresponding lesion. In the example of process 200 shown in Figure 2A, the analysis controller 110 may include an extraction engine 111 configured to extract a first volume 230a containing a first lesion and a second volume 230b containing a second lesion from the positron emission tomography (PET) scan and optionally a jointly aligned computed tomography (CT) scan 220.

[0065] In some exemplary embodiments, the extraction engine 111 may extract a first volume 230a and a second volume 230b by identifying at least the corresponding lesions within the positron emission tomography (PET) scan 210 and optionally within a jointly aligned computed tomography (CT) scan 220. For example, optionally, the extraction engine 111 may segment within the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220 to identify a first set of pixels corresponding to the first lesion and a second set of pixels corresponding to the second lesion (for example, by applying a segmentation model 112 such as an artificial neural network). Furthermore, the extraction engine 111 may determine the first centroid of the first lesion and the second centroid of the second lesion. A first volume 230a extracted from a positron emission tomography (PET) scan 210 and, optionally, a jointly aligned computed tomography (CT) scan 220 may contain a first set of two-dimensional patches (e.g., axial patches, coronal patches, etc.) centered on the first centroid of the first lesion. A second volume 230b, on the other hand, may contain a second set of two-dimensional patches (e.g., axial patches, coronal patches, etc.) centered on the second centroid of the second lesion.

[0066] Referring again to Figure 2A, in some exemplary embodiments, the analysis controller 110 may apply the origin cell classification model 113 to determine the first origin cells 240a of a first lesion contained in the first volume 230a, based on at least the first volume 230a. Also, in the example of process 200 shown in Figure 2A, the analysis controller 110 may apply the origin cell classification model 113 to determine the second origin cells 240b of a second lesion contained in the second volume 230b, based on at least the second volume 230b. In some cases, the first origin cells 240a of a first lesion may include the probability that the first lesion depicted in the first volume 230a has the corresponding origin cells for each possible origin cell (e.g., origin cell label), while the second origin cells 240b of a second lesion may include the probability that the second lesion depicted in the second volume 230b has the corresponding origin cells for each possible origin cell (e.g., origin cell label). For example, in the case of diffuse large B-cell lymphoma (DLBCL), the first origin cells 240a may include a first probability that the first lesion is germinal center B cells (GCB) and / or a second probability that the first lesion is activated B cells (ABC) (or non-germinal center B cells (non-GCB)). Similarly, the second origin cells 240b may include a third probability that the second lesion is germinal center B cells (GCB) and / or a fourth probability that the second lesion is activated B cells (ABC) (or non-germinal center B cells (non-GCB)). As shown in Figure 2A, the analysis controller 110 may further include an evaluation engine 115 that determines molecular subtype profiles 180 of multiple cancer cells depicted in the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220, based on at least the first origin cells 240a of the first lesion and the second origin cells 240b of the second lesion.For example, in the case of diffuse large B-cell lymphoma (DLBCL), the molecular subtype profile 180 can be determined based on at least the first probability that the first lesion is germinal center B-cell (GCB), the second probability that the first lesion is activated B-cell (ABC) (or non-germinal center B-cell (non-GCB)), the third probability that the second lesion is germinal center B-cell (GCB), and / or the fourth probability that the second lesion is activated B-cell (ABC) (or non-germinal center B-cell (non-GCB)).

[0067] In some exemplary embodiments, instead of extracting one or more volumes centered on the centroid of one or more corresponding lesions, the extraction engine 111 may extract one or more features of each lesion from a positron emission tomography (PET) scan 210 and optionally a jointly aligned computed tomography (CT) scan 220. For further illustration, Figure 2B depicts another example of a process 250 in which the extraction engine 111 extracts a first plurality of features 235a of a first lesion and a second plurality of features 235b of a second lesion from a positron emission tomography (PET) scan 210 and optionally a jointly aligned computed tomography (CT) scan 220. The first plurality of features 235a and the second plurality of features 235b may include one or more radiomic features. For example, the first plurality of features 235a and the second plurality of features 235b may each include the size, shape, and / or texture of the corresponding lesion.

[0068] Alternatively and / or additionally, the first set of features 235a and the second set of features 235b may also include, for each corresponding lesion, one or more primary statistics such as the range, maximum, minimum, median, mode, and / or mean pixel values ​​of one or more pixels depicting the lesion in a positron emission tomography (PET) scan 210 and, optionally, a jointly aligned computed tomography (CT) scan 220. In the case of the positron emission tomography (PET) scan 210, these primary statistics may correspond to primary statistics of the level of metabolic activity (e.g., standardized uptake value (SUV)) indicated by the lesion. On the other hand, in the case of the computed tomography (CT) scan 220, these primary statistics may correspond to primary statistics of tissue density (or X-ray attenuation) observed across the lesion. In some cases, the first set of features 235a and the second set of features 235b may further include, for the corresponding lesion, a gray level co-occurrence matrix, a gray level size zone matrix, and / or a gray level run-length matrix for one or more pixels depicting the lesion in a positron emission tomography (PET) scan 210 and optionally a jointly aligned computed tomography (CT) scan 220.

[0069] Referring again to Figure 2B, the analysis controller 110 may apply the origin cell classification model 113 to determine the first origin cells 240a of the first lesion based on at least a first set of features 23Sa. Furthermore, the analysis controller 110 may apply the origin cell classification model 113 to determine the second origin cells 240b of the second lesion based on at least a second set of features 235b. As shown in Figure 2B, the evaluation engine 115 may then determine the molecular subtype profiles 180 of the multiple cancerous cells depicted in the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220 based on at least the first origin cells 240a of the first lesion and the second origin cells 240b of the second lesion.

[0070] In some exemplary embodiments, the evaluation engine 115 may determine a molecular subtype profile 180 of a plurality of cancerous cells to include an overall originating cell determined based on the originating cells of each lesion present in a positron emission tomography (PET) scan 210 and a computed tomography (CT) scan 220. For example, in FIGS. 2A-2B, the overall originating cell included in the molecular subtype profile 180 may be determined based on a first originating cell 240a of a first lesion and a second originating cell 240b of a second lesion. In some cases, the overall originating cell of a plurality of cancerous cells may be determined based on the following formula (1). For example, for each possible originating cell (e.g., originating cell label) m, the overall originating cell may include the corresponding probability P m thereof. That is, in some cases, the overall originating cell of a plurality of cancerous cells may be represented as (P0,..., P m ). In the case of diffuse large B-cell lymphoma (DLBCL), the overall originating cell of a plurality of cancerous cells may be represented as (P GCB , P ABC ), where P GCB represents the probability that the overall originating cell is a germinal B cell, and P ABC represents the probability that the overall originating cell is an activated B cell (or non-germinal center B cell (non-GCB)).

[0071] According to the following formula (1), the probability P m associated with the originating cell m may be determined based on the probability P n that each lesion n is the originating cell m. For example, formula (1) shows an example where the probability Pm associated with the originating cell m is the average of the individual probabilities (p0,..., p n ). Alternatively, the probability P m associated with the originating cell m may be the maximum value, minimum value, median value, and / or most frequent value of the probabilities (p_{0},..., p n ) associated with the individual lesions n. The following formula (1) may, in some cases, take into account that the probability p n of each lesion n corresponds to a weight w nIt is further shown that these can be weighted by the following: Examples of such characteristics include one or more dimensions of each lesion (e.g., length, width, volume), the location of each lesion (e.g., spatial coordinates, distance to other lesions), etc. In some cases, the overall origin cells of multiple cancerous cells may be identified as a specific origin cell if the probability of a threshold amount of lesions associated with that particular origin cell (e.g., percentage, ratio, etc.) satisfies one or more thresholds. For example, in diffuse large B-cell lymphoma (DLBCL), the overall origin cells of multiple cancerous cells may be identified as germinal center B cells (GCBs) if more than 25% of the lesions depicted in positron emission tomography (PET) scans 210 and computed tomography (CT) scans 220 indicate the possibility that more than 50% are germinal center B cells (GCBs). TIFF2026516465000002.tif9170

[0072] In some cases, the evaluation engine 115 may apply a machine learning model (e.g., a recurrent neural network) to determine the overall origin of multiple cancer cells. For example, for each lesion, the evaluation engine 115 may generate an embedding sequence that includes the probability of each possible origin of the constituent cancer cells. In the case of diffuse large B-cell lymphoma (DLBCL), the embedding sequence may include, for each lesion, a first probability that the constituent cancer cells are embryonic B cells (GCBs) and a second probability that the constituent cancer cells are activated B cells (ABCs) (or non-embryonic B cells (non-GCBs)). To determine the overall origin of multiple cancer cells for inclusion in the molecular subtype profile 180, a machine learning model (e.g., a recurrent neural network) may be applied to the embedding sequence. In some cases, the embedding sequence may be generated to include one or more additional characteristics of each lesion, such as the dimensions of each lesion (e.g., length, width, volume) and the location of each lesion (e.g., spatial coordinates, distance to other lesions).

[0073] In some cases, in addition to or instead of the overall origin cells, the evaluation engine 115 may determine the molecular subtype profile 180 to include one or more metrics indicating heterogeneity and / or homogeneity of origin cells for different lesions (e.g., first origin cells 240a for the first lesion and second origin cells 240b for the second lesion). For example, in some cases, the molecular subtype profile 180 may include, for each possible origin cell, the proportion (e.g., percentage, ratio, etc.) of lesions identified as presenting the origin cells. In the case of diffuse large B-cell lymphoma (DLBCL), the molecular subtype profile 180 may include a first proportion of lesions identified as germinal center B cells (GCBs) and a second proportion of lesions identified as activated B cells (ABCs) (or non-germinal center B cells (non-GCBs)).

[0074] In some exemplary embodiments, the origin cell classification model 113 may include one or more machine learning models trained to determine the origin cell associated with a lesion based on a volume containing one or more features of the lesion extracted from the lesion and / or positron emission tomography (PET) scans and optionally computed tomography (CT) scans from the same time point. For example, in some cases, the origin cell classification model 113 may include one or more of the following: artificial neural networks (ANNs) (e.g., visual transformer models such as visual transformer models with shifted patch tokenization and local autoattention), tree-based classifiers (e.g., gradient boosted decision trees, random forests, extreme gradient boosted decision trees (XGBoost)), ridge classifiers, etc. In some cases, the origin cell classification model 113 may include visual transformer models implemented with shifted patch tokenization and local autoattention to increase the locality-inducing bias of the visual transformer model (e.g., assumptions about relationships between nearby pixels) and enable the visual transformer model to learn from a limited number of training data.

[0075] If the origin cell classification model 113 operates on positron emission tomography (PET) scans without corresponding computed tomography (CT) scans from the same time point, the origin cell classification model 113 may include one or more tree-based classifiers (e.g., gradient-boosted decision trees, random forests, extreme gradient-boosted decision trees (XGBoost), etc.). For example, in some cases, the analysis engine 110 may apply the origin cell classification model 113 to determine the origin cells of the corresponding lesion depicted in the positron emission tomography scan, based on at least one or more radiomic features extracted from the tumor mask and the positron emission tomography scan. In this regard, the tumor mask may include multiple pixels, each of which has a first value (e.g., "1") indicating that the pixel is part of the lesion or a second value (e.g., "0") indicating that the pixel is not part of the lesion. One or more radiomic features may include, for example, the size, shape, and / or texture of the lesion.

[0076] Alternatively and / or additionally, one or more radiomic features may include one or more primary statistics for one or more pixels of a positron emission tomography (PET) scan depicting a lesion, a gray-level co-occurrence matrix, a gray-level size zone matrix, and / or a gray-level run-length matrix.

[0077] In some exemplary embodiments, the origin cell classification model 113 may be trained on a training set comprising one or more annotated training samples. For example, the training set may be generated to include a first training sample having a first volume containing lesions depicted by positron emission tomography (PET) scans and a first ground truth annotation of the origin cells of the lesions. Furthermore, the training set may be generated to include a second training sample having a second volume containing the same lesions depicted by positron emission tomography (PET) scans and a second ground truth annotation of the origin cells of the lesions. In some cases, the second training sample may be generated by applying one or more data augmentation techniques. For example, in some cases, the second volume may be generated by modifying the first volume, for example, by changing one or more of the first volume's normalization, rotation, inversion, and zoom. Furthermore, in some cases, the modification of the first volume may be limited to one or more slices of the first volume that are within a threshold distance of the centroids of the lesions contained in the first volume.

[0078] Figure 3 shows a flowchart illustrating an example of Process 300 for predicting machine learning-enabled origin cells, based on several exemplary embodiments.

[0079] Referring to Figure 3, process 300 may be performed by the analysis controller 110 to determine molecular subtype profiles 180 of multiple cancer cells depicted by positron emission tomography (PET) scans and, optionally, corresponding computed tomography (CT) scans from the same time point.

[0080] In 302, the analysis controller 110 may receive positron emission tomography (PET) scans and computed tomography (CT) scans depicting multiple cancer cells. For example, the analysis controller 110 may receive a positron emission tomography (PET) scan 210 and, optionally, a computed tomography (CT) scan 220 from one or more imaging devices 120. When the analysis controller 110 receives a positron emission tomography (PET) scan 210 and a computed tomography (CT) scan 220, the positron emission tomography (PET) scan 210 and the computed tomography (CT) scan 220 may be from the same point in time. In those examples, a positron emission tomography (PET) scan 210 and a computed tomography (CT) scan 220 are superimposed (or aligned together) so that pixels in the PET scan 210, whose values ​​indicate levels of metabolic activity (e.g., standard uptake values ​​(SUV)), are mapped to pixels in the CT scan 220, whose values ​​indicate tissue density (or X-ray attenuation).

[0081] In 304, the analysis controller 110 can determine molecular subtype profiles of multiple cancer cells based on at least positron emission tomography (PET) scans and computed tomography (CT) scans. In some exemplary embodiments, the analysis controller 110 may apply an origin cell classification model 113 to determine the corresponding origin cell for each lesion identified within the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220. As will be described in more detail below, the origin cell classification model 113 may be applied to determine the origin cell for each lesion based on the volume containing the lesion extracted from the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220. Alternatively, the origin cell classification model 113 may be applied to determine the origin cell for each lesion based on one or more features of the lesion extracted from the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220. The molecular subtype profiles 180 of multiple cancerous cells depicted in positron emission tomography (PET) scans 210 and computed tomography (CT) scans 220 can be determined based on the origin cells of each lesion.

[0082] In 306, the analysis controller 110 may determine one or more of the following based on molecular subtype profiles of at least multiple cancer cells: diagnosis of disease, prognosis of disease, disease progression, treatment, and treatment response. In some exemplary embodiments, the analysis controller 110 may determine one or more of the following based on at least molecular subtype profiles 180 of multiple cancer cells depicted in positron emission tomography (PET) scans 210 and computed tomography (CT) scans 220. For example, in some cases, multiple cancer cells may be associated with diffuse large B-cell lymphoma (DLBCL), and the molecular subtype profiles 180 of the multiple cancer cells may indicate the proportion of lesions showing the overall origin cells (e.g., germinal center B cells (GCB) or activated B cells (ABC) (or non-germinal center B cells (non-GCB))) and / or each possible origin cell. Furthermore, in some cases, the analysis controller 110 may determine the patient's overall survival (OS) and / or progression-free survival (PFS) associated with positron emission tomography (PET) scans 210 and computed tomography (CT) scans 220 depicting multiple cancer cells, based at least on the molecular subtype profiles 180 of the multiple cancer cells. Alternatively and / or additionally, the analysis controller 110 may determine whether the patient is suitable for treatment (e.g., the probability that the patient is a responder (or non-responder) to treatment), based at least on the molecular subtype profiles 180 of the multiple cancer cells.

[0083] Figure 4A shows a flowchart illustrating an example of process 400 for machine learning-enabled origin cell prediction in several exemplary embodiments. Referring to Figure 4A, process 400 may be performed by the analysis controller 110 to determine molecular subtype profiles 180 of multiple cancer cells depicted by positron emission tomography (PET) scans and, optionally, corresponding computed tomography (CT) scans from the same time point. Optionally, process 400 may perform operation 304 of process 300.

[0084] In some cases, the analysis controller 110 may, as part of performing process 400, apply the origin cell classification model 113 to the positron emission tomography (PET) scan 210 and, optionally, the computed tomography (CT) scan 220 to determine the origin cells of each lesion depicted therein. The origin cell classification model 113 may be a machine learning model (e.g., a visual transducer, a tree-based classifier, a ridge classifier, etc.) that can distinguish lesions having different origin cells (e.g., germinal center B cells (GCB) and activated B cells (ABC) (or non-germinal center B cells (non-GCB)) in diffuse large B-cell lymphoma (DLBCL)) based on features (e.g., radiomic features) present in the positron emission tomography (PET) scan 210 and, optionally, the computed tomography (CT) scan 220.

[0085] As the performance metrics shown in Figures 5-6 demonstrate, the machine learning-based origin cell classification model 113 applied by the analysis controller 110 when running process 400 can accurately distinguish between lesions having different origin cells (e.g., germinal center B cells (GCBs) and activated B cells (ABCs) (or non-germinal center B cells (non-GCBs) in diffuse large B-cell lymphoma (DLBCL))). The fact that the origin cell classification model 113 can determine origin cells non-invasively, based on medical images acquired as part of routine patient care (e.g., positron emission tomography (PET) scans, and possibly corresponding computed tomography (CT) scans), means that process 400 can be run to perform accurate and precise origin cell determination non-invasively and with minimal additional resources. In contrast, conventional techniques for determining origin cells (COOs) require invasive assays and data analysis to extract tumor tissue, which are accompanied by either limited availability or limited performance, limited reproducibility, and high inter-leader variability.

[0086] Furthermore, the analysis controller 110 may perform process 400 to generate multiple molecular subtype profiles 180 of cancerous cells based on data associated with individual lesions. If multiple lesions are present in the positron emission tomography (PET) scan 210 and computed tomography (CT) scan 220, the molecular subtype profiles 180 can explain the different origin cells that may be present in the different lesions. In contrast, the aforementioned conventional techniques rely on data associated with a single tumor sample, which means that conventional techniques cannot capture valuable biological insights such as heterogeneity and / or homogeneity of origin cells across different lesions.

[0087] In 402, the analysis controller 110 may identify one or more lesions within positron emission tomography (PET) scans and computed tomography (CT) scans depicting multiple cancer cells. In some exemplary embodiments, the analysis controller 110 may segment the positron emission tomography (PET) scan 210 and, optionally, the computed tomography (CT) scan 220 to identify one or more lesions depicted therein. For example, the analysis controller 110 may identify one or more lesions by applying a segmentation model 112 (e.g., an artificial neural network) trained to assign labels to each pixel of the positron emission tomography (PET) scan 210 and / or the computed tomography (CT) scan 220, having a first value indicating that the pixel depicts a lesion and a second value indicating that the pixel does not depict a lesion. In some cases, pixels in a positron emission tomography (PET) scan 210 and corresponding pixels in a co-aligned computed tomography (CT) scan 220 may be identified as depicting a lesion if the values ​​of these pixels meet one or more corresponding thresholds (e.g., for metabolic activity levels, tissue density, and / or X-ray attenuation). Alternatively and / or additionally, the analysis controller 110 may identify one or more lesions by first applying thresholds to identify one or more objects present in the positron emission tomography (PET) scan 210 and the corresponding computed tomography (CT) scan 220 before applying one or more machine learning models (e.g., logistic regression models, tree-based classifiers, fully connected neural networks, etc.) trained to perform object classification.

[0088] In 404, the analysis controller 110 may extract volumes for each of one or more lesions from positron emission tomography (PET) scans and computed tomography (CT) scans. In some exemplary embodiments, the analysis controller 110 may extract a first volume 230a containing a first lesion and a second volume 230b containing a second lesion from a positron emission tomography (PET) scan 210 and optionally a computed tomography (CT) scan 220. Each of the first volume 230a and the second volume 230b may be a three-dimensional volume having multiple two-dimensional patches (e.g., axial patches, coronal patches, etc.). In some cases, the analysis controller 110 may extract a first set of two-dimensional patches centered on the first centroid of the first lesion in order to extract the first volume 230a containing the first lesion. The analysis controller 110 may also extract a second set of two-dimensional patches centered on the second centroid of the second lesion in order to extract the second volume 230b.

[0089] In 406, the analysis controller 110 may apply an origin cell classification model 113 to determine the origin cells of a corresponding lesion based on each volume extracted from positron emission tomography (PET) scans and computed tomography (CT) scans. In some exemplary embodiments, the origin cell classification model 113 may be trained to determine the origin cells of a lesion based on at least a three-dimensional volume containing the lesion from a positron emission tomography (PET) scan 210, and optionally a corresponding three-dimensional volume containing the lesion from a jointly aligned computed tomography (CT) scan 220. The intensity values ​​of pixels within the three-dimensional volume extracted from the positron emission tomography (PET) scan 210 may correspond to levels of cellular metabolic activity (e.g., standardized uptake values ​​(SUV) corresponding to glucose intake), while the intensity values ​​of pixels within the three-dimensional volume extracted from the computed tomography (CT) scan 220 may correspond to tissue density or X-ray attenuation. Therefore, in some cases, the origin cell classification model 113 can be trained to determine the origin cells of a lesion based on the level of metabolic activity (e.g., standardized uptake value (SUV)), tissue density, and / or X-ray attenuation indicated by the lesion and its surrounding environment. For example, in some exemplary embodiments, the analysis controller 110 may apply the origin cell classification model 113 to determine the first origin cells 240a of a first lesion contained in the first volume 230a, based on at least the first volume 230a. The analysis controller 110 may also apply the origin cell classification model 113 to determine the second origin cells 240b of a second lesion contained in the second volume 230b, based on at least the second volume 230b.

[0090] In 408, the analysis controller 110 may determine molecular subtype profiles for multiple cancer cells based on the origin cells of each lesion present in at least the positron emission tomography (PET) scan and computed tomography (CT) scan. In some exemplary embodiments, the analysis controller 110 may determine molecular subtype profiles 180 of multiple cancer cells depicted in the positron emission tomography (PET) scan 210 and optionally the computed tomography (CT) scan 220 based on the first origin cells 240a of the first lesion contained in the first volume 230a and the second origin cells 240b of the second lesion contained in the second volume 230b. For example, optionally, the analysis controller 110 may determine that the molecular subtype profiles 180 of multiple cancer cells include a total origin cell set, based at least on the first origin cells 240a of the first lesion and the second origin cells 240b of the second lesion. The overall origin of multiple cancer cells may, in some cases, be identified as a specific origin cell (e.g., germinal center B cells (GCBs) or activated B cells (ABCs) (or non-germinal center B cells (non-GCBs)) if the probability of a threshold amount (e.g., percentage, ratio, etc.) of the lesion associated with that particular origin cell satisfies one or more thresholds. In some cases, the analysis engine 110 may determine the overall origin of multiple cancer cells by applying at least a machine learning model (e.g., a neural network such as a recurrent neural network (RNN)) that operates on an embedding sequence containing the probability that the lesion indicates the corresponding origin cell for each possible origin cell and each lesion. Alternatively and / or additionally, the analysis controller 110 may determine one or more metrics indicating heterogeneity and / or homogeneity between the first origin cell 240a of the first lesion and the second origin cell 240b of the second lesion for inclusion in the molecular subtype profile 180.

[0091] Figure 4B shows a flowchart illustrating another example of process 450 for machine learning-enabled origin cell prediction, according to several exemplary embodiments. Referring to Figure 4B, process 450 may be performed by the analysis controller 110 to generate molecular subtype profiles 180 of multiple cancer cells depicted by positron emission tomography (PET) scans and, optionally, corresponding computed tomography (CT) scans from the same time point. Optionally, process 400 may perform operation 304 of process 300.

[0092] In some cases, the analysis controller 110 may apply the origin cell classification model 113 to positron emission tomography (PET) scans 210 and, optionally, computed tomography (CT) scans 220 as part of performing process 450 to determine the origin cells of each lesion depicted therein. As the performance metrics shown in Figures 5-6 indicate, the machine learning-based origin cell classification model 113 applied by the analysis controller 110 when performing process 450 can accurately distinguish between lesions having different origin cells (e.g., germinal center B cells (GCBs) and activated B cells (ABCs) (or non-germinal center B cells (non-GCBs) in diffuse large B-cell lymphoma (DLBCL))). The ability of the origin cell classification model 113 to determine origin cells non-invasively and based on medical images (e.g., positron emission tomography (PET) scans and, optionally, corresponding computed tomography (CT) scans) acquired as part of routine patient care is a way to perform accurate and precise origin cell determination non-invasively and with minimal additional resources. This means that process 450 can be performed. In contrast, conventional techniques for determining the origin cell (COO) require invasive assays and data analysis to extract tumor tissue, which are accompanied by either limited availability or performance, limited reproducibility, and high inter-leader variability. Furthermore, while conventional techniques rely on data associated with a single tumor sample, the analysis controller 110 can perform process 450 to generate molecular subtype profiles 180, which can describe different origin cells that may be present in different lesions. Thus, the molecular subtype profiles 180 generated by the analysis controller 110 performing process 450 can capture valuable biological insights, such as heterogeneity and / or homogeneity of origin cells across different lesions, which escape conventional techniques.

[0093] In 452, the analysis controller 110 may identify one or more lesions within positron emission tomography (PET) scans and computed tomography (CT) scans depicting multiple cancerous cells. As noted, in some exemplary embodiments, the analysis controller 110 may segment the positron emission tomography (PET) scan 210 and optionally the computed tomography (CT) scan 220 to identify one or more lesions depicted therein. In some cases, the analysis controller 110 may identify one or more lesions by applying a segmentation model 112 (e.g., an artificial neural network) or machine learning model (e.g., a logistic regression model, a tree-based classifier, a fully connected neural network) trained to classify one or more objects identified in the positron emission tomography (PET) scan 210 and / or computed tomography (CT) scan 220 (e.g., via thresholding of pixel values).

[0094] In 454, the analysis controller 110 may extract multiple features of each of one or more lesions from positron emission tomography (PET) scans and computed tomography (CT) scans. The origin cells of the lesion may be determined based on one or more features present in the positron emission tomography (PET) scan 210 and, optionally, in the computed tomography (CT) scan 220 depicting the lesion. Thus, in some exemplary embodiments, the analysis controller 110 may extract first multiple features 235a of a first lesion and second multiple features 235b of a second lesion from the positron emission tomography (PET) scan 210 and, optionally, in the computed tomography (CT) scan 220. The first multiple features 235a and second multiple features 235b may include, for example, various radiomic features including the size, shape, and / or texture of the corresponding lesion. Other examples of radiomic features include one or more primary statistics such as range, maximum, minimum, median, mode, and / or mean pixel value of one or more pixels depicting a lesion in a positron emission tomography (PET) scan 210 and optionally a computed tomography (CT) scan 220. In some cases, the first and second sets of features 235a and 235b may further include other radiomic features such as gray level co-occurrence matrices, gray level size zone matrices, and / or gray level run-length matrices of one or more pixels depicting the corresponding lesion in the positron emission tomography (PET) scan 210 and optionally a computed tomography (CT) scan 220.

[0095] In 456, the analysis controller 110 may apply an origin cell classification model 113 to determine the origin cells of a corresponding lesion based on several features extracted from positron emission tomography (PET) scans and computed tomography (CT) scans. For example, in some cases, the origin cell classification model 113 may be applied to determine the first origin cells 240a of a first lesion based on at least a first set of features 235a. Furthermore, the origin cell classification model 113 may be applied to determine the second origin cells 240b of a second lesion based on at least a second set of features 235b. As noted, the origin cell classification model 113 may include one or more of the following: artificial neural networks (ANNs) (e.g., visual transformer models such as a visual transformer model with shifted patch tokenization and local autoattention), tree-based classifiers (e.g., gradient boosted decision trees, random forests, extreme gradient boosted decision trees (XGBoost)), ridge classifiers, etc. In the example of process 450, the origin cell classification model 113 can be trained to learn the nexus between the origin cells of a lesion and the various radiomic features of the lesion present in positron emission tomography (PET) and / or computed tomography (CT) scans that depict the lesion.

[0096] In 458, the analysis controller 110 may determine molecular subtype profiles of multiple cancer cells based on the origin cells of each lesion present in at least the positron emission tomography (PET) scan and computed tomography (CT) scan. In some exemplary embodiments, the analysis controller 110 may determine an overall origin cell determined based at least on the first origin cell 240a of the first lesion contained in the first volume 230a and the second origin cell 240b of the second lesion contained in the second volume 230b, in order to include in the molecular subtype profiles 180 of multiple cancer cells depicted in the positron emission tomography (PET) scan 210 and optionally the computed tomography (CT) scan 220. As noted, optionally, the overall origin cell may include, for each possible origin cell, a corresponding probability determined based on the probability of each individual lesion having that origin cell (e.g., the maximum, minimum, mean, median, and / or mode of the first probability of the first lesion and the second probability of the second lesion having the origin cell). In some cases, the analysis engine 110 may determine the overall origin cells by applying at least one machine learning model (e.g., a neural network such as a recurrent neural network (RNN)) that operates on an embedding sequence containing the corresponding probability that the lesion represents the origin cells for each possible origin cell and each lesion. Alternatively and / or additionally, molecular subtype profiles 180 of multiple cancer cells may be generated to include one or more metrics indicating heterogeneity and / or homogeneity of origin cells present across different lesions (e.g., between the first origin cells 240a of the first lesion and the second origin cells 240b of the second lesion).

[0097] Table 1 below illustrates various performance metrics (e.g., sensitivity, specificity, and area under the curve (AUC)) for distinguishing between germinal center B cell (GCB) molecular subtypes and activated B cell (ABC) molecular subtypes across various clinical datasets (e.g., GOYA holdout, Cavalli, Gather) 113 origin cell classification models (e.g., gradient boosted decision trees (GB), random forests (RF), extreme gradient boosted decision trees (XGBoost), and visual transducer models with shifted patch tokenization and localized autoattention (SPT / LSA ViT)).

[0098] [Table 1]

[0099] Table 2 below illustrates various performance metrics (e.g., sensitivity, specificity, and area under the curve (AUC)) for distinguishing between germinal center B cell (GCB) molecular subtypes and non-germinal center B cell (non-GCB) molecular subtypes for various embodiments of origin cell classification models113 (e.g., gradient-boosted decision trees (GB), random forests (RF), extreme gradient-boosted decision trees (XGBoost), and visual transformer models with shifted patch tokenization and localized autoattention (SPT / LSA ViT)) across various clinical datasets (e.g., GOYA holdout, Cavalli, Gather).

[0100] [Table 2]

[0101] Figure 5A plots the accuracy of different embodiments of the origin cell classification model 113 (e.g., SPT / LSA ViT, extreme gradient boosted decision tree (XG Boost), random forest (RF), and left-to-right gradient boosted decision tree (GB)) across different clinical datasets (e.g., GOYA holdout, Cavalli, and Gather). Figure 5B shows the F1 scores for each embodiment of the origin cell classification model 113 (e.g., SPT / LSA ViT, extreme gradient boosted decision tree (XG Boost), random forest (RF), and left-to-right gradient boosted decision tree (GB)) for activated B cell (ABC) molecular subtypes across different clinical datasets (e.g., GOYA holdout, Cavalli, and Gather). Figure 5C shows the F1 scores for each embodiment of the origin cell classification model 113 (e.g., SPT / LSA ViT, extreme gradient boosted decision tree (XG Boost), random forest (RF), and left-to-right gradient boosted decision tree (GB)) for germinal center B cell (GCB) molecular subtypes across different clinical datasets (e.g., GOYA holdout, Cavalli, and Gather). Figure 5D plots the area under the curve (AUC) achieved by different embodiments of the origin cell classification model 113 (e.g., SPT / LSA ViT, extreme gradient boosted decision tree (XG Boost), random forest (RF), and left-to-right gradient boosted decision tree (GB)) across different clinical datasets (e.g., GOYA holdout, Cavalli, and Gather).

[0102] Figure 6A shows a graph illustrating the receiver operating characteristic (ROC) curves of the origin cell classification model 113 when distinguishing between activated B cell (ABC) molecular subtypes and germinal center B cell (GCB) molecular subtypes across different test sets. Figure 6B shows a graph illustrating the receiver operating characteristic (ROC) curves of the origin cell classification model 113 when distinguishing between germinal center B cell (GCB) molecular subtypes and non-germinal center B cell (non-GCB) molecular subtypes across different test sets.

[0103] Figure 7A shows a graph illustrating the importance of various radiomic features in the origin cell classification model 113 implemented using gradient-boosted decision trees. Figure 7B shows a graph illustrating the importance of various radiomic features in the origin cell classification model 113 implemented using random forests (RF). Figure 7C shows a graph illustrating the importance of various radiomic features for the origin cell classification model 113 implemented using extreme gradient-boosted decision trees (XGBoost).

[0104] Table 3 below illustrates various performance metrics (e.g., sensitivity, specificity, and area under the curve (AUC)) for different embodiments of origin cell classification models (e.g., visual transformer models with ridge classifiers and fully connected convolutional neural networks (FCCC)) for distinguishing between germinal center B cell (GCB) molecular subtypes and activated B cell (ABC) molecular subtypes across various clinical datasets (e.g., GOYA holdout, Cavalli, Gather).

[0105] [Table 3]

[0106] Taking into consideration the above-described embodiments of the subject matter, the present application discloses the following list of examples, which are further examples included in the disclosure of the present application, by combining one feature of a single example or two or more features of the aforementioned examples, or optionally, one or more features of one or more further examples.

[0107] Item 1: A computer implementation method comprising: receiving a first positron emission tomography (PET) scan depicting multiple cancer cells; identifying a first lesion depicted in the first PET scan; applying an origin cell classification model to determine a first origin cell associated with the first lesion based on at least the first lesion depicted in the first PET scan; and determining the molecular subtype profiles of multiple cancer cells depicted in the first PET scan based on at least the first origin cell of the first lesion.

[0108] Item 2: The method according to Item 1, wherein the first origin cell of the first lesion includes, for each possible origin cell, the probability that one or more cancerous cells forming the first lesion originate from that origin cell.

[0109] Item 3: The method of Item 1 or Item 2, further comprising: identifying a second lesion depicted in a first PET scan; applying an origin cell classification model to determine a second origin cell associated with the second lesion, based at least on the second lesion depicted in the first PET scan; and determining molecular subtype profiles of multiple cancer cells depicted in the first PET scan, further based on the second origin cells of the second lesion.

[0110] Item 4: The method of Item 3, wherein the molecular subtype profiles of multiple cancer cells depicted in the first PET scan include, for each possible origin cell, the probability that the overall origin cell of the multiple cancer cells is that origin cell.

[0111] Item 5: The method according to Item 4, wherein the probability that the overall origin of multiple cancerous cells is a specific origin is the maximum, minimum, mean, median, and / or mode of the respective probabilities that each of the first and second lesions has that specific origin.

[0112] Item 6: The method of Item 3 or 4, wherein the molecular subtype profiles of multiple cancerous cells depicted in the first PET scan include, for each possible origin cell, the corresponding percentage of lesions having that origin cell.

[0113] Item 7: The method according to any one of items 3 to 5, wherein the molecular subtype profiles of multiple cancer cells depicted in a first PET scan are determined by generating an embedding sequence that includes at least the first origin cells of the first lesion and the second origin cells of the second lesion, and by applying a machine learning model to determine the overall origin cells of the multiple cancer cells in the first PET scan, based at least on the embedding sequence.

[0114] Item 8: The method described in Item 7, wherein the machine learning model is a recurrent neural network.

[0115] Item 9: The method of any one of items 1 to 8, further comprising extracting a volume containing the primary lesion identified in the primary PET scan from the primary PET scan, and applying an origin cell classification model to determine the primary origin cells associated with the primary lesion, based on the volume extracted from at least the primary PET scan.

[0116] Item 10: The method of Item 9, wherein the volume containing the first lesion is extracted by determining the centroid of the first lesion within at least the first PET scan, and extracting the volume based at least on the centroid of the first lesion.

[0117] Item 11: The method of Item 10, wherein the volume is a three-dimensional volume containing multiple two-dimensional patches centered on the centroid of the primary lesion.

[0118] Item 12: The method according to Item 11, wherein the multiple two-dimensional patches include multiple axial patches or multiple coronal patches.

[0119] Item 13: The method according to any one of items 1 to 12, wherein the first lesion is identified by at least applying a segmentation model to identify multiple pixels corresponding to the first lesion within the first PET scan.

[0120] Item 14: The method according to any one of items 1 to 13, further comprising: extracting from a first PET scan several features associated with a first lesion identified in the first PET scan; and applying an origin cell classification model to determine the first origin cells associated with the first lesion based on at least the several features extracted from the first PET scan.

[0121] Item 15: The method of Item 14, wherein multiple features include the size of the primary lesion, the shape of the primary lesion, and / or the texture of the primary lesion.

[0122] Item 16: The method of Item 14 or 15, wherein multiple features include one or more primary statistics associated with one or more pixels that depict a primary lesion in a primary PET scan.

[0123] Item 17: The method according to any one of items 14 to 16, wherein multiple features include a gray-level co-occurrence matrix, a gray-level size zone matrix, and / or a gray-level run-length matrix of one or more pixels depicting a first lesion in a first PET scan.

[0124] Item 18: The method of any one of items 1 to 17, further comprising receiving a computed tomography (CT) scan from the same time point as the first PET scan, identifying the first lesion depicted in the CT scan, and applying an origin cell classification model to determine the first origin cell associated with the first lesion based on the first PET scan and the first lesion depicted in the CT scan.

[0125] Item 19: The method according to Item 18, wherein each pixel included in a CT scan is associated with an intensity value corresponding to the density of tissue or X-ray attenuation.

[0126] Item 20: The method according to Item 18, wherein each pixel in the first PET scan is associated with an intensity value corresponding to the level of metabolic activity.

[0127] Item 21: The method of Item 18, further comprising determining a tumor mask corresponding to a first lesion based on at least one of a CT scan and a first PET scan, and further applying an origin cell classification model to determine the origin cells of the first lesion, based at least on the tumor mask.

[0128] Item 22: The method according to Item 21, further comprising determining an organ mask corresponding to one or more organs depicted in a CT scan and a first PET scan, based on at least one of a CT scan and a first PET scan, and further applying an origin cell classification model to determine the origin cells of a first lesion, based at least on the organ masks.

[0129] Item 23: The method according to any one of items 1 to 22, further comprising receiving a second positron emission tomography (PET) scan from a different time point than the first PET scan, identifying the first lesion depicted in the second PET scan, and applying an origin cell classification model to determine the first origin cell of the first lesion based on the first and second PET scans of the first lesion.

[0130] Item 24: The method of any one of items 1 to 23, further comprising: receiving a first computed tomography (CT) scan from the same time point as the first PET scan and a second CT scan from the same time point as the second PET scan; identifying the first lesion depicted in the first CT scan and the first lesion depicted in the second CT scan; and applying an origin cell classification model to determine the first origin cell of the first lesion based on at least the first PET scan, the first CT scan, the second PET scan, and the second CT scan of the first lesion.

[0131] Item 25: The method according to any one of Items 1 to 24, further comprising determining the disease diagnosis, prognosis, disease progression, treatment, and / or treatment response of a patient associated with a first PET scan, based on molecular subtype profiles of multiple cancer cells depicted in at least one first PET scan.

[0132] Item 26: The method described in any one of items 1 through 25, wherein multiple cancerous cells are associated with diffuse large B-cell lymphoma (DLBCL).

[0133] Item 27: The method described in any one of items 1 through 26, wherein the origin cell classification model is trained to distinguish between multiple origin cells.

[0134] Item 28: The method according to Item 27, wherein the multiple origin cells include germinal center B cells (GCBs) and activated B cells (ABCs).

[0135] Item 29: The method according to Item 27, wherein the multiple origin cells include germinal center B cells (GCBs) or non-germinal center B cells (non-GCBs).

[0136] Item 30: The method described in any one of items 1 through 29, wherein the origin cell classification model includes an artificial neural network (ANN).

[0137] Item 31: The method according to any one of items 1 to 30, wherein the origin cell classification model includes a visual transducer.

[0138] Item 32: The method according to any one of Items 1 to 31, wherein the origin cell classification model includes a visual converter with shifted patch tokenization and localized autoattention.

[0139] Item 33: The method according to any one of items 1 through 32, wherein the origin cell classification model includes a tree-based classifier.

[0140] Item 34: The method described in any one of items 1 through 33, wherein the origin cell classification model includes a ridge classifier.

[0141] Item 35: The method according to any one of items 1 to 34, further comprising training an origin cell classification model based on at least a training set to determine the origin cells of at least one lesion depicted in a positron emission tomography (PET) scan based on at least a portion of the PET scan.

[0142] Item 36: The method of Item 35, further comprising: extracting a first volume from a second positron emission tomography (PET) scan containing the second lesion depicted in the second PET scan; generating a first training sample containing the first volume and a first ground truth annotation of the second origin cells of the second lesion; and generating a training set containing the first training sample.

[0143] Item 37: The method of Item 36, further comprising generating a second volume based on at least a first volume, and generating a second training sample containing the second volume and a first ground truth annotation of the second origin cells of the second lesion for inclusion in a training set.

[0144] Item 38: The method described in Item 37, which produces the second volume by modifying the first volume.

[0145] Item 39: Modifications made in any way described in Item 37 or 38, including one or more of the following: normalization, rotation, inversion, and zoom modification of the primary volume.

[0146] Item 40: The method according to any one of items 37 to 39, wherein the modification of the first volume includes modifying one or more slices of the first volume that are within a threshold distance of the centroid of the second lesion.

[0147] Item 40: A system comprising at least one data processor and at least one memory for storing instructions that, when executed by the at least one data processor, result in an operation including the method described in any one of Items 1 to 40.

[0148] Item 41: A non-temporary computer-readable medium that stores instructions that, when executed by at least one data processor, result in an operation including the method described in any one of Items 1 through 40.

[0149] Figure 8 is a block diagram illustrating an example of a computing system 800 in the implementation form of this subject. Referring to Figures 1 to 8, the computing system 800 can be used to implement an analysis controller 110, one or more imaging devices 120, a client device 130, and / or any of its components.

[0150] As shown in Figure 8, the computing system 800 may include a processor 810, memory 820, storage device 830, and input / output device 840. The processor 810, memory 820, storage device 830, and input / output device 840 can be interconnected via a system bus 850. The processor 810 is capable of processing instructions for execution within the computing system 800. Such instructions for execution may implement, for example, one or more components of an analysis controller 110, one or more imaging devices 120, and a client device 130. In some exemplary embodiments, the processor 810 may be a single-threaded processor. Alternatively, the processor 810 may be a multi-threaded processor. The processor 810 is capable of processing instructions stored in memory 820 and / or storage device 830 to display graphical information for a user interface provided via the input / output device 840.

[0151] Memory 820 is a computer-readable medium, such as volatile or non-volatile, that stores information within the computing system 800. Memory 820 can store, for example, data structures representing a configuration object database. Storage device 830 can provide persistent storage for the computing system 800. Storage device 830 may be a solid-state drive, floppy disk device, hard disk device, optical disk device, or tape device, or other suitable persistent storage means. Input / output device 840 provides input / output operations for the computing system 800. In some exemplary embodiments, input / output device 840 includes a keyboard and / or pointing device. In various embodiments, input / output device 840 includes a display device for displaying a graphical user interface.

[0152] According to some exemplary embodiments, the input / output device 840 can provide input / output operations for network devices. For example, the input / output device 840 may include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., local area networks (LANs), wide area networks (WANs), the Internet).

[0153] In some exemplary embodiments, the computing system 800 can be used to run various interactive computer software applications that can be used for organizing, analyzing, and / or storing various forms of data. Alternatively, the computing system 800 can be used to run any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., generating, managing, and editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, and so on. The application may include various add-in functions or may be a standalone computing product and / or function. When activated within the application, functionality can be used to generate a user interface provided via the input / output device 840. The user interface is generated by the computing system 800 and can be presented to the user (e.g., on a computer screen monitor).

[0154] One or more aspects or features of the subject matter described herein may be realized in digital electronic circuits, integrated circuits, specially designed ASICs, field-programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementations in one or more computer programs executable and / or interpretable on a programmable system which includes at least one programmable processor, which may be for special or general purposes, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device. A programmable system or computing system may include a client and a server. The client and server are generally located geographically separated from each other and typically interact through a communication network. The client-server relationship arises from computer programs running on each computer having a client-server relationship with each other.

[0155] These computer programs, sometimes called programs, software, software applications, applications, components, or code, contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, as well as / or assembly / machine languages. As used herein, the term “machine-readable medium” means any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, including, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), and includes machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signals” means any signals used to provide machine instructions and / or data to a programmable processor. Machine-readable medium can store such machine instructions non-temporarily, for example, non-temporarily, solid-state memory, magnetic hard drives, or any equivalent storage medium. Machine-readable medium can, alternatively or additionally, store such machine instructions temporarily, for example, processor caches or other random-access memories associated with one or more physical processor cores.

[0156] To provide user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having, for example, a display device such as a cathode ray tube (CRT), liquid crystal display (LCD), or light-emitting diode (LED) monitor for displaying information to the user, and a keyboard, and a pointing device such as a mouse or trackball by which the user can provide input to the computer. User interaction may also be provided using other types of devices. For example, repetition provided to the user may be any form of sensory repetition, such as visual repetition, auditory repetition, or tactile repetition, and input from the user may be received in any form, including acoustic input, voice input, and tactile input. Other possible input devices include touchscreens, or other touch-sensing devices such as single-point or multi-point resistive or capacitive trackpads, speech recognition hardware and software, optical scanners, optical pointers, digital image acquisition devices, and associated interpretation software.

[0157] In the above specification and claims, phrases such as “at least one of ~” or “one or more of ~” may precede a list of consecutive elements or features. The term “and / or” may also be used in an enumeration of two or more elements or features. Unless implicitly or explicitly contradicted by the context in which it is used, such phrases are intended to mean any of the enumerated elements or features individually, or any of the enumerated elements or features in combination with any of the other enumerated elements or features. For example, the phrases “at least one of A and B,” “one or more of A and B,” and “A and / or B” are intended to mean “A only,” “B only,” or “A and B together,” respectively. A similar interpretation is intended for enumerations containing three or more items. For example, the phrases “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, and / or C” are intended to mean, respectively, “A only, B only, C only, A and B together, A and C together, B and C together, or A, B, and C together.” The use of the term “based on” in the above and claims is intended to mean “at least partially based on,” thereby allowing features or elements that are not enumerated.

[0158] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles, depending on the desired configuration. The embodiments described above do not necessarily represent all embodiments of the subject matter described herein. Rather, they are merely some examples that correspond to aspects related to the described subject matter. While some variations have been detailed above, other modifications and additions are possible. In particular, further features and / or variations can be provided in addition to those described herein. For example, the embodiments described above can be directed to various combinations and subcombinations of the disclosed features, and / or combinations and subcombinations of some of the further features disclosed above. In addition, the logical flows depicted in the accompanying drawings and / or described herein do not necessarily require a specific order or sequence shown to achieve the desired result.

[0159] Other embodiments may fall within the scope of the following claims.

Claims

1. A computer implementation method, Receiving a first positron emission tomography (PET) scan depicting multiple cancer cells, Identifying the first lesion depicted by the first PET scan, Applying an origin cell classification model to determine the first origin cell associated with the first lesion, based at least on the first lesion depicted in the first PET scan, and A computer implementation method comprising determining the molecular subtype profiles of the plurality of cancer cells depicted in the first PET scan, based at least on the first origin cells of the first lesion.

2. The method according to claim 1, wherein the first origin cell of the first lesion includes, for each possible origin cell, the probability that one or more cancerous cells forming the first lesion originate from that origin cell.

3. Identifying the second lesion depicted by the first PET scan, Applying the origin cell classification model to determine the second origin cell associated with the second lesion, based at least on the second lesion depicted in the first PET scan, and The method according to claim 1 or 2, further comprising determining the molecular subtype profiles of the plurality of cancerous cells depicted in the first PET scan, based on the second origin cells of the second lesion.

4. The method according to claim 3, wherein the molecular subtype profiles of the plurality of cancer cells depicted by the first PET scan include, for each possible origin cell, the probability that the overall origin cell of the plurality of cancer cells is derived from that origin cell.

5. The method according to claim 4, wherein the probability that the overall origin of the plurality of cancerous cells is a specific origin is the maximum, minimum, mean, median, and / or mode of the probability that each of the first lesion and the second lesion has that specific origin.

6. The method according to claim 3 or 4, wherein the molecular subtype profile of the plurality of cancer cells depicted by the first PET scan includes, for each possible origin cell, the corresponding proportion of lesions having that origin cell.

7. The molecular subtype profiles of the plurality of cancer cells depicted in the first PET scan are at least, To generate an embedding sequence that includes the first origin cells of the first lesion and the second origin cells of the second lesion, and The method according to any one of claims 3 to 5, wherein the determination is made by applying a machine learning model to determine the overall origin of the plurality of cancerous cells in the first PET scan, based at least on the implantation sequence.

8. The method according to claim 7, wherein the machine learning model is a recurrent neural network.

9. Extracting a volume from the first PET scan that includes the first lesion identified in the first PET scan, and The method according to any one of claims 1 to 8, further comprising applying the origin cell classification model to determine the first origin cells associated with the first lesion based on the volume extracted from at least the first PET scan.

10. The volume containing the first lesion is at least, Within the first PET scan, the center of gravity of the first lesion is determined, and The method according to claim 9, wherein the volume is extracted based on the centroid of at least the first lesion.

11. The method according to claim 10, wherein the volume is a three-dimensional volume comprising a plurality of two-dimensional patches centered on the centroid of the first lesion.

12. The method according to claim 11, wherein the plurality of two-dimensional patches include a plurality of axial patches or a plurality of coronal patches.

13. The method according to any one of claims 1 to 12, wherein the first lesion is identified by at least applying a segmentation model to identify a plurality of pixels corresponding to the first lesion within the first PET scan.

14. Extracting from the first PET scan a plurality of features associated with the first lesion identified in the first PET scan, and The method according to any one of claims 1 to 13, further comprising applying the origin cell classification model to determine the first origin cells associated with the first lesion based on the plurality of features extracted from at least the first PET scan.

15. The method according to claim 14, wherein the plurality of features include the size of the first lesion, the shape of the first lesion, and / or the texture of the first lesion.

16. The method according to claim 14 or 15, wherein the plurality of features include one or more primary statistics associated with one or more pixels depicting the first lesion in the first PET scan.

17. The method according to any one of claims 14 to 16, wherein the plurality of features include a gray-level co-occurrence matrix, a gray-level size zone matrix, and / or a gray-level run-length matrix of one or more pixels depicting the first lesion in the first PET scan.

18. Receiving a computed tomography (CT) scan from the same point in time as the first PET scan, identifying the first lesion depicted in the CT scan, and The method according to any one of claims 1 to 17, further comprising applying the origin cell classification model to determine the first origin cells associated with the first lesion based on the first PET scan and the first CT scan.

19. The method according to claim 18, wherein each pixel included in the CT scan is associated with an intensity value corresponding to tissue density or X-ray attenuation.

20. The method according to claim 18, wherein each pixel in the first PET is associated with an intensity value corresponding to a level of metabolic activity.

21. Determining a tumor mask corresponding to the first lesion based on at least one of the CT scan and the first PET scan, and The method of claim 18, further comprising applying the origin cell classification model to determine the origin cells of the first lesion, based at least on the tumor mask.

22. Based on at least one of the CT scan and the first PET scan, determine an organ mask corresponding to one or more organs depicted in the CT scan and the first PET scan, and The method according to claim 21, further comprising applying the origin cell classification model to determine the origin cells of the first lesion, based at least on the organ mask.

23. Receiving a second positron emission tomography (PET) scan from a different point in time than the first PET scan, Identifying the first lesion depicted by the second PET scan, and The method according to any one of claims 1 to 22, further comprising applying the origin cell classification model to determine the first origin cells of the first lesion based on the first lesion depicted in the first PET scan and the second PET scan.

24. To receive a first computed tomography (CT) scan from the same time as the first PET scan, and a second CT scan from the same time as the second PET scan, Identifying the first lesion depicted in the first CT scan and the first lesion depicted in the second CT scan, and The method according to any one of claims 1 to 23, further comprising applying the origin cell classification model to determine the first origin cells of the first lesion based on at least the first PET scan, the first CT scan, the second PET scan, and the first lesion depicted in the second CT scan.

25. The method according to any one of claims 1 to 24, further comprising determining the disease diagnosis, prognosis, disease progression, treatment, and / or treatment response of a patient associated with the first PET scan, based on the molecular subtype profiles of the plurality of cancer cells depicted in at least the first PET scan.

26. The method according to any one of claims 1 to 25, wherein the plurality of cancerous cells are associated with diffuse large B-cell lymphoma (DLBCL).

27. The method according to any one of claims 1 to 26, wherein the origin cell classification model is trained to distinguish between multiple origin cells.

28. The method according to claim 27, wherein the plurality of origin cells include germinal center B cells (GCBs) and activated B cells (ABCs).

29. The method according to claim 27, wherein the plurality of origin cells include germinal center B cells (GCBs) or non-germinal center B cells (non-GCBs).

30. The method according to any one of claims 1 to 29, wherein the origin cell classification model includes an artificial neural network (ANN).

31. The method according to any one of claims 1 to 30, wherein the origin cell classification model includes a visual transducer.

32. The method according to any one of claims 1 to 31, wherein the origin cell classification model includes a visual transducer with shifted patch tokenization and localized autoattention.

33. The method according to any one of claims 1 to 32, wherein the origin cell classification model includes a tree-based classifier.

34. The method according to any one of claims 1 to 33, wherein the origin cell classification model includes a ridge classifier.

35. The method according to any one of claims 1 to 34, further comprising training the origin cell classification model based on at least a training set to determine the origin cells of at least one lesion depicted in the positron emission tomography (PET) scan based on at least a portion of the PET scan.

36. Extracting a first volume from a second positron emission tomography (PET) scan, which includes the second lesion depicted in the second PET scan. To generate a first training sample that includes the first volume and the first ground truth annotation of the second origin cells of the second lesion, and The method of claim 35, further comprising generating the training set to include the first training sample.

37. To generate a second volume based at least on the first volume, and The method of claim 36, further comprising generating a second training sample comprising a second volume and the first ground truth annotation of the second origin cells of the second lesion, for inclusion in the training set.

38. The method of claim 37, wherein the second volume is generated by modifying the first volume.

39. The method according to claim 37 or 38, wherein the modification includes one or more of the normalization, rotation, inversion, and zoom modification of the first volume.

40. The method according to any one of claims 37 to 39, wherein the modification of the first volume includes modifying one or more slices of the first volume that are within a threshold distance of the centroid of the second lesion.

41. It is a system, At least one data processor, A memory for storing instructions that, when executed by the at least one data processor, result in an operation including the method according to any one of claims 1 to 40, A system equipped with these features.

42. A non-transient computer-readable medium storing instructions that, when executed by at least one data processor, result in an operation including the method according to any one of claims 1 to 40.