Multicellular representations of single-cell transcriptomics for machine learning enabled patient-level disease analytics

The embedding computation model aggregates cell embeddings from subsets of cells to generate patient-level representations, addressing data heterogeneity challenges and enabling comprehensive disease analytics across multiple diseases.

WO2026090285A1PCT designated stage Publication Date: 2026-04-30GENENTECH INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Conventional single-cell transcriptome analytics are limited to specific cell types or indications, disregarding the overall cellular ecosystem and struggling to model multiple diseases due to data heterogeneity, batch effects, and imbalanced composition across different tissue types and diseases.

Method used

An embedding computation model generates patient-level multicellular representations by aggregating cell embeddings from subsets of cells using attention-based mechanisms, trained on large-scale single-cell expression studies to enable downstream tasks like disease analytics, classification, and clustering.

Benefits of technology

Provides biologically informed multicellular representations at the patient level, enabling fine-grained gene and cell-type prioritization, disease severity inference, and holistic disease mechanism interrogation across various diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025052061_30042026_PF_FP_ABST
    Figure US2025052061_30042026_PF_FP_ABST
Patent Text Reader

Abstract

A training dataset may be generated from a set of sample gene expression profiles associated with different tissue types, cell types, and / or diseases. An embedding computation model may be trained based on the training dataset. A gene expression profile of a patient may include a gene expression level of a plurality of genes expressed by different cells. The embedding computation model may be applied to generate a multicellular representation of one or more subsets of cells from the gene expression profile of the patient. The embedding computation model may generate the multicellular representation by aggregating the cell embeddings generated by embedding the gene expression level of each cell in the subset of cells. One or more downstream tasks, such as dimensionality reduction, biological feature prioritization, treatment response prediction, disease severity classification, and patient subgroup discovery, may be performed based on the multicellular representation of the patient.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1MULTICELLULAR REPRESENTATIONS OF SINGLE-CELL TRANSCRIPTOMICS FOR MACHINE LEARNING ENABLED PATIENT-LEVEL DISEASE ANALYTICS TECHNICAL FIELD

[0001] This application claims priority to U. S. Provisional Application No.63 / 711011, entitled “SINGLE-CELL TRANSCRIPTOMICS BASED MODELING OF DISEASES” and filed on October 23, 2024, the disclosure of which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] The subject matter described herein relates generally to artificial intelligence and more specifically to multicellular representations of single-cell transcriptomics for machine learning enabled patient-level disease analytics.INTRODUCTION

[0003] Sequencing refers to various methodologies for determining the sequence of nucleotide bases in one or more polynucleotides including, for example, nucleic acid molecules such as deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and variants or derivatives thereof (e.g., single stranded DNA). When performed in bulk, sequencing may determine the nucleic acid sequence of polynucleotides extracted from a population of cells but without any differentiation between the different individual cells or cell types. Contrastingly, single-cell sequencing leverages next-generation sequencing (NGS) technologies to examine nucleic acid sequence information from individual cells. The resulting single-cell genome or transcriptome profile may provide high resolution insights into the characteristics and functions of individual cells in the context of its microenvironment.1NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1SUMMARY

[0004] Systems, methods, and articles of manufacture, including computer program products, are provided for generating multicellular representations of single-cell transcriptomics for machine learning enabled patient-level disease analytics. For example, in some cases, an embedding computation model may generate, based at least on the gene expression profile of a patient, a patient-level multicellular representation for use with various downstream tasks for machine learning enabled disease analytics. Examples of downstream tasks for machine learning enabled disease analytics include dimensionality reduction and visualization, biological feature (e.g., cell, gene, and / or the like) prioritization, treatment response and disease severity prediction, patient subgroup discovery, and / or the like. In some cases, the patient-level multicellular representation may be an embedding summarizing the patient’s cellular context. In some cases, the embedding computation model may generate the patient-level multicellular representation from one or more subsets of cells present in the patient’s gene expression profile, with each subset of cells including some but not all of the cells present in the patient’s gene expression profile. In some cases, each subset of cells include cells of a single type of cell (or cell type) or multiple cell types present in the patient’s gene expression profile. In some cases, each subset of cells may be represented as a gene expression matrix in which each row corresponds to an individual cell in the subset of cells, each column corresponds to a specific gene, and each value corresponds to the gene expression level of a gene in a cell. In some cases, the embedding computation model may generate the patient-level multicellular representation by at least embedding the gene expression levels in the cell expression matrix of one or more subsets of cells before the resulting cell embeddings are aggregated into a single vector, for example, by weighting the cell embeddings with a cell-level attention-based aggregation2NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1mechanism. Moreover, in some cases, a disease inference model may perform, based at least on the patient-level multicellular representation of the patient, one or more downstream tasks for machine learning enabled disease analytics. In some cases, the one or more downstream tasks for machine learning enabled disease analytics may include one or more of comparing, clustering, or classifying the patient-level multicellular representation. In some cases, the embedding computation model and the disease inference model may be trained in an end-to-end fashion with annotated training samples, such as gene expression profiles with ground-truth labels for disease classification. Alternatively, the embedding computation model may be trained in a self-supervised manner to reduce (or minimize) a reconstruction loss and a contrastive loss associated with the patient-level multicellular representations generated by the embedding computation model.

[0005] In one aspect, there is provided a system for generating multicellular representations of single-cell transcriptomics for machine learning enabled patient-level disease analytics. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that result in operations when executed by the at least one data processor. The operations may include: generating, based at least on one or more single-cell transcriptome datasets, a training dataset, wherein the one or more single-cell transcriptome datasets include sample gene expression profiles associated with a plurality of different tissue types, cell types, and / or diseases; training, based at least on the training dataset, an embedding computation model; receiving a gene expression profile of a patient, wherein the gene expression profile includes a gene expression level of a plurality of genes expressed by each cell of a plurality of cells; determining a subset of cells from the gene expression profile of the patient; applying the embedding computation model to generate, based at least on the subset3NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1of cells from the gene expression profile of the patient, a multicellular representation for the patient, wherein the embedding computation model generates the multicellular representation by at least aggregating a plurality of cell embeddings generated by embedding the gene expression level of each cell included in the subset of cells from the gene expression profile of the patient; and performing, based at least on the multicellular representation of the patient, one or more downstream tasks.

[0006] In another aspect, there is provided a computer-implemented method for generating multicellular representations of single-cell transcriptomics for machine learning enabled patient-level disease analytics. The method may include: generating, based at least on one or more single-cell transcriptome datasets, a training dataset, wherein the one or more singlecell transcriptome datasets include sample gene expression profiles associated with a plurality of different tissue types, cell types, and / or diseases; training, based at least on the training dataset, an embedding computation model; receiving a gene expression profile of a patient, wherein the gene expression profile includes a gene expression level of a plurality of genes expressed by each cell of a plurality of cells; determining a subset of cells from the gene expression profile of the patient; applying the embedding computation model to generate, based at least on the subset of cells from the gene expression profile of the patient, a multicellular representation for the patient, wherein the embedding computation model generates the multicellular representation by at least aggregating a plurality of cell embeddings generated by embedding the gene expression level of each cell included in the subset of cells from the gene expression profile of the patient; and performing, based at least on the multicellular representation of the patient, one or more downstream tasks.4NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0007] In another aspect, there is provided a computer program product for generating multicellular representations of single-cell transcriptomics for machine learning enabled patient-level disease analytics. The computer program product may include a non-transitory computer readable medium storing instructions that result in operations when executed by at least one data processor. The operations may include: generating, based at least on one or more single-cell transcriptome datasets, a training dataset, wherein the one or more single-cell transcriptome datasets include sample gene expression profiles associated with a plurality of different tissue types, cell types, and / or diseases; training, based at least on the training dataset, an embedding computation model; receiving a gene expression profile of a patient, wherein the gene expression profile includes a gene expression level of a plurality of genes expressed by each cell of a plurality of cells; determining a subset of cells from the gene expression profile of the patient; applying the embedding computation model to generate, based at least on the subset of cells from the gene expression profile of the patient, a multicellular representation for the patient, wherein the embedding computation model generates the multicellular representation by at least aggregating a plurality of cell embeddings generated by embedding the gene expression level of each cell included in the subset of cells from the gene expression profile of the patient; and performing, based at least on the multicellular representation of the patient, one or more downstream tasks.

[0008] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0009] In some variations, the subset of cells from the gene expression profile is represented as a gene expression matrix.5NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0010] In some variations, the gene expression matrix includes one or more rows corresponding to one or more individual cells in the subset of cells from the gene expression profile, one or more columns corresponding to one or more genes, and one or more values corresponding to one or more gene expression levels.

[0011] In some variations, the embedding computation model generates each cell embedding of the plurality of cell embeddings by at least embedding the one or more gene expression levels of a corresponding cell in the subset of cells.

[0012] In some variations, the embedding computation model generates the multicellular representation of the patient by applying a cell-level attention-based aggregator to aggregate the plurality of cell embeddings.

[0013] In some variations, the cell-level attention-based aggregator aggregates the plurality of cell embeddings by at least determining a plurality of weights, and aggregating the plurality of cell embeddings by at least applying the plurality of weights to determine a weighted sum of the plurality of cell embeddings.

[0014] In some variations, the cell-level attention-based aggregator includes one or more of a mean aggregator, a linear aggregator, a non-linear aggregator, or a transformer aggregator.

[0015] In some variations, the cell-level attention-based aggregator includes a softmax-attention pooling layer.

[0016] In some variations, the cell-level attention-based aggregator is applied to aggregate the plurality of cell embeddings into a single vector corresponding to the multicellular representation of the patient.6NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0017] In some variations, the embedding computation model is trained in an end-to-end fashion with a disease inference model applied to perform the one or more downstream tasks.

[0018] In some variations, an end-to-end training of the embedding computation model and the disease inference model includes adjusting one or more parameters of the embedding computation model and the disease inference model to reduce a loss function quantifying a discrepancy in an output of the disease inference model.

[0019] In some variations, an end-to-end training of the embedding computation model and the disease inference model includes generating the training dataset to include a plurality of annotated training samples, wherein each annotated training sample of the plurality of annotated training samples includes a subset of cells from a sample gene expression profile and a corresponding ground-truth label, applying the embedding computation model to determine, for each annotated training sample included in the training dataset, a corresponding sample multicellular representation, applying the disease inference model to perform, based at least on the corresponding sample multicellular representation, the one or more downstream tasks, and adjusting one or more parameters of the embedding computation model and the disease inference model to reduce a loss function quantifying a discrepancy between an output of the disease inference model and the corresponding ground-truth label of each annotated training sample in the training dataset.

[0020] In some variations, the embedding computation model is trained in a self-supervised manner using a plurality of unannotated training samples.

[0021] In some variations, the training the embedding computation model in the self-supervised manner includes generating the training dataset to include a plurality of unannotated7NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1training samples, wherein each unannotated training sample of the plurality of unannotated training samples includes a subset of cells from a sample gene expression profile without a corresponding ground-truth label, applying the embedding computation model to determine, for each unannotated training sample included in the training dataset, a corresponding sample multicellular representation, and adjusting one or more parameters of the embedding computation model to reduce a reconstruction loss and a contrastive loss associated with the corresponding sample multicellular representation of each annotated training sample included in the training dataset.

[0022] In some variations, the reconstruction loss of a sample multicellular representation quantifies a similarity between a plurality of cell embeddings generated for the sample gene expression profile and a reconstruction of the plurality of cell embeddings recovered from the sample multicellular representation.

[0023] In some variations, the contrastive loss of a sample multicellular representation of one subset of cells from a sample gene expression profile quantifies a similarity between the sample multicellular representation and a sample multicellular representation of a different subset of cells from a same sample gene expression profiles.

[0024] In some variations, the training dataset is generated to include an oversampling of one or more underrepresented tissue types, cell types, and / or diseases.

[0025] In some variations, wherein the oversampling includes determining a subset of cells from a sample gene expression profile to include a different proportion of the one or more underrepresented tissue types, cell types, and / or diseases than an original proportion of the one or more underrepresented tissue types, cell types, and / or diseases present in the sample gene expression profile, and generating, for inclusion in the training dataset, a training sample8NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1including the subset of genes having an oversampling of the one or more underrepresented tissue types, cell types, and / or diseases.

[0026] In some variations, the oversampling includes generating, for inclusion in the training dataset, a different proportion of training samples associated with the one or more underrepresented tissue types, cell types, and / or diseases than an original proportion of the one or more underrepresented tissue types, cell types, and / or diseases present in the one or more singlecell transcriptome datasets.

[0027] In some variations, a disease inference model is applied to perform the one or more downstream tasks. A plurality of gradient attributions are generated by at least determining, for each cell-gene combination present in the gene-expression profile of the patient, a gradient attribution corresponding to a rate of change in an output of the disease inference model performing the one or more downstream tasks.

[0028] In some variations, the plurality of gradient attributions are averaged across the plurality of cells present in the gene expression profile to determine, for each cell of the plurality of cells, an importance metric indicative of a contribution of the cell to the output of the disease inference model.

[0029] In some variations, the plurality of gradient attributions are averaged across the plurality of genes present in the gene expression profile to determine, for each gene of the plurality of genes, an importance metric indicative of a contribution of the gene to the output of the disease inference model.

[0030] In some variations, the plurality of gradient attributions are averaged across a plurality of cell types present in the gene expression profile to determine, for each cell types of9NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1the plurality of cell types, an importance metric indicative of a contribution of the cell type to the output of the disease inference model.

[0031] In some variations, the plurality of gradient attributions are averaged across a plurality of combinations of genes and cell types present in the gene expression profile to determine, for each gene-cell type combination of the plurality of combinations of genes and cell types, an importance metric indicative of a contribution of the gene-cell type combination to the output of the disease inference model.

[0032] In some variations, the disease inference model is trained to perform a disease classification task.

[0033] In some variations, the disease inference model is adapted to determine disease severity by at least identifying, based at least on the plurality of gradient attributions, one or more prioritized biological features, and determining, based at least on a proportion of the one or more prioritized biological features, disease severity.

[0034] In some variations, the one or more prioritized biological features include one or more biological features identified as having a threshold contribution to an output of the disease inference model performing the disease classification task.

[0035] In some variations, the one or more prioritized biological features include one or more cells, cell types, genes, or genes expressed by certain cell types.

[0036] In some variations, the one or more downstream tasks are performed by comparing, clustering, and / or classifying the multicellular representation of the patient.

[0037] In some variations, the one or more downstream tasks include one or more of dimensionality reduction and visualization, biological feature prioritization, treatment response prediction, disease severity prediction, or patient subgroup identification.10NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0038] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computing systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.

[0039] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes in relation to single-cell sequencing, notably single-cell ribonucleic acid (RNA) sequence (scRNA-seq), and the sequencing data derived therefrom it11NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.DESCRIPTION OF DRAWINGS

[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,

[0041] FIG. 1 depicts a system diagram illustrating an example of a transcriptome analysis system, in accordance with some example embodiments; and

[0042] FIG. 2A depicts a block diagram illustrating an example of an embedding computation model trained in an end-to-end fashion with a disease inference model, in accordance with some example embodiments;

[0043] FIG. 2B depicts a block diagram illustrating an example of an embedding computation model trained in a self-supervised manner to reduce (or minimize) reconstruction loss and contrastive loss, in accordance with some example embodiments;

[0044] FIG. 3 depicts a flowchart illustrating an example of a process for machine learning enabled patient-level disease analytics with patient-level multicellular representations of single-cell transcriptome data, in accordance with some example embodiments;

[0045] FIG. 4A depicts a flowchart illustrating an example of a process for training an embedding computation model, in accordance with some example embodiments;

[0046] FIG. 4B depicts a flowchart illustrating another example of a process for training an embedding computation model, in accordance with some example embodiments;12NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0047] FIG. 5A depicts a schematic diagram illustrating the dataflow in an example of a transcriptome analysis platform, in accordance with some example embodiments;

[0048] FIG. 5B depicts a schematic diagram illustrating the architecture of an example of a transcriptome analysis platform in which an embedding computation model and a disease inference model are trained in an end-to-end manner, in accordance with some example embodiments;

[0049] FIG. 5C depicts a schematic diagram illustrating the architecture of an example of a transcriptome analysis system in which an embedding computation model and a disease inference model are trained in a self-supervised manner, in accordance with some example embodiments;

[0050] FIG. 5D depicts a schematic diagram illustrating the interpretability of multicellular representations generated by an example of an embedding computation model, in accordance with some example embodiments;

[0051] FIG. 6A depicts a schematic diagram illustrating the heterogeneity of an example of a training dataset, in accordance with some example embodiments;

[0052] FIG. 6B depicts a bar graph illustrating the statistics of disease states in an example of a training dataset, in accordance with some example embodiments;

[0053] FIG. 6C depicts a bar graph illustrating the statistics of tissue types in an example of a training dataset, in accordance with some example embodiments;

[0054] FIG. 7A depicts a candlestick chart illustrating the weighted Fl -score results of an example of the embedding computation model of the present disclosure and various baseline models, in accordance with some example embodiments;13NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0055] FIG. 7B depicts a candlestick chart illustrating the weighted Fl -score results of various example embodiments of the embedding computation model of the present disclosure, in accordance with some example embodiments;

[0056] FIG. 8A depicts a uniform manifold approximation and projection (UMAP) visualization of patient-level multicellular representations organized by disease, in accordance with some example embodiments;

[0057] FIG. 8B depicts a uniform manifold approximation and projection (UMAP) visualization of patient-level multicellular representations organized by tissue type, in accordance with some example embodiments;

[0058] FIG. 9(a) depicts a candlestick chart illustrating the prioritization of cell types with attributions aggregated by cell types and ranked by mean attribution, in accordance with some example embodiments;

[0059] FIG. 9(b) depicts a candlestick chart illustrating the prioritization of cell type specific genes with attributions aggregated over classical monocytes and genes ranked by mean attribution, in accordance with some example embodiments;

[0060] FIG. 9(c) depicts a candlestick chart illustrating the prioritization of cell type specific genes with attributions aggregated over platelets and genes ranked by mean attribution, in accordance with some example embodiments;

[0061] FIG. 10A depicts a uniform manifold approximation and projection (UMAP) visualization of patient-level multicellular representations organized by disease severity and pseudobulk gene expression levels organized by disease severity, in accordance with some example embodiments;14NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0062] FIG. 10B depicts a candlestick chart illustrating the magnitude of integrated gradient attributions averaged across all myeloid dendritic cells for each patient gene expression profile, grouped by disease severity, in accordance with some example embodiments;

[0063] FIG. 10C depicts a candlestick chart illustrating the predicted probability of COVID-19 diagnosis stratified by disease severity, in accordance with some example embodiments;

[0064] FIG. 11 depicts a candlestick chart illustrating the benchmarking results for binary classification between an example of the embedding computation model described herein and various baseline models for classifying COVID 19 versus healthy condition, in accordance with some example embodiments;

[0065] FIG. 12A depicts a candlestick chart illustrating the performance of an example of the embedding computation model described herein under different learning rates for binary classification, in accordance with some example embodiments;

[0066] FIG. 12B depicts a candlestick chart illustrating the performance of an example of the embedding computation model described herein under different dropout rates for binary classification, in accordance with some example embodiments;

[0067] FIG. 12C depicts a candlestick chart illustrating the performance of an example of the embedding computation model described herein under different weight decay for binary classification, in accordance with some example embodiments;

[0068] FIG. 12D depicts a candlestick chart illustrating the performance of an example of the embedding computation model described herein under different quantities of sampled cells for binary classification, in accordance with some example embodiments;15NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0069] FIG. 12E depicts a candlestick chart illustrating the performance of an example of the embedding computation model described herein under the cumulative setting of different disease states, in accordance with some example embodiments;

[0070] FIG. 12F depicts a candlestick chart illustrating the performance of an example of the embedding computation model described herein classifying healthy gene expression profiles and other gene expression profiles with different diseases, in accordance with some example embodiments;

[0071] FIG. 12G depicts a candlestick chart illustrating the performance of an example of the embedding computation model described herein under different width neural network layers, in accordance with some example embodiments;

[0072] FIG. 12H depicts a candlestick chart illustrating the performance of an example of the embedding computation model described herein under different depths of neural network layers, in accordance with some example embodiments;

[0073] FIG. 13 depicts a heatmap illustrating the distribution of correlation coefficients of patient-level multicellular representations with different diseases across tissue types, in accordance with some example embodiments;

[0074] FIG. 14A depicts a uniform manifold approximation and projection (UMAP) visualization of patient-level multicellular representations organized by disease, in accordance with some example embodiments;

[0075] FIG. 14B depicts a uniform manifold approximation and projection (UMAP) visualization of patient-level multicellular representations organized by tissue type, in accordance with some example embodiments;16NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0076] FIG. 14C depicts a uniform manifold approximation and projection (UMAP) visualization of patient-level multicellular representations organized by disease and tissue type pairs, in accordance with some example embodiments;

[0077] FIG. 15A depicts correlation matrices of patient-level multicellular representations computed based on different rules for disease similarity across diseases, in accordance with some example embodiments;

[0078] FIG. 15B depicts a scatter plot illustrating the relationship between the correlation computed with patient-level multicellular representations generated by an example of the embedding computation model described herein and the correlation computed with text embeddings for describing the similarity of lung cancer in lung tissue and other disease-tissue pairs, in accordance with some example embodiments;

[0079] FIG. 15C depicts a scatter plot illustrating the relationship between the correlation computed with patient-level multicellular representations generated by an example of the embedding computation model described herein and the correlation computed with text embeddings for describing the similarity of COVID-19 in lung tissue and other disease-tissue pairs, in accordance with some example embodiments;

[0080] FIG. 16 depicts a bubble chart illustrating the Gene Ontology Enrichment Analysis (GOEA) results of representative immune cells based on prioritized genes, in accordance with some example embodiments;

[0081] FIG. 17 depicts candlestick charts illustrating the results of Wilcoxon rank sum tests for the association between attribution scores and CO VID 19 severity for each cell type, in accordance with some example embodiments;17NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0082] FIG. 18A depicts a schematic diagram illustrating an overview of an example of a treatment prediction task, in accordance with some example embodiments;

[0083] FIG. 18B depicts a uniform manifold approximation and projection (UMAP) visualization of patient-level multicellular representations organized by treatment response, in accordance with some example embodiments;

[0084] FIG. 18C depicts a bar graph illustrating the benchmarking results between using patient-level multicellular representations and pseudobulk gene expression levels as inputs for the treatment prediction task, in accordance with some example embodiments;

[0085] FIG. 18D depicts a bar graph illustrating a comparison of the performance for various disease analytics tasks using a random approach, a pseudo-bulk approach, an embedding computation model trained in an end-to-end manner, and an embedding computation model trained in a self-supervised manner, in accordance with some example embodiments; and

[0086] FIG. 19 depicts a block diagram illustrating an example of a computing system, in accordance with some example embodiments.

[0087] When practical, similar reference numbers denote similar structures, features, or elements.DETAILED DESCRIPTION

[0088] Unlike bulk sequencing, which examines nucleic acid sequences or gene expression levels at a cell population level without any differentiation between individual cells or cell types, single-cell sequencing technologies are capable of providing this information at the resolution of individual cells. For example, single cell deoxyribonucleic acid (DNA) sequencing may include isolating a single cell and amplifying at least a portion of its genome for sequencing. In oncological applications, genetic mutations giving rise to cancerous cells may be detected by18NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1sequencing the deoxyribonucleic acid (DNA) molecules of individual cells. Contrastingly, single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) may include isolating and amplifying the RNA molecules of individual cells, each of which being tagged with a DNA barcode to enable differentiating between RNA molecules originating from different cells during subsequent sequencing. In some cases, transcriptome data from single-cell RNA sequencing (scRNA-seq) may include gene expression levels (or counts) derived by measuring the abundance (or relative abundance) of the RNA molecules.

[0089] Advances in sequencing technologies, including the aforementioned singlecell sequencing techniques, have given rise to a vast and exponentially growing collection of data for biological research. For example, by profiling the expression level of genes across hundreds of millions of individual cells, single-cell RNA sequencing (scRNA-seq) has yielded insights into the heterogeneity of cell states and functions. Understanding disease processes at the patient level with the granularity of single-cell transcriptomics may uncover new cell types, genes programs associated with response to therapy or drug resistance, specific marker genes, unique patient subsets, and / or the like. However, conventional single-cell transcriptome analytics first partitions cells into categories (e.g., types, subtypes, states, and / or the like) before each category is separately analyzed, thus disregarding the overall cellular ecosystem and the impact that the homeostasis disruptions caused by diseases has across a multitude of cells.Existing state-of-the-art machine learning models capable of identifying disease states across broader cellular ecosystems have been limited to specific indications, such as COVID-19.Despite the growing number of single-cell RNA sequencing (scRNA-seq) studies yielding single-cell transcriptome data from a sufficiently large patient population to realistically support machine learning approaches capable of modeling disease biology at a patient level, efforts to19NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1develop a single machine learning model capable of leveraging the full scope of available data and jointly model multiple diseases are thwarted by a number of significant challenges arising from the heterogeneity of single-cell RNA sequencing (scRNA-seq) data, including the confounding and batch effects of data from different studies, an imbalanced composition of data from different tissue types, cell types, and diseases, and the noise present in the data.

[0090] Various example embodiments of the present disclosure overcome the limitations of conventional single-cell transcriptome analytics, including those leveraging machine learning, by training an embedding computation model on a gene expression profile from different tissue types and diseases. In some cases, the embedding computation model may be trained to generate multicellular representations at the individual patient level for use in one or more downstream tasks for machine learning enabled disease analytics. For example, in some cases, the embedding computation model may generate a patient-level multicellular representation from one or more subsets of cells present in a patient’s gene expression profile. In some cases, each subset of cells may include some but not all of the cells present in the patient’s gene expression profile. Moreover, in some cases, each subset of cells in this context may include cells of a single type of cells (or cell type) or multiple cell types present in the patient’s gene expression profile. In some cases, each subset of cells may be represented as a gene expression matrix in which each row corresponds to an individual cell in that subset of cells, each column corresponds to a specific gene, and each value corresponds to an expression level of a gene in a cell. In some cases, the embedding computation model may generate the patientlevel multicellular representation by at least embedding the gene expression levels of the individual cells in each subset of cells to generate cell embeddings. In some cases, the cell embeddings may be aggregated, for example, into a single vector, by weighting with a cell-level20NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1atention-based aggregation mechanism. As described in more detail below, the patient-level multicellular representation of the patient may be used for one or more downstream tasks for machine learning enabled disease analytics. For instance, in some cases, the one or more downstream tasks for machine learning enabled disease analytics may include one or more of comparing, clustering, or classifying the patient-level multicellular representation. Unlike conventional single-cell transcriptome analytics, which are limited to specific cell types or indications, various example embodiments of the embedding computation model described herein leverage large scale single-cell expression studies across different tissue types and diseases to provide biologically informed multicellular representations at the patient level.

[0091] In some example embodiments, a disease inference model may perform, based at least on one or more patient-level multicellular representations generated by the embedding computation model, one or more downstream tasks for machine learning enabled disease analytics. For example, in some cases, the one or more downstream tasks for machine learning enabled disease analytics may include one or more of comparing, clustering, or classifying the one or more patient-level multicellular representations. In some cases, the one or more patientlevel multicellular representations may be leveraged to enable fine-grained gene or cell-type prioritization for different diseases as well as support biological discovery, including the holistic interrogation of disease mechanisms at the patient level, for single or multiple genes and cell types. As described in more detail below, the output of the disease inference model, such as disease, severity, subtype, and / or subgroup classifications, may be interpreted based on the underlying patient-level multicellular representations. For instance, in some cases, the disease inference model may be applied to the patient-level multicellular representation of the gene21NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1expression profile of a patient exhibiting a disease (e.g., COVID- 19) to infer the patient’s disease severity subgroup as well as prioritize cell type specific genes associated with the disease.

[0092] In some example embodiments, the output of the disease inference model, such as disease, severity, subtype, and / or subgroup classification for a patient, may be interpreted based at least on the patient-level multicellular representation used by the disease inference model. For example, in some cases, an importance metric may be determined for individual cells, sets of cells, cell types, genes, or sets of genes to quantify the corresponding contribution to the output of the disease inference model. In some cases, the output of the disease inference model may be interpreted based on integrated gradients (IG). In some cases, integrated gradients may be applied to generate, for each cell-gene combination in the gene expression profile of the patient, a corresponding gradient attribution. In some cases, doing so may generate a matrix of gradient attributions, which can then be averaged along different dimensions to support different levels of interpretability. For instance, averaging gradient attributions over genes may generate the importance metric of each individual cell whereas whether averaging gradient attributions over cells may yield importance metrics of individual genes. In some cases, gradient attributions may also be averaged over sets of cells, cell types, and / or genes to generate importance metrics at the corresponding resolution.

[0093] In some example embodiments, the disease inference model may be trained to perform a classification task, which may include assigning one or more labels to the multicellular representation of a patient. For example, in some cases, the disease inference model may be trained to perform a binary classification task that includes assigning, to the multicellular representation of the patient, a label identifying a patient as healthy or exhibiting a particular disease (e.g., COVID). In some cases, although the disease inference model is trained to perform22NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1the classification task, the disease inference model may nevertheless be adapted to also determine the severity of the disease present in the patient. In some cases, one or more prioritized biological features, such as cells, cell types, genes, and / or genes expressed by certain cell types, identified (e.g., based on integrated gradients (IGs)) as exhibiting a threshold contribution to the output of the disease inference model performing the classification task (e.g., binary classification task) may be leveraged to determine the severity of the disease. In some cases, the proportion of the one or more prioritized biological features (e.g., cells, cell types, genes, genes expressed by certain cell types, and / or the like) present in the subset of cells used by the embedding computation model to generate the multicellular representation of the patient may correspond to the severity of the disease. For instance, in some cases, the proportion of the one or more prioritized biological features (e.g., cells, cell types, gene, genes expressed by certain cell types, and / or the like) may be directly proportional to the severity of the corresponding disease such that the greater the proportion of the one or more prioritized biological features, the more severe the corresponding disease.

[0094] In some example embodiments, the embedding computation model and the disease inference model may be trained in an end-to-end fashion with annotated training samples, such as gene expression profiles with ground-truth labels (e.g., for disease classification, biological feature prioritization, patient subgroup discovery, and / or the like). Accordingly, in some cases, the training of the embedding computation model may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model along with one or more parameters (e.g., weights, biases, and / or the like) of the disease inference model to reduce (or minimize) a loss function quantifying a discrepancy between an output of the disease inference model and the ground-truth label of the corresponding23NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1annotated training sample. Alternatively, the embedding computation model may be trained in a self-supervised manner (e.g., without unannotated training samples) to reduce (or minimize) a loss function quantifying a reconstruction loss and a contrastive loss associated with the patientlevel multicellular representations generated by the embedding computation model. As described in more detail below, in some cases, the reconstruction loss of a patient-level multicellular representation generated from a patient’s gene expression profile may quantify a discrepancy between the cell embeddings generated for at least a portion of original gene expression profile and the cell embeddings reconstructed from the patient-level multicellular representation. Furthermore, in some cases, contrastive loss may quantify the similarity between the patient-level multicellular representations of similar gene expression profiles (or portions thereof). In some cases, reducing (or minimizing) contrastive loss may include increasing (or maximizing) the similarity between the patient-level multicellular representations of similar gene expression profiles (or subsets of cells from the same gene expression profile) and decreasing (or minimizing) the similarity between patient-level multicellular representations of dissimilar gene expression profiles (or subsets of cells from different gene expression profiles).

[0095] In some example embodiments, the embedding computation model may be trained on subsets of cells from gene expression profiles with different tissue types, cell types, diseases, and / or the like. However, data heterogeneity, including an imbalanced composition of data from different tissue types, cell types, and diseases, may be thwart efforts to integrate data from large scale single-cell expression studies across different tissue types and diseases when training the embedding computation model. Accordingly, various example embodiments of the present disclosure mitigate the challenges of integrating heterogeneous data, including single-cell transcriptome data with an imbalanced composition of different tissue types, cell types, diseases,24NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1and / or the like. For example, in some cases, the training set for training the embedding computation model may be generated to include a different proportion of cells and / or subsets of cells of underrepresented tissue types, cell types, and / or diseases than what is present in the available single-cell transcriptome datasets.

[0096] FIG. 1 depicts a system diagram illustrating an example of a transcriptome analysis system 100, in accordance with some example embodiments. In some example embodiments, the transcriptome analysis system 100 may leverage single-cell transcriptome data from different tissue types, cell types, and diseases to generate patient-level multicellular representations for various machine learning enabled disease analytics, such as dimensionality reduction and visualization, biological feature (e.g., cell, gene, and / or the like) prioritization, treatment response and disease severity prediction, patient subgroup discovery, and / or the like.

[0097] Referring again to FIG. 1, the transcriptome analysis system 100 may include a transcriptome analysis platform 110, a sequencing platform 120, and a client device 130. As shown in FIG. 1, in some cases, the transcriptome analysis platform 110, the sequencing system 120, and the client device 130 may be communicatively coupled via a network 140. In some cases, the sequencing platform 120 may include one or more library preparation kits, polymerase chain reaction (PCR) machines, centrifuges, reagents and chemicals (e.g., nucleic acid polymerase, nucleotides, primers), sequencers, flow cells, microarrays, computational tools (e.g., sequence alignment, differential gene expression analysis), and / or the like. In some cases, the client device 130 may be a processor-based device, including, for example, a smartphone, a tablet computer, a wearable apparatus, a virtual assistant, an Intemet-of-Things (loT) appliance, and / or the like. In some cases, the network 140 may be a wired network and / or a wireless network, including, for example, a wide area network (WAN), a local area network (LAN), a25NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1virtual local area network (VLAN), a public land mobile network (PLMN), the Internet, and / or the like. In some cases, the transcriptome analysis platform 110, the sequencing platform 120, and / or the client device 130 may be contained within and / or operate on the same platform and / or device. For example, in some cases, the client device 130 may form a part of, include, and / or be coupled to one or both of the transcriptome analysis platform 110 and the sequencing platform 120.

[0098] In some example embodiments, the transcriptome analysis platform 110 may include an embedding computation model 112 and a disease inference model 114. In some cases, the embedding computation model 112 may be trained to generate, from a gene expression profile 111 of a patient, a multicellular representation 113. For example, in some cases, the gene expression profile 111 of the patient may include one or more subsets of cells, each of which including some but not all of the cells present in the gene expression profile 111 of the patient. In some cases, a single subset of cells may include cells of a single type of cell (or cell type) or multiple cell types present in the gene expression profile 111. In some cases, each subset of cells from the gene expression profile 111 may be represented as a gene expression matrix in which each row corresponds to a cell the corresponding subset of cells, each column corresponds to a gene, and each value corresponds an expression level of a gene in a cell.

[0099] In some cases, the embedding computation model 112 may generate, based at least on the gene expression matrix of one or more subsets of cells from the gene expression profile 111 of the patient, the multicellular representation 113. For example, in some cases, the embedding computation model 112 may generate the multicellular representation 113 by at least embedding the gene expression levels of the individual cells in each subset of cells before the resulting cell embeddings are aggregated into a single vector corresponding to the multicellular26NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1representation 113. In some cases, the embedding computation model 112 may aggregate the cell embeddings by at least applying a cell-level attention-based aggregation mechanism such that the cell embeddings are aggregated by weighting the individual cell embeddings.

[0100] Referring again to FIG. 1, in some cases, the disease inference model 114 may operate on the multicellular representation 113 to perform one or more downstream tasks for machine learning enabled disease analytics. Examples of the one or more downstream tasks for machine learning enabled disease analytics may include dimensionality reduction and visualization, biological feature (e.g., cell, gene, and / or the like) prioritization, treatment response and disease severity prediction, patient subgroup discovery, and / or the like. For example, in some cases, the disease inference model 114 may perform the one or more downstream tasks for machine learning enabled disease analytics by at least comparing, clustering, and / or classifying the multicellular representation 113. In the example shown in FIG.1, in some cases, the disease inference model 114 may generate, for example, for display at a user interface 135 of the client device 130, a predictive output 115 corresponding to a result of comparing, clustering, and / or classifying the multicellular representation. For instance, in some case, the predictive output 115 may include a classification of a disease severity subgroup for the patient associated with the gene expression profile 111 as well as a prioritization of cell type specific genes associated with a disease exhibited by the patient.

[0101] In some example embodiments, the embedding computation model 112 and the disease inference model 114 may be trained in an end-to-end fashion using annotated training samples, such as gene expression profiles with ground-truth labels (e.g., for disease classification, biological feature prioritization, patient subgroup discovery, and / or the like). In some cases, training the embedding computation model 112 and the disease inference model 11427NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1in the end-to-end fashion may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model 112 and the disease inference model 114 while reducing (or minimizing) a loss function quantifying a discrepancy between an output of the disease inference model 114 and the ground-truth label of the corresponding annotated training sample.

[0102] To further illustrate, FIG. 2A depicts a block diagram illustrating an example of the embedding computation model 112 trained in an end-to-end fashion with the disease inference model 114, in accordance with some example embodiments. As shown in FIG. 2 A, the transcriptome analysis platform 110 may receive an annotated training sample 201 including a subset of cells 203 from a sample gene expression profile 204 and a corresponding ground truth label 205. In some cases, the embedding computation model 112 may generate a sample multicellular representation 207 by at least embedding the gene expression levels of individual cells included in each set of cells included in the subset of cells 203 from the sample gene expression profile 204 before aggregating the resulting cell embeddings into a single vector corresponding to the sample multicellular representation 207. In some cases, the disease inference model 114 may determine, based at least on the sample multicellular representation 207, a sample predictive output 209 including, for example, one or more of a classification for disease, severity, subtype, subgroup, and / or the like. In some cases, the end-to-end training of the embedding computation model 112 and the disease inference model 114 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model 112 and the disease inference model 114 to reduce (or minimize) a loss function quantifying a discrepancy between the sample predictive output 209 of the disease28NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1inference model 114 and the ground truth label 205 associated with the subset of cells 203 from the sample gene expression profile 204.

[0103] In some example embodiments, instead of training the embedding computation model 112 and the disease inference model 114 in an end-to-end fashion as shown in FIG. 2A, the embedding computation model 112 may be trained in a self-supervised manner to obviate the need to annotated training samples, such as the annotated training sample 201 shown in FIG. 2A as including the subset of cells 203 from the sample gene expression profile 204 and the ground truth label 205 (e.g., of a corresponding classification for disease, severity, subtype, subgroup, and / or the like). For example, in some cases, the training of the embedding computation model 112 may include reducing (or minimizing) a loss function quantifying the reconstruction loss and contrastive loss associated with the patient-level multicellular representations generated by the embedding computation model 112.

[0104] In some example embodiments, instead of being trained end-to-end with the disease inference 114, the embedding computation model 112 may be trained in a self-supervised manner. To further illustrate, FIG. 2B depicts a block diagram illustrating an example of the embedding computation model 112 trained in a self-supervised manner to reduce (or minimize) reconstruction loss and contrastive loss, in accordance with some example embodiments. As shown in FIG. 2B, the transcriptome analysis platform 110 may receive an unannotated training sample 211, which may include a first subset of cells 213a from a sample gene expression profile 214 without any ground-truth labels (e g., of a corresponding classification for disease, severity, subtype, subgroup, and / or the like). In some cases, the embedding computation model 112 may generate a sample multicellular representation 215 by at least embedding the gene expression levels of individual cells included in in the first subset of cells 213a from the sample gene29NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1expression profile 214 before aggregating the resulting cell embeddings into a single vector corresponding to the sample multicellular representation 215. Furthermore, in some cases, the disease inference model 114 may determine, based at least on the sample multicellular representation 215, a sample predictive output 217 including, for example, one or more of a classification for disease, severity, subtype, subgroup, and / or the like.

[0105] Referring again to FIG. 2B, in some cases, the training of the embedding computation model 112 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model 112 to reduce (or minimize) a loss function quantifying a contrastive loss and a reconstruction loss. For example, as shown in FIG. 2B, the contrastive loss of the embedding computation model 112 may include a similarity (e.g., cosine distance) between the sample multicellular representation 215 and a sample multicellular representation 219 of one or more similar subsets of cells. In some cases, the first subset of cells 213a may be considered similar to a second subset of cells 213b from the same sample gene expression profile 214. In some cases, the reduction (or minimization) of the contrastive loss may train the embedding computation model 212 to increase (or maximize) the similarity between the patient-level multicellular representations of similar subsets of cells, such as those originating from the same gene expression profiles, and decrease (or minimize) the similarity between patient-level multicellular representations of dissimilar subsets of cells, such as those originating from different gene expression profiles.

[0106] As noted, in some cases, the self-supervised training of the embedding computation model 112 may further include reducing a reconstruction loss of the embedding computation model 112. As shown in FIG. 2B, the reconstruction loss of the embedding computation model 112 may include a similarity between the cell embeddings 219 generated by30NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1the embedding computation model 112 for each cell included in the first subset of cells 213a from the sample gene expression profile 214 and reconstructed cell embeddings 221 generated, for example, by a decoder 216, decoding the sample multicellular representation 215.

[0107] FIG. 3 depicts a flowchart illustrating an example of a process 300 for machine learning enabled patient-level disease analytics with patient-level multicellular representations of single-cell transcriptome data, in accordance with some example embodiments. In some example embodiments, the process 300 may be performed by the transcriptome analysis platform 110 to train the embedding computation model 112 to generate the multicellular representation 113 for one or more subsets of cells from the gene expression profile 111 of a patient. As shown in FIG. 1, the disease inference model 114 may determine, based at least on the multicellular representation 113 of the one or more subsets of cells from the gene expression profile 111, the predictive output 115. In some cases, the predictive output 115 may be determined by one or more of comparing, clustering, or classifying the multicellular representation 113 of the one or more subsets of cells from the gene expression profile 111. As described in more detail below, the embedding computation model 112 may be trained in an end-to-end fashion with the disease inference model 114 using annotated training samples (e.g., the annotated training sample 201 shown in FIG. 2A) or, alternatively, in a self-supervised manner to reduce (or minimize) a loss function quantifying a contrastive loss and a reconstruction loss of the multicellular representation (e.g., the sample multicellular representation 215 shown in FIG.2B) generated by the embedding computation model 112.

[0108] At 302, a training dataset is generated, based at least on one or more singlecell transcriptome datasets, to include a set of sample gene expression profiles associated with a plurality of different tissue types, cell types, and / or disease. In some example embodiments the31NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1training dataset may be generated to include single-cell transcriptome data from across different studies. Accordingly, in some cases, the underlying single-cell transcriptome data may exhibit heterogeneity, including an imbalanced composition in the form of under- and / or overrepresentation of some tissue types, cell types, and / or diseases. Nevertheless, generating the training dataset to encompass heterogeneous data may improve the generalizability (e.g., across tissue types, cell types, diseases, and / or the like) of an embedding computation model trained therewith. Moreover, in some cases, the deleterious impact of an imbalanced composition of tissue types, cell types, and / or diseases on the performance of the embedding computation model may be mitigated by oversampling at least some underrepresented tissue types, cell types, and / or diseases when generating the training dataset.

[0109] At 304, an embedding computation model is trained based at least on the training dataset. In some example embodiments, a single gene expression profile may include the gene expression levels of cells from multiple different types of cells (or cell types). In some cases, one or more subsets of cells may be formed from the cells in the gene expression profile. In some cases, each subset of cells may include the gene expression levels of some but not all of the cells present in the gene expression profile. In some cases, each subset of cells may be represented as a gene expression matrix in which each row corresponds to an individual cell present in the subset of cells, each column corresponds to a specific gene, and each value corresponds to the gene expression level of a gene in a cell. In some cases, the embedding computation model may be trained to generate a patient-level multicellular representation by at least embedding the gene expression levels of the individual cells in one or more subsets of cells before the resulting cell embeddings are aggregated into a single vector corresponding to a multicellular representation of the gene expression profile. For instance, in some cases, the32NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1embedding computation model may be trained to aggregate the cell embeddings of different cells by weighting the cell embeddings with a cell-level attention-based aggregation mechanism.

[0110] As described in more detail below, the embedding computation model may be trained in an end-to-end fashion with a disease inference model that uses the multicellular representations generated by the embedding computation model to perform one or more downstream tasks for machine learning enabled disease analytics including, for example, disease classification, biological feature prioritization, patient subgroup discovery, and / or the like. In some cases, the embedding computation model and the disease inference model may be trained using annotated training samples to reduce (or minimize) a loss function quantifying a discrepancy between the output of the disease inference model and the ground-truth label of the corresponding annotated training sample. Alternatively, the embedding computation model may be trained in a self-supervised manner (e.g., without any annotated training samples) to reduce (or minimize) a loss function quantifying the reconstruction loss and contrastive loss associated with the multicellular representations generated by the embedding computation model.

[0111] At 306, a gene expression profile of a patient is received. In some example embodiments, the patient’s gene expression profile may be received from a sequencing platform. In some cases, the sequencing platform may perform single-cell sequencing, such as single-cell ribonucleic (RNA) sequencing (scRNA-seq) to measure the gene activity of individual cells. For example, in some cases, the sequencing platform may generate the gene expression profile of the patient by isolating individual cells within a biological sample (e.g., tissue sample, fluid sample, and / or the like) associated with the patient. In some cases, the ribonucleic acid (RNA) of the cells are converted to barcoded complementary deoxyribonucleic acid (cDNA), which then undergoes sequencing (e.g., next-generation sequencing) to determine the expression levels of33NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1different genes by individual cells. Accordingly, in some cases, the patient’s gene expression profile may include the gene expression levels of genes expressed by different types of cells (or cell types) present in the biological sample (e.g., tissue sample, fluid sample, and / or the like) associated with the patient. For instance, in some cases, the gene expression profile of the patient may include one or more subsets of cells, each of which including some but not all of the cells present in the gene expression profile. Moreover, in some cases, each subset of cells may be represented as a gene-expression matrix in which each row corresponds to an individual cell present in the subset of cells, each column corresponds to a specific gene, and each value corresponds to the gene expression level of a gene in a cell.

[0112] At 308, the trained embedding computation model is applied to generate a multicellular representation of the gene expression profile of the patient. In some example embodiments, the trained embedding computation model may generate the multicellular representation of the patient’s gene expression profile by at least embedding the gene expression levels of the cells in one or more subsets of cells to generate one or more corresponding cell embeddings. In some cases, to generate the multicellular representation of the patient’s gene expression profile, the trained embedding computation model may aggregate the one or more cell embeddings. For example, in some cases, the embedding computation model may aggregate the one or more cell embeddings by at least weighting the one or more cell embeddings with a celllevel attention-based aggregation mechanism.

[0113] At 310, one or more downstream or additional tasks for machine learning enabled disease analytics are performed based at least of the multicellular representation of the patient. In some example embodiments, a disease inference model may be trained to perform, based at least on the multicellular representation of the patient, the one or more downstream34NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1tasks for machine learning enabled disease analytics. For example, in some cases, the disease inference model may perform the one or more downstream tasks by comparing, clustering, and / or classifying the multicellular representation of the patient’s gene expression profile. In some cases, the one or more downstream tasks may include dimensionality reduction and visualization, biological feature (e.g., cell, gene, and / or the like) prioritization, treatment response and disease severity prediction, patient subgroup discovery, and / or the like. Moreover, in some cases, the output of the disease inference model may be interpreted, based on the multicellular representation of the patient’s gene expression profile, to indicate a contribution (e.g., importance) of one or more cells, genes, cell types, and / or groupings to the output of the disease inference model. For instance, in some cases, the output of the disease inference model may be interpreted by at least applying integrated gradients (IG) to generate, for each cell-gene combination in the gene expression profile of the patient, a corresponding gradient attribution. In some cases, the gradient attribution of a particular cell-gene combination may be determined by calculating the gradient (or rate of change) in output of the disease inference model attributable to the cell-gene combination, thus indicating how changes in that cell-gene combination impact the output of the disease inference model. In some cases, the resulting matrix of gradient attributions may be averaged along different dimensions to support different levels of interpretability. In some cases, averaging gradient attributions over genes may generate the importance metric of each individual cell whereas whether averaging gradient attributions over cells may yield importance metrics of individual genes. In some cases, gradient attributions may also be averaged over sets of cells, cell types, and / or genes to generate importance metrics at the corresponding resolution.35NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0114] In some example embodiments, the disease inference model may be trained to perform a classification task, which may include assigning one or more labels to the multicellular representation of a patient. For example, in some cases, the disease inference model may be trained to perform a binary classification task that includes assigning, to the multicellular representation of the patient, a label identifying a patient as healthy or exhibiting a particular disease (e.g., COVID). In some cases, although the disease inference model is trained to perform the classification task, the disease inference model may nevertheless be adapted to also determine the severity of the disease present in the patient. For instance, in some cases, one or more biological features prioritized by the disease inference model when performing the classification task may be leveraged toward determining the severity of the disease.

[0115] In some cases, the one or more prioritized biological features may include one or more cells, cell types, and genes present in the subset of cells from which the embedding computation model generates the multicellular representation of the patient. In some cases, the one or more prioritized biological features may be identified based on integrated gradients (IGs). In some cases, a biological feature identified through integrated gradients (IGs) as exhibiting a threshold contribution to the output of the disease inference model performing the classification task (e.g., binary classification task) may be a prioritized biological feature leveraged to determine the severity of the disease. For example, in some cases, the proportion of a prioritized biological feature (e.g., cell, cell type, gene, genes expressed by a certain cell type, and / or the like) present in the subset of cells used by the embedding computation model to generate the multicellular representation of the patient may correspond to the severity of the disease. In some cases, the proportion of the prioritized biological feature (e.g., cell, cell type, gene, genes expressed by a certain cell type, and / or the like) may be directly proportional to the severity of36NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1the corresponding disease such that the greater the proportion of the prioritized biological feature, the greater the severity the corresponding disease.

[0116] FIG. 4A depicts a flowchart illustrating an example of a process 400 for training an embedding computation model, in accordance with some example embodiments. In some example embodiments, the process 400 may be performed by the transcriptome analysis platform 110 to train the embedding computation model 112 and the disease inference model 114 in an end-to-end manner. In some cases, the process 400 may implement operation 302 of the process 300 shown in FIG. 3. As described in more detail below, in some cases, the end-to-end training of the embedding computation model 112 and the disease inference model 114 may use annotated training samples, such as the annotated training sample 201 shown in FIG. 2 A as including the sample gene expression profile 203 and the corresponding ground-truth label 205 (e.g., for disease, severity, subtype, and / or subgroup classification).

[0117] At 402, one or more annotated training samples are generated in which each annotated training sample includes a subset of cells from a sample gene expression profile and a corresponding ground-truth label. In some example embodiments, the subset of cells may include the gene expression levels of some but not all of the cells included in the sample gene expression profile. In some cases, the sample gene expression profiles may originate from large scale single-cell expression studies across different tissue types and disease. In some cases, sample gene expression profiles from different tissue types and diseases may exhibit certain heterogeneity, such as an imbalanced composition of different tissue types, cell types, diseases, and / or the like. In some cases, the challenges of integrating heterogeneous data may be mitigated by oversampling underrepresented tissue types, cell types, and / or diseases when generating each annotated training sample. In some cases, a tissue type, cell type, and / or disease37NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1may be underrepresented in the available single-cell transcriptome datasets if the constituent proportion gene expression levels of cells of that tissue type, cell type, and / or disease fail to satisfy one or more thresholds. In some cases, the underrepresented tissue type, cell type, and / or disease may be oversampled when generating the training dataset by at least including, in the subset of cells from the sample gene expression profile included with each annotated training sample, a different proportion of the cells associated with the underrepresented tissue type, cell type, and / or disease. For example, in some cases, the subset of cells included with each annotated training sample may be generated to include a larger proportion of the cells associated with the underrepresented tissue type, cell type, and / or disease that what is present in the available single-cell transcriptome datasets.

[0118] At 404, a training dataset may be generated to include the one or more annotated training samples. In some example embodiments, the one or more annotated training samples may include the gene expression levels of one or more subsets of cells from gene expression profiles from different tissue types and diseases. In some cases, tissue types, cell types, and / or diseases that are underrepresented in the available single-cell transcriptome datasets from which the gene expression profiles originate may be oversampled in terms of the annotated training samples included in the training dataset. For example, in some cases, the training dataset may be generated to include a larger proportion of annotated training samples associated with underrepresented tissue types, cell types, and / or disease than the original proportion of these tissue types, cell types, and / or diseases in the available single-cell transcriptome datasets.

[0119] At 406, an embedding computation model is applied to generate a sample multicellular embedding of the subset of cells included with each annotated training sample in the training dataset. In some example embodiments, to train the embedding computation model,38NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1the embedding computation model may be applied to generate the sample multicellular embedding of the subset of cells by at least embedding the gene expression levels of each cell included in the subset of cells. For example, in some cases, the subset of cells may be represented as a gene expression matrix in which each row corresponds to an individual cell in the subset of cells, each column corresponds to a specific gene, and each value corresponds to the gene expression level of a gene in a cell. Accordingly, in some cases, the embedding computation model may operate on the gene expression matrix to generate the cell embeddings before the aggregating the cell embeddings into a single vector corresponding to the sample multicellular representation of the subset of cells. In some cases, the embedding computation model may aggregate the cell embeddings by at least weighting each cell embedding with a celllevel attention-based aggregation mechanism.

[0120] At 408, a disease inference model is applied to generate, based at least on the sample multicellular embedding, an output including a result of one or more disease analytics tasks. In some example embodiments, the disease inference model may be trained by at least applying the disease inference model to perform one or more downstream tasks for disease analytics. In some cases, the one or more downstream or additional tasks for disease analytics may include comparing, clustering, and / or classifying the sample multicellular embedding generated by the embedding computation model. Examples of the one or more downstream tasks for disease analytics may include dimensionality reduction and visualization, biological feature (e.g., cell, gene, and / or the like) prioritization, treatment response and disease severity prediction, patient subgroup discovery, and / or the like.

[0121] At 410, one or more parameters of the embedding computation model and the disease inference model are adjusted to reduce a discrepancy between the output of the disease39NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1inference model and the ground-truth label associated included with a corresponding annotated training sample. In some example embodiments, the embedding computation model and the disease inference model are trained in an end-to-end manner. For example, in some cases, the embedding computation model and the disease inference model may be trained by at least reducing a loss function quantifying a discrepancy between an output of the disease computation model operating on a sample multicellular embedding of a subset of cells from a sample gene expression profile generated by the embedding computation model for each annotated training sample in the training dataset and a ground-truth label associated with the sample gene expression profile. For example, in some cases, the end-to-end training may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model and the disease inference model to reduce (or minimize) the loss function.

[0122] FIG. 4B depicts a flowchart illustrating another example of a process 450 for training an embedding computation model, in accordance with some example embodiments. In some example embodiments, the process 400 may be performed by the transcriptome analysis platform 110 to train the embedding computation model 112 in a self-supervised manner, thus obviating the need to annotated training samples. In some cases, the process 450 may implement operation 302 of the process 300 shown in FIG. 3. As described in more detail below, in some cases, the self-supervised training of the embedding computation model 112 may include reducing (or minimizing) a loss function for one or more unannotated training samples, such as the unannotated training sample 211 shown in FIG. 2B as including the sample gene expression profile 213. In some cases, the loss function may quantify a reconstruction loss corresponding to a discrepancy between the original sample gene expression profile 213 and the reconstructed gene expression profile 221 recovered from the sample multicellular representation 21540NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1generated by the embedding computation model 112. Furthermore, in some cases, the loss function may also quantify a contrastive loss quantifying a similarity between the sample multicellular representation 215 generated by the embedding computation model 112 and the one or more similar multicellular representation 219 of gene expression profiles similar to the sample gene expression profile 213.

[0123] At 452, one or more unannotated training samples are generated in which each unannotated training sample includes a subset of cells from a sample gene expression profile without a corresponding ground-truth label. In some example embodiments, the subset of cells may include the gene expression levels of some but not all of the cells included in the sample gene expression profile. In some cases, the subset of cells may be generated to include an oversampling of cells of one or more tissue types, cell types, and / or disease that are underrepresented in the available single-cell transcriptome datasets from which the sample gene expression profile originates. For example, in some cases, the subset of cells may be generated to include a different proportion of the cells associated with the underrepresented tissue types, cell types, and / or diseases than the proportion of these tissue types, cell types, and / or diseases present in the available single-cell transcriptome datasets. In some cases, each unannotated training sample may be generated without a ground-truth label, such as the ground-truth label of one or more downstream disease analytics tasks performed using a sample multicellular representation of the corresponding subset of cells from the gene expression profile included with the unannotated training sample. In some cases, obviating ground-truth labels may reduce annotation burden as well as increase the applicability of single-cell transcriptome datasets without ground-truth labels, thus increasing the volume of single-cell transcriptome datasets available to train the embedding computation model.41NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0124] At 454, a training dataset may be generated to include the one or more unannotated training samples. In some example embodiments, the one or more unannotated training samples may include the gene expression levels of one or more subsets of cells from gene expression profiles from different tissue types and diseases. In some cases, tissue types, cell types, and / or diseases that are underrepresented in the available single-cell transcriptome datasets from which the gene expression profiles originate may be oversampled in terms of the unannotated training samples included in the training dataset such that the resulting training dataset includes a larger proportion of annotated training samples associated with underrepresented tissue types, cell types, and / or disease than the original proportion of these tissue types, cell types, and / or diseases in the available single-cell transcriptome datasets.

[0125] At 456, an embedding computation model is applied to generate a sample multicellular embedding of a sample gene expression profile included with each unannotated training sample in the training dataset. In some example embodiments, to train the embedding computation model in a self-supervised manner, the embedding computation model may first be applied to generate the sample multicellular embedding of the subset of cells by at least embedding the gene expression levels of each cell included in the subset of cells. For example, in some cases, the subset of cells may be represented as a gene expression matrix in which each row corresponds to an individual cell in the subset of cells, each column corresponds to a specific gene, and each value corresponds to the gene expression level of a gene in a cell.Accordingly, in some cases, the embedding computation model may operate on the gene expression matrix to generate the cell embeddings before the aggregating the cell embeddings into a single vector corresponding to the sample multicellular representation of the subset of cells. In some cases, the embedding computation model may aggregate the cell embeddings by42NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1at least weighting each cell embedding with a cell-level attention-based aggregation mechanism. At 458, a decoder is applied to reconstruct the subset of cells included with each unannotated training sample by at least decoding a corresponding sample multicellular embedding generated by the embedding computation model. In some example embodiments, the embedding computation model is trained in a self-supervised manner by at least reducing a loss function quantifying a contrastive loss and a reconstruction loss associated with the sample multicellular embedding of the subset of cells included with each unannotated training sample in the training dataset. As noted, in some cases, the self-supervised training of the embedding computation model may include applying the embedding computation model to generate, for each unannotated training sample included in the training dataset, a sample multicellular representation of the corresponding subset of cells from a sample gene expression profile. In some cases, in order to reduce (or minimize) the reconstructive loss of the embedding computation model, a decoder may be applied to reconstruct, from each sample multicellular embedding generated by the embedding computation model, the subset of cells included with the corresponding unannotated training sample. For example, in some cases, the decoder may decode the sample multicellular embedding generated by the embedding computation model, thereby generating a reconstruction of the subset of cells from the corresponding unannotated training sample.

[0126] At 460, one or more parameters of the embedding computation model are adjusted to reduce a reconstruction loss quantifying a discrepancy between the subset of cells included with each unannotated training sample and a reconstruction of the subset of cells generated by decoding a corresponding sample multicellular embedding. In some example embodiments, the self-supervised training of the embedding computation model may include43NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1training the embedding computation model to generate multicellular embeddings that enable the gene expression levels of the corresponding subsets of cells to be recovered accurately therefrom. For example, in some cases, the reconstruction loss associated with each unannotated training sample may include a difference (e.g., cosine distance) between the gene expression levels of the subset of cells included with the unannotated training sample and those reconstructed from the multicellular representation of the subset of cells. In some cases, one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model may be adjusted in order to reduce (or minimize) the difference (e.g., cosine distance) between the gene expression levels of the subset of cells included with the unannotated training sample and those reconstructed from the multicellular representation of the subset of cells.

[0127] At 462, one or more parameters of the embedding computation model are adjusted to reduce a contrastive loss quantifying a discrepancy between the sample multicellular embeddings of similar subsets of cells. In some example embodiments, the self-supervised training of the embedding computation model may further include training the embedding computation model to reduce (or minimize) contrastive loss, which quantifies the discrepancy between the multicellular embeddings of similar subsets of cells. In some cases, two or more subsets of cells may be identified as similar if the two or more subsets of cells originate from the same sample gene expression profile. In some cases, the embedding computation model may be trained to generate similar sample multicellular embeddings for subsets of cells originating from the same sample expression profile by at least adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model to reduce (or minimize) a difference (e.g., cosine distance) between two or more multicellular representations of subsets of cells derived from the same gene expression profile.44NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0128] FIG. 5A depicts a schematic diagram illustrating the dataflow in an example of the transcriptome analysis platform 110, in accordance with some example embodiments. Referring to FIG. 5 A, in some cases, the embedding computation model 112 may generate, based at least on one or more subsets of cells 501 from the gene expression profile 111 of a patient 500, the multicellular representation 113. As further shown in FIG. 5 A, the disease inference model 114 may perform, based at least on the multicellular representation 113 of the patient 500, one or more downstream tasks for machine learning enabled disease analytics. Examples of the one or more downstream tasks shown in FIG. 5A include dimensionality reduction, biological feature prioritization, treatment response prediction, and disease severity classification.

[0129] Referring now to FIG. 5B, which depicts a schematic diagram illustrating the architecture of an example of the transcriptome analysis platform 110, in accordance with some example embodiments. In the example shown in FIG. 5B, the transcriptome analysis platform 110 may include the embedding computation model 112 and the disease inference model 114. In some cases, the gene expression profile 111 of the patient 500 may include cells of different types of cells (or cell types). In some cases, the one or more subsets of cells 501 may be selected from the gene expression profile 111, with each subset of cells including some but not all of the cells present in the gene expression profile 111. In some cases, each subset of cells 501 may include cells of a single cell type or, alternatively, multiple cell types. Moreover, in some cases, each subset of cells 501 may be represented as a gene expression matrix 502. For example, in some cases, the gene expression matrix 502 of a particular subset of cells 501 may include rows corresponding to the individual cells in the subset of cells 501, columns corresponding to specific genes, and values corresponding to the gene expression levels of the genes in each cell. In some cases, the embedding computation model 112 may generate, based at least on the gene45NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1expression matrix 502 of the one or more subsets of cells 501 included in the gene expression profile 111 of the patient 500, the multicellular representation 113.

[0130] In some cases, the embedding computation model 112 may generate the multicellular representation 113 by at least generating, for each cell included in the gene expression matrix 502 of the one or more subsets of cells 501 from the gene expression profile 111, a corresponding cell embedding 504. For example, the example of the embedding computation model 112 shown in FIG. 5B includes a cell embedder 503, which embeds the gene expression values of each cell in the gene expression matrix 502 into the cell embedding 504. Furthermore, as shown in FIG. 5B, the embedding computation model 112 may include a cell aggregator 505, which aggregates the cell embedding 504 of each cell to generate the multicellular representation 113. For instance, in some cases, the cell aggregator 505 may weigh each cell embedding 504 using a cell-level attention-based aggregation mechanism to generate the multicellular representation 113. In some cases, the disease inference model 114 may operate on the multicellular representation 113 to generate the predictive output 115 which, in the example shown in FIG. 5B, includes a disease classification corresponding to the gene expression profile 111 of the patient 500.

[0131] To further illustrate, consider an aggregated single-cell transcriptome dataset that includes N gene expression profiles denoted D = sxs2,...,sN, where strepresents the gene expression profile of the ithpatient. In some cases, the gene expression profile s, of patient i may include the gene expression levels of genes across different cells. In some cases, the gene expression profile s(- of patient i may be denoted s, = {c c2,...where Cj represents the jthcell in the gene expression profile Sj. In some cases, each cell Cj in the gene expression profile may be a vector whose features are gene expression levels (or counts). Accordingly, in some 46NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1cases, the gene expression profile s, of patient i may be represented as a matrix X, G IKM[XCis the number of cells for patient i and may vary across patients, and dgis the number of genes (e.g., dg= 28). In some cases, the gene expression profile stof patient i may be associated with (or include) patient-level metadata, such as a disease label y a tissue label tband / or the like.

[0132] In some cases, the embedding computation model 112 may include a cell encoder fg( ), a cell aggregator (•), and a classifier gg(y). In some cases, one or more of the cell encoder the cell aggregator hg(-), and the classifier gg( ) may be implemented as neural networks. In some cases, the cell encodermay generate a cell embedding for each cell Cj in the gene expression profile s, of patient i, the cell aggregator / ie(9 may aggregate the resulting cell embeddings into a multicellular representation of patient i, and the classifier gg(’) may determine a class (or label) based on the multicellular representation of patient i.

[0133] In some cases, the cell encodermay be a linear layer. In some cases, the embedding computation model 112 may encode each cell in the gene expression profile using the cell encoder fg:-> dh, where dhdenotes the dimension of the resulting cell embeddings. For example, in some cases, the output of the cell encoder fe( ) for patient i may be denoted as Zj — fe(ct), where z;- is a set of vectors {z,: j — 1,.... M of size dhand each constituent vector corresponding to a cell embedding. In some cases, this set of vectors zy- may be represented as a matrix ZLG ^Mtxdh

[0134] In some cases, a patient-level multicellular representation of patient i may be generated by the cell aggregator / i0(') having an output of e, = gg(Zi) with z =[zltz2,..., ZMI - Insome cases, the cell aggregator he( ) may generate the multicellular47NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1f the cellembeddings Zy. In some cases, different aggregators may differ in the way the weights w =[w1(..., wMj] are computed. For example, in some cases a mean aggregator uses Wj = — alinear attention aggregator uses w = softmax(z), a non-linear attention aggregator uses softmax(ci0(z)) with a being a neural network that operates on each cell embedding Zy independently, and a gated-attention uses iv = softmax(I / 0(z)) O Sigmoid(Fe(z))) with ueand VQ both being neural networks. In some cases, a transformer aggregator may differ in its architecture as it updates each cell embedding Zy according to the entire sample and sums the resulting embeddings. It should be appreciated that according to some example embodiments, the cell aggregatormay be implemented using a softmax-attention pooling layer as follows: w = softmax(a0(Zj)) (1)wherein adenotes a neural network acting on each row of the cell embedding matrix Zi independently.

[0135] In some cases, the multicellular representation patient-level embedding eLof patient i may be applied towards one or more downstream tasks for machine learning enabled disease analytics including, for example, the classification of disease, severity, subtype, subgroup, and / or the like. For example, in instances where the one or more downstream tasks include a classification task, the multicellular representation patient-level embedding etof patient i may be provided to the classifier ge('where dcdenotes the number of possible disease classes (or labels). In some cases, the disease classification of patient i may be determined as follows:48NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1p;= softmax(^(^(ef)), (3) wherein p;denotes the predicted probabilities for each disease class (or label).

[0136] In some cases, the classifier gemay be implemented as a multi-layer perceptron (MLP) with a final softmax activation. Furthermore, in some cases, the cell encoder the cell aggregator / i0(-), and the classifier g ( ) may be trained in an end-to-end manner by at least reducing (or minimizing) the discrepancy (e.g., cross-entropy) between the output p(of the classifier g0and the corresponding ground-truth label Y; associated with the gene expression profile S of patient i.

[0137] As noted, in some cases, the embedding computation model 112 may be trained end-to-end with the disease inference model 114 using a training dataset of annotated training samples, such as the annotated training sample 201 shown in FIG. 2A as including the sample gene expression profile 203 and the corresponding ground-truth label 205 (e.g., groundtruth class (or label) for a downstream classification task). In some cases, the end-to-end training of the embedding computation model 112 and the disease inference model 114 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model 112 and the disease inference model 114 to reduce (or minimize) a loss function quantifying a discrepancy (e.g., cross-entropy) between the predictive output 115 of the disease inference model 114 and the ground-truth label 205 associated with the sample gene expression profile 203.

[0138] Alternatively, instead of being trained end-to-end with the disease inference model 114, the embedding computation model 112 may be trained in a self-supervised manner to reduce (or minimize) reconstruction loss and contrastive loss. FIG. 5C depicts a schematic diagram illustrating the architecture of an example of a transcriptome analysis system in which49NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1an embedding computation model and a disease inference model are trained in a self-supervised manner, in accordance with some example embodiments. As shown in FIG. 5C, in some cases, multiple subsets of cells may be derived from the gene expression profile 111 including, for example, a first subset of cells 501a, a second subset of cells 501b, and / or the like. In the example shown in FIG. 5C, may include N cells and G genes. In some cases, each of the first subset of cells 501a and the second subset of cells 501b may include the gene expression levels G genes expressed by N cells. In some cases, the N cells included in each of the first subset of cells 501a and the second subset of cells 501b may include some but not all of the cells included in the gene expression profile 111.

[0139] Referring again to FIG. 5C, in some cases, the training of the embedding computation model 112 may include applying the embedding computation model 112 to generate a first sample multicellular representation 507a of the first subset of cells 501a. For example, in some cases, the embedding computation model 112 may be trained by at least being applied to generate, for the gene expression levels of the first subset of cells 501a from the gene expression profile 111, a first sample multicellular representation 508a. In some cases, the embedding computation model 112 may include the cell embedder 503, which generates the cell embeddings 504 by at least encoding the gene expression levels of each cell in the first subset of cells 501a. In some cases, each cell embedding 504 may be a Q-dimensional embedding generated by the cell embedder 503, where Q > G. Furthermore, in some cases, the embedding computation model 112 may include the cell aggregator 505, which generates the first sample multicellular representation 508a of the first subset of cells 501a by at least aggregating the cell embedding 504 of each cell included in the first subset of cells 501a. For instance, in some cases, the cell aggregator 505 may aggregate the cell embedding 504 of the cells included in the first subset of50NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1cells 501a by at least weighting the cell embeddings 504 with a cell-level attention-based aggregation mechanism.

[0140] As noted, in some cases, the embedding computation model 112 may be trained by at least reducing (or minimizing) a reconstruction loss. In the example shown in FIG.5C, the reconstruction loss of the embedding computation model 112 may include a discrepancy between the cell embeddings 504 generated by the cell embedder 503 and a reconstruction 510 of the cell embeddings 504 generated by a decoder 509 decoding the first sample multicellular representation 508a. In the example shown in FIG. 5C, the decoder 509 may be trained using one or more masked embeddings 511 generated by applying, for example, a mask 507 to the cell embeddings 504 to obscure one or more of the values present in the cell embeddings 504. In some cases, the decoder 509 may be trained to recover the one or more masked values in the masked embeddings 511. In some cases, once trained, the decoder 509 may be applied to generate the reconstruction 510. Moreover, in some cases, the training of the embedding computation model 112 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model 112 to reduce (or minimize) a difference (e.g., cosine distance) between the cell embeddings 504 and the reconstruction 510 of the cell embeddings 504.

[0141] In some cases, the embedding computation model 112 may be further trained by at least reducing (or minimizing) a contrastive loss. In some cases, the contrastive loss of the embedding computation model 112 may include a discrepancy between the sample multicellular representations of similar subsets of cells. In some case, two or more subsets of cells may be considered similar in this context if the two or more subsets of cells originate from the same gene expression profile 111. For example, as shown in FIG. 5C, in some cases, the first subset of cells51NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1501a and the second subset of cells 501b may be identified as similar subsets of cells for purposes of contrastive loss based at least on the first subset of cells 501a and the second subset of cells 501b originating from the same gene expression profile 111. Accordingly, in some cases, the training of the embedding computation model 112 may include adjusting one or more parameters (e.g., weights, biases, and / or the like) of the embedding computation model 112 to reduce (or minimize) a difference (e.g., cosine distance) between the first sample multicellular representation 508a of the first subset of cells 501a and a second multicellular representation 508b of the second subset of cell 501b.

[0142] To further illustrate the training of the embedding computation model 112 to reduce (or minimize) a joint contrastive and reconstruction loss, at each training step, the embedding computation model 112 may operate on a set of batch size 2 x S gene expression matrices, each of which including C cells and G genes matrices from S different sample gene expression profiles. In some cases, the genes may be aligned to a predetermined set of G = 33,680 genes, with the gene expression levels of any missing genes set to zero. In some cases, the gene expression levels of each of the G genes expressed by each of the C cells may be encoded into a Q-dimensional vector by the cell embedder 503, with Q < G. In some cases, each of the resulting C X Q table may be randomly masked (e.g., by applying the mask 507) to replace, for example, randomly, a portion of the values with zeros. In some cases, the resulting masked embeddings 511 may be passed through the cell aggregator 505 to generate a 1 x Q vector corresponding to, for example, the first sample multicellular representation 508a, the second multicellular representation 508b, and / or the like. In some cases, the 1 X Q vector may be decoded by the decoder 509 to generate the reconstruction 510 of the cell embeddings 504, which is a C x Q matrix. In some cases, the reconstruction loss of the cell embedding model 11252NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1may be the difference between the C x Q matrix corresponding to the reconstruction 510 and the original C X Q matrix of the cell embeddings 504. In some cases, the contrastive loss C of the embedding computation model 112 may be the cross entropy between the distance (e.g., cosine distance) between the first sample multicellular representation 508a of the first subset of cells 501a and the second sample multicellular representation 508b of the second subset of cells 501b. In this context, cross-entropy, which is also known as logarithmic loss or log loss, measures the difference between the probability distribution of the output of the cell embedding model 112 and the probability distribution of the multicellular representations of similar subsets of cells. It should be appreciated that a lower cross-entropy value generally indicates better performance, hence the training of the cell embedding model 112 may include reducing (or minimizing) crossentropy loss. In some cases, the overall loss function of the embedding computation model 112 may be expressed as a weighted sum in accordance with the expression below:aR + (1 — <z) x C (4) wherein R denotes the reconstruction loss and C denotes the contrastive loss,

[0143] In some example embodiments, the predictive output 115 of the disease inference model 114 may be interpreted using integrated gradients (IG). For example, in some cases, where the disease inference model 114 is applied to perform a classification task, the predictive output 115 of the disease inference model 114 may include the probabilities of each possible class (or label). Moreover, in some cases, the predictive output 115 may be interpreted based on an importance metric of each cell, cell type, gene, or sets of genes to quantify the corresponding contribution to the predictive output 115 of the disease inference model 114. To further illustrate, FIG. 5D depicts a schematic diagram illustrating the interpretability of the multicellular representations generated by an example of the embedding computation model 112,53NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1in accordance with some example embodiments. For example, in some cases, a gradient attribution may be determined, using integrated gradients, for each cell-gene combination present in the gene expression profile 111 of the patient 500 to generate an attribution matrix Rt£M,xds having the same dimensions as the matrix XtE [RM'xds forming the gene expression profile Si of patient i.

[0144] In some cases, the attributions in the matrix of attributions may be averaged across different dimensions to support different levels of interpretability. For instance, in some cases, averaging the attributionsover genes may yield the importance metrics of each individual cell while averaging the attributions R[ over cells may yield the importance metrics of individual genes. As shown in FIG. 5D, in some cases, the attributions Ri may also be averaged over groups of cells, cell types, groups of genes, or individual genes within a group of cells (or cell type) to yield importance metrics for individual cells, cell types, genes, genes expressed by certain cell types, and / or the like.

[0145] Experimental Examples

[0146] The performance of various example embodiments of the embedding computation model described herein was evaluated for different downstream tasks for machine learning enabled disease analytics. The embedding computation model was trained on a heterogeneous dataset including 24.3 million gene expression profiles from single cell ribonucleic acid (RNA) sequencing (scRNA-seq) data from over 5,000 patient samples and spanning 135 unique disease-state labels, across 413 studies, and 189 tissues (organs). Each patient contributed to a single sample, thus allowing patient and sample to be used interchangeably in the following text. Cells were profiled using droplet based single-cell RNA sequencing (scRNA-seq) from 10X Genomics. The dataset was split into a training set (60%), a54NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1validation set (20%), and test set (20%), ensuring that all samples from a given study are in the same split. The dataset was preprocessed by removing profiles with no gene expression levels and normalizing the remaining profiles to the corrected sequencing depth followed by a log(% + 1) transformation. A visual summary of the composition of the dataset and the corresponding splits are shown in FIGS. 6A-C. FIG. 6A shows an overview of the composition of the training dataset across tissue types and diseases. The diameter of the bubbles in FIG. 6A corresponds to the quantity of cells included for a given tissue type or disease. FIG. 6B and C depicts graphs illustrating statistics of different disease states (FIG. 6B) and tissue types (FIG. 6C). As shown in FIGS. 6B-C, the dataset was imbalanced in terms of diseases and tissue types. For example, while COVID-19 samples accounted for ~9% of the samples, multiple sclerosis accounted for merely ~2% of the samples (Extended Data Fig. 2 (a) and (b)).

[0147] Disease Classification from Gene expression profiles

[0148] The embedding computation model was trained to predict the disease label associated with each sample in the dataset. The performance of the embedding computation model was quantified in terms of a weighted Fl -score and compared against different embedding baselines, such as a s pseudo-bulk approach, cell-type proportions (CTP), as well as state-of-the-art single-cell foundation models (CellPLM and SCimilarity). It should be appreciated that Fl score is robust to class imbalance and reflects both precision and recall across all classes. For each of these methods, two classifiers were used to predict the label from the patient embedding: a fc -Nearest Neighbor Classifier (kNN) and a multi-layer perceptron (MLP). FIG. 7A depicts a comparison of the performance of the embedding computation model (ECM) on multi-disease classification relative to the baseline methodologies. Table 1 below depicts a summary of the results shown in FIG. 7A.55NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0149] Table 1CellPLM SCimilarity Pseudobulk MultiMIL ECM 0.302 ± 0.029 0.304 ± 0.016 0.448 ± 0.018 0.500 ± 0.000 0.789 ± 0.016

[0150] As shown in FIG. 7A, the embedding computation model described herein outperformed every baseline by a significant margin. Notably, the pseudo-bulk approach outperforms more complicated foundation models in this task. Additional results on a binary classification task (e.g., COVID-19 vs. healthy) are shown in FIG. 11. The results depicted in FIG. 11 show that the embedding computation model even outperformed the most recent domain-expert model ScRAT by a large margin.

[0151] The performance impact of different aggregation mechanisms for aggregating cell embeddings into a patient-level multicellular representation, including mean-pooling, transformer, gated attention, linear attention, and non-linear attention mechanisms, was also investigated. The results are shown in FIG. 7B, which indicates that non-linear attention performed best, improving the weighted Fl-score by 16.6% compared to a mean-pooling mechanism. The transformer approach, albeit more expressive, results in poor performance, likely due to the number of parameters (e.g., weights, biases, and / or the like) exceeding what is necessary for this task. A softmax-attention layer was found to be the most effective in these ablation studies.

[0152] The performance impact of introducing cell-type information was also investigated. The baseline uses cell-type proportions to represent each patient and classifies binary conditions using a fc-nearest neighbor (kNN) classifier (CTP-kNN). The loss function used to train the embedding computation model includes a contrastive loss to bring cells from the 56NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1same type closer together in embedding space, while pushing apart cells of different types (ECM-CT). Such an approach can reduce the potential batch effect existing in the training data.

[0153] To address the class imbalance in the dataset, the performance impact of different resampling per disease class and per tissue type was also investigated. The weighted Fl scores yielded by resampling per disease-class and per tissue-class are shown in FIG. 7B. As shown in FIG. 7B, oversampling the training set for both disease and tissue resulted in a significant improvement compared to baseline.

[0154] The patient embedding space determined by the embedding computation model organized by disease state is shown in FIG. 8A while the patient embedding space organized by tissue type is shown in FIG. 8B. Notably, COVID-19 patients partition into clusters corresponding to blood and lung tissue samples. Additional analyses of the patient embedding space, aggregated per disease and tissue type, are shown in FIGS. 14A-14C.

[0155] The impact that different training and hyper-parameter have on the sensitivity of the embedding computation model is shown in FIGS. 12A-12H. During the training process of the embedding computation model, different factors that could affect the training process, including hyper-parameters, size of sampled cells, composition of diseases, and scaling law were explored. The sensitivity analyses focusing on these factors shed light on the difficulties of patient modeling on a broader scale. As shown in FIG. 12A, a learning rate closer to le — 4 reduces the negative effects associated with over-fitting. Meanwhile, a smaller number of epochs (< 40 for binary classification and < 5 for multi-label classification, based on the epoch with the best validation accuracy) also contributed to better model performances (FIG. 12B). The dropout rate and weight decay rate also performed better with a small value (FIG. 12B and 12C). Increasing the number of sampled cells does not always enhance the performances of,57NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1shown in FIG. 12D for the experiments based on binary classification. A suitable range of sampled cell numbers was found to be in (100, 2000).

[0156] Biological Feature Prioritization

[0157] Disease classifications made based on the multicellular embeddings generated by the cell embedding model was interpreted using integrated gradients (IG), which yields importance metrics for individual cells and genes. The importance metric enabled a fine-grained analysis of the individual cells and genes that contribute most to a disease of interest, which was COVID- 19 in this case.

[0158] Using integrated gradients (IG), cell type level attributions was determined to uncover what cell types were contributing most to the COVID-19 label for each patient (FIG. 9A). The highest average attributions (computed over all patients) are found for classical monocytes and platelets, suggesting the importance of these cell types in COVID-19.Remarkably, this fine-grained analysis enabled further exploration within cell types of interest, including what genes had the most impact on COVID-19 prediction for each of these specific cell types. For each patient, an importance metric was computed for each gene in monocytes (FIG. 9B) and in platelets (FIG. 9C). Doing so determined the significance of each gene in a given cell type. Ranking genes by averaging the corresponding importance metrics revealed that S100A8, IFITM3, and IFI27 are the most pertinent genes in monocytes while IFI27, HBB, and CAI are the most pertinent genes in platelets. These genes were identified as being associated with COVID-19 severity or treatment response. A set of important genes was also uncovered by the embedding computation model by measuring the overlap with the set of differentially expressed genes from ToppCell. A Fisher’s exact test indicates strong overlap for both classical monocytes (p-value = 2.1e — 22) and platelets (p — value = 2.5e — 20).58NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1

[0159] FIG. 16 depicts a bubble chart illustrating the Gene Ontology Enrichment Analysis (GOEA) results of representative immune cells based on prioritized genes. With a p-value of 0.05 set as threshold, FIG. 16 provides a visualization of the results of Gene Ontology Enrichment Analysis (GOEA) based on selected gene sets from classical monocytes and non-classical monocytes. As shown in FIG. 16, the prioritized genes for classical monocytes and non-classical monocytes can form meaningful pathways related to immunological function, such as defense response to virus and inflammatory response. Therefore, the identified pathways are representative enough to support the usefulness of the embedding computation model for multiple diseases and tissues (cell states) simultaneously.

[0160] FIG. 17 depicts the results of Wilcoxon rank sum tests for the association between attribution scores and COVID-19 severity for each cell-type, with p-values shown being Bonferroni corrected.

[0161] Disease Severity Classification

[0162] To investigate the effectiveness of the multicellular representations generated by the embedding computation model, four single-cell RNA sequencing (scRNA-seq) datasets were collected from COVID-19 patients where a severity label is available (mild or severe) and that were not included during training. Visualizing the multicellular represented generated by the embedding computation model, the landscape was shown as being primarily organized by disease severity and not by study (FIG. 10A). Conversely, a principal components analysis (PCA) representation of pseudo-bulk data is organized primarily by study rather than severity, highlighting batch effects.

[0163] Moreover, the importance metrics given by the embedding computation model to different cell types correlates with disease severity, with significant associations (corrected p<59NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-10.01) for natural killer (NK) cells, B cells, myeloid dendritic cells, and mucosal-associated invariant T (MAIT) cells. There is a significant difference in the magnitude of the integrated gradients (1G) attributions of the embedding computation model, averaged over all myeloid dendritic cells in each patient sample, between mild and severe patient groups (FIG. 10B, Bonferroni-corrected p-value=0.001, rank sum test). Similarly, there is a significant association between the disease severity and the magnitude of the probability of CO VID-19 diagnosis predicted using the multicellular representations generated by the embedding computation model (FIG. 10C). Together, these results show that the embedding computation model described herein can implicitly represent the disease severity of each patient.

[0164] Differences and Similarities of Different Classes (or Labels) in Representation Space

[0165] Modeling patients is a complex problem. As such, merely using classifier performance to measure patient-level multicellular representations may be insufficient to show that the embedding computation model was successfully trained to generate meaningful patientlevel multicellular representations. Accordingly, a given group of patient-level multicellular representations was categorized by disease as well as by tissue type of their origin to study the effects of the same disease on different tissues. The patient-level multicellular embeddings were averaged across both diseases and tissue type to generate the correlation matrix shown in FIG.13. As shown in FIG. 13, the multicellular representations of COVID-19 and lung adenocarcinoma have a high correlation across different tissue types, which implied that the embedding computation model successfully learned multicellular representations across different tissue types with the same disease. The multicellular representations generated by the embedding computation model also captured signals of different tissue types within the same60NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1disease demonstrated by the results of hierarchical clustering. Such findings are consistent with the multi-tissue damages caused by COVID-1957 and lung adenocarcinoma, as these diseases tended to affect different tissue types jointly.

[0166] Rigorously evaluating the quality of the multicellular representations produced by a given method is a challenging task. For example, this evaluation can require metadata annotations that is not always available. To address this challenge, a disease similarity measure was constructed based on the text descriptions of each disease, extracted from the Kyoto Encyclopedia of Genes and Genomes (KEGG) database and National Center for Biotechnology Information (NCBI). Each disease description was converted to an embedding using the OpenAI text-embedding tool. Similarity between diseases was then obtained by computing the Pearson Correlation Coefficients (PCCs) between their respective embedding. The resulting similarity matrix is given in FIG. 15 A.

[0167] A similarity measure between diseases from the multicellular representations generated by the embedding computation model was also constructed by averaging the multicellular representations of all patients with a given disease and computing the pairwise Pearson Correlation Coefficients (PCCs). A similar procedure was used for computing disease similarities from pseudo-bulk data. The resulting similarity matrices are presented in FIG. 15A as well.

[0168] To quantitatively assess the discrepancy between the text-based disease similarities and embedding computation model-based similarities, pairwise Pearson Correlation Coefficients (PCCs) were computed between both similarity matrices. The correlation from the result of the embedding computation model (PCC=0.65, p-value=9.5e-33) was higher than the correlation from pseudo-bulk gene expression levels (PCC=0.28, p-value=2.6e-6), suggesting61NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1that the embedding computation model was able to learn better patient representations by considering both the differences and similarities across different diseases.

[0169] Furthermore, the ability of the embedding computation model in capturing the similarity of certain disease-tissue embeddings with other embeddings was investigated by visualizing the relationship between the correlation from text embeddings and the correlation from the embedding computation model. These results are shown in FIGS. 15B-C. As shown, the embedding computation model aligned with the understanding of diseases with patient representations from transcriptomic data for the lung samples with lung cancer (PCC=0.83, p-value=1.2e-4) and for the lung samples with COVID-19 (PCC=0.72, p-value=2.0e-3). Overall, these preliminary results show that text embeddings can be a promising metric for evaluating the reliability of learned patient representations.

[0170] Identification of Treatment Responders

[0171] To complement the quality -assessment of the patient-level multicellular representations generated by the embedding computation model, whether these multicellular representation can be used to identify treatment responders for T-cell immunotherapy in melanoma was also investigated using two single-cell RNA sequencing (scRNA-seq) datasets from patients with melanoma treated with T-cell immunotherapy, and for which a binary treatment outcome label was available. The former dataset was used as a training dataset while the latter as a testing dataset. For reference (or baseline comparison), pseudo-bulk data from the original gene expression profiles by patients was used with a Support Vector Classifier (SVC). Patient-level multicellular representations were generated by querying the embedding computation model with the original gene expression profiles. A treatment response classification was performed using the resulting patient-level multicellular representations. The62NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1whole workflow is shown schematically in FIG. 18A. Sample multicellular representations generated with gene expression profiles in the training dataset are visualized FIG. 18B, which further depicts annotations to differentiate between the clusters of responders and nonresponders. The classification results depicted in FIG. 18C shows that classifications using the multicellular representations from the embedding computation model as input outperform the baseline model under different classification metrics, especially because the baseline model predicted all patient samples as drug-responsive.

[0172] FIG. 18D depicts a bar graph illustrating a comparison of the performance for various disease analytics tasks using a random approach, a pseudo-bulk approach, an embedding computation model trained in an end-to-end manner, and an embedding computation model trained in a self-supervised manner, in accordance with some example embodiments. The disease analytics tasks shown in FIG. 18D include SEA- AD ADNC scoring (evaluated using quadratic weighted kappa), longitudinal COVID timepoint (evaluated using weighted Fl), taurus treatment (evaluated using weighted Fl), and HLCA IPF X-study (evaluated using weighted Fl). As shown in FIG. 18D, the self-supervised trained embedding computation model generally outperformed state-of-the-art techniques as well as the end-to-end trained embedding computation model.

[0173] FIG. 19 depicts a block diagram illustrating an example of a computing system 1900, in accordance with some example embodiments. Referring to FIGS. 1, 2A-2B, and 19, the computing system 1900 may be used to implement the transcriptome analysis platform 110, the sequencing platform 120, the client device 130, and / or any components therein.

[0174] As shown in FIG. 19, the computing system 1900 can include a processor 1910, a memory 1920, a storage device 1930, and input / output devices 1940. The processor63NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-11910, the memory 1920, the storage device 1930, and the input / output devices 1940 can be interconnected via a system bus 1950. The processor 1910 is capable of processing instructions for execution within the computing system 1900. Such executed instructions can implement one or more components of, for example, the transcriptome analysis platform 110, the sequencing platform 120, the client device 130, and / or the like. In some example embodiments, the processor 1910 can be a single-threaded processor. Alternately, the processor 1910 can be a multi -threaded processor. The processor 1910 is capable of processing instructions stored in the memory 1920 and / or on the storage device 1930 to display graphical information for a user interface provided via the input / output device 1940.

[0175] The memory 1920 is a computer readable medium such as volatile or nonvolatile that stores information within the computing system 1900. The memory 1920 can store data structures representing configuration object databases, for example. The storage device 1930 is capable of providing persistent storage for the computing system 1900. The storage device 1930 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 1940 provides input / output operations for the computing system 1900. In some example embodiments, the input / output device 1940 includes a keyboard and / or pointing device. In various implementations, the input / output device 1940 includes a display unit for displaying graphical user interfaces.

[0176] According to some example embodiments, the input / output device 1940 can provide input / output operations for a network device. For example, the input / output device 1940 can include Ethernet ports or other networking ports to communicate with one or more wired64NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0177] In some example embodiments, the computing system 1900 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various formats. Alternatively, the computing system 1900 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via the input / output device 1940. The user interface can be generated and presented to a user by the computing system 1900 (e.g., on a computer screen monitor, etc.).

[0178] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of65NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0179] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.

[0180] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For66NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.

[0181] In the descriptions above and in the claims, phrases such as “at least one of’ or “one or more of’ may occur followed by a conjunctive list of elements or features. The term “and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at least one of A and B;” “one or more of A and B;” and “A and / or B” are each intended to mean “A alone, B alone, or A and B together.” A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.

[0182] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related67NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desired results. Other implementations may be within the scope of the following claims.68NAI-5004501617v1

Claims

Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1CLAIMSWhat is claimed is:

1. A computer-implemented method, comprising:generating, based at least on one or more single-cell transcriptome datasets, a training dataset, wherein the one or more single-cell transcriptome datasets include sample gene expression profiles associated with a plurality of different tissue types, cell types, and / or diseases;training, based at least on the training dataset, an embedding computation model; receiving a gene expression profile of a patient, wherein the gene expression profile includes a gene expression level of a plurality of genes expressed by each cell of a plurality of cells;determining a subset of cells from the gene expression profile of the patient; applying the embedding computation model to generate, based at least on the subset of cells from the gene expression profile of the patient, a multicellular representation for the patient, wherein the embedding computation model generates the multicellular representation by at least aggregating a plurality of cell embeddings generated by embedding the gene expression level of each cell included in the subset of cells from the gene expression profile of the patient; andperforming, based at least on the multicellular representation of the patient, one or more downstream tasks.

2. The method of claim 1, wherein the subset of cells from the gene expression profile is represented as a gene expression matrix.69NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-13. The method of claim 2, wherein the gene expression matrix includes one or more rows corresponding to one or more individual cells in the subset of cells from the gene expression profile, one or more columns corresponding to one or more genes, and one or more values corresponding to one or more gene expression levels.

4. The method of any of claims 2 to 3, wherein the embedding computation model generates each cell embedding of the plurality of cell embeddings by at least embedding the one or more gene expression levels of a corresponding cell in the subset of cells.

5. The method of any of claims 1 to 4, wherein the embedding computation model generates the multicellular representation of the patient by applying a cell-level attention-based aggregator to aggregate the plurality of cell embeddings.

6. The method of claim 5, wherein the cell-level attention-based aggregator aggregates the plurality of cell embeddings by at leastdetermining a plurality of weights, andaggregating the plurality of cell embeddings by at least applying the plurality of weights to determine a weighted sum of the plurality of cell embeddings.

7. The method of any of claims 5 to 6, wherein the cell-level attention-based aggregator comprises one or more of a mean aggregator, a linear aggregator, a non-linear aggregator, or a transformer aggregator.

8. The method of any of claims 5 to 7, wherein the cell-level attention-based aggregator comprises a softmax-att ention pooling layer.

9. The method of any of claims 5 to 8, wherein the cell-level attention-based aggregator is applied to aggregate the plurality of cell embeddings into a single vector corresponding to the multicellular representation of the patient.70NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-110. The method of any of claims 1 to 9, wherein the embedding computation model is trained in an end-to-end fashion with a disease inference model applied to perform the one or more downstream tasks.

11. The method of claim 10, wherein an end-to-end training of the embedding computation model and the disease inference model includes adjusting one or more parameters of the embedding computation model and the disease inference model to reduce a loss function quantifying a discrepancy in an output of the disease inference model.

12. The method of any of claims 10 to 11, wherein an end-to-end training of the embedding computation model and the disease inference model includesgenerating the training dataset to include a plurality of annotated training samples, wherein each annotated training sample of the plurality of annotated training samples includes a subset of cells from a sample gene expression profile and a corresponding ground-truth label, applying the embedding computation model to determine, for each annotated training sample included in the training dataset, a corresponding sample multicellular representation, applying the disease inference model to perform, based at least on the corresponding sample multicellular representation, the one or more downstream tasks, andadjusting one or more parameters of the embedding computation model and the disease inference model to reduce a loss function quantifying a discrepancy between an output of the disease inference model and the corresponding ground-truth label of each annotated training sample in the training dataset.

13. The method of any of claims 1 to 12, wherein the embedding computation model is trained in a self-supervised manner using a plurality of unannotated training samples.71NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-114. The method of claim 13, wherein the training the embedding computation model in the self-supervised manner includesgenerating the training dataset to include a plurality of unannotated training samples, wherein each unannotated training sample of the plurality of unannotated training samples includes a subset of cells from a sample gene expression profile without a corresponding groundtruth label,applying the embedding computation model to determine, for each unannotated training sample included in the training dataset, a corresponding sample multicellular representation, and adjusting one or more parameters of the embedding computation model to reduce a reconstruction loss and a contrastive loss associated with the corresponding sample multicellular representation of each annotated training sample included in the training dataset.

15. The method of claim 14, wherein the reconstruction loss of a sample multicellular representation quantifies a similarity between a plurality of cell embeddings generated for the sample gene expression profile and a reconstruction of the plurality of cell embeddings recovered from the sample multicellular representation.

16. The method of any of claims 14 to 15, wherein the contrastive loss of a sample multicellular representation of one subset of cells from a sample gene expression profile quantifies a similarity between the sample multicellular representation and a sample multicellular representation of a different subset of cells from a same sample gene expression profiles.

17. The method of any of claims 1 to 16, wherein the training dataset is generated to include an oversampling of one or more underrepresented tissue types, cell types, and / or diseases.

18. The method of claim 17, wherein the oversampling includes72NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1determining a subset of cells from a sample gene expression profile to include a different proportion of the one or more underrepresented tissue types, cell types, and / or diseases than an original proportion of the one or more underrepresented tissue types, cell types, and / or diseases present in the sample gene expression profile, andgenerating, for inclusion in the training dataset, a training sample including the subset of genes having an oversampling of the one or more underrepresented tissue types, cell types, and / or diseases.

19. The method of any of claims 17 to 18, wherein the oversampling includes generating, for inclusion in the training dataset, a different proportion of training samples associated with the one or more underrepresented tissue types, cell types, and / or diseases than an original proportion of the one or more underrepresented tissue types, cell types, and / or diseases present in the one or more single-cell transcriptome datasets.

20. The method of any of claims 1 to 19, further comprising:applying a disease inference model to perform the one or more downstream tasks; and generating a plurality of gradient attributions by at least determining, for each cell-gene combination present in the gene-expression profile of the patient, a gradient attribution corresponding to a rate of change in an output of the disease inference model performing the one or more downstream tasks.

21. The method of claim 20, further comprising:averaging the plurality of gradient attributions across the plurality of cells present in the gene expression profile to determine, for each cell of the plurality of cells, an importance metric indicative of a contribution of the cell to the output of the disease inference model.

22. The method of any of claims 20 to 21, further comprising:73NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-1averaging the plurality of gradient attributions across the plurality of genes present in the gene expression profile to determine, for each gene of the plurality of genes, an importance metric indicative of a contribution of the gene to the output of the disease inference model.

23. The method of any of claims 20 to 22, further comprising:averaging the plurality of gradient attributions across a plurality of cell types present in the gene expression profile to determine, for each cell types of the plurality of cell types, an importance metric indicative of a contribution of the cell type to the output of the disease inference model.

24. The method of any of claims 20 to 23, further comprising:averaging the plurality of gradient attributions across a plurality of combinations of genes and cell types present in the gene expression profile to determine, for each gene-cell type combination of the plurality of combinations of genes and cell types, an importance metric indicative of a contribution of the gene-cell type combination to the output of the disease inference model.

25. The method of any of claims 20 to 24, wherein the disease inference model is trained to perform a disease classification task.

26. The method of claim 25, further comprising:adapting the disease inference model to determine disease severity by at least identifying, based at least on the plurality of gradient attributions, one or more prioritized biological features, anddetermining, based at least on a proportion of the one or more prioritized biological features, disease severity.74NAI-5004501617v1Attorney Ref.: 14786-081-228 (103963-228081) / P39726-WO-127. The method of claim 26, wherein the one or more prioritized biological features include one or more biological features identified as having a threshold contribution to an output of the disease inference model performing the disease classification task.

28. The method of any of claims 26 to 27, wherein the one or more prioritized biological features include one or more cells, cell types, genes, or genes expressed by certain cell types.

29. The method of any of claims 1 to 28, wherein the one or more downstream tasks are performed by comparing, clustering, and / or classifying the multicellular representation of the patient.

30. The method of any of claims 1 to 29, wherein the one or more downstream tasks include one or more of dimensionality reduction and visualization, biological feature prioritization, treatment response prediction, disease severity prediction, or patient subgroup identification.

31. A system, comprising:at least one data processor; andat least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1 to 30.

32. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 1 to 30.75NAI-5004501617v1

Citation Information

Patent Citations

  • US202463711011P