Visualization of intracellular co-expression of cellular components using dimensionality reduction weighted by cell type.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2026-08-14
Smart Images

Figure CN122580702A_ABST
Abstract
Description
[0001] Related applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 619,659, filed January 10, 2024, entitled “VISUALIZING INTRA-CELL CO-EXPRESSION OF CELLULAR CONSTITUENTS USING DIMENSIONALITY REDUCTION THAT ISWEIGHTED BY CELL TYPE”, which is hereby incorporated herein by reference in its entirety.
[0002] In the event of any conflict between this application and any document incorporated by reference, this application shall prevail. Background Technology
[0003] Single-cell analysis techniques enable the measurement of expression levels of different components within individual cells, including components such as proteins and RNA transcripts. For example, flow cytometry passes single cells through a laser path and examines them using a variety of visible and fluorescent light sources to allow for the assessment of protein composition; mass flow cytometry uses heavy metal ion tags instead of fluorescent dyes as markers and reads them using time-of-flight spectroscopy; and instruments that index transcriptomes and epitopes through sequencing (“CITE-Seq”) use DNA barcoded antibodies to detect proteins.
[0004] Common uses of these single-cell analysis techniques involve collecting cell samples from a specific subject; using an instrument to apply one of the analysis techniques to each cell to obtain the expression level of each of one or more components; and outputting a table identifying the measured expression level of each component in each cell. Attached Figure Description
[0005] Figure 1 It is a block diagram showing some of the components of a computer system and other equipment that are typically incorporated into a facility on which it operates.
[0006] Figure 2 It is a flowchart illustrating a process performed by a facility in some implementations to visualize the intracellular co-expression of cellular components.
[0007] Figure 3 This is a table illustrating an example of the contents of a cell component expression level table, which in some embodiments is used by a facility to store the expression levels of multiple cell components detected in each of multiple individual cells of a cell sample by a single-cell analyzer.
[0008] Figure 4This is a hierarchy diagram, showing example hierarchies used by the facility in some implementations.
[0009] Figure 5 This is a schematic diagram illustrating an example acyclic hierarchy diagram generated by a facility in some implementation schemes.
[0010] Figure 6 An example of a cell type adjacency matrix table used by a facility in some implementations is shown, which is used to identify cell types whose nodes in a graph are directly connected.
[0011] Figure 7 This is a table illustrating an example of a cell type frequency table used by a facility in some embodiments, which stores the frequency of each assigned cell type in a sample—that is, the percentage of cells in a cell sample that are assigned to each of the assigned cell types.
[0012] Figure 8 This is a diagram illustrating an example of the contents of a weighted cell-type adjacency matrix table used by a facility in some implementations. The weighted cell-type adjacency matrix table is used to store the weights assigned to edges that are directly connected pairs of nodes in the graph.
[0013] Figure 9 This is a diagram illustrating an example of a weighted cell type distance table used by a facility in some implementations. The weighted cell type distance table is used to store the weighted distances between pairs of nodes representing corresponding cell type pairs in the graph.
[0014] Figure 10 The table illustrates an example of the contents of a normalized weighted cell type distance table used by the facility in some implementations, which is used to store variance-normalized weighted cell type distances.
[0015] Figure 11 This is a table diagram illustrating an example of the contents of a cell type low-dimensional embedding table used by the facility in some embodiments, which is used to store low-dimensional embeddings of cell type representations created by the facility in action 208.
[0016] Figure 12 This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.
[0017] Figure 13 This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.
[0018] Figure 14This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.5.
[0019] Figure 15 This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.75.
[0020] Figure 16 This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.99.
[0021] Figure 17 It is a flowchart illustrating a process performed by a facility in some implementations to generate alternative visualizations, where cellular visual representations of colors indicate the expression levels of specific components.
[0022] Figure 18 It is a visualization that shows according to Figure 17 The visualization generated by the process shown. Detailed Implementation
[0023] The inventors have recognized the limitations of conventional methods for single-cell analysis. Specifically, they have determined that the ability to visualize intracellular co-expression levels in samples in two or more dimensions would be of significant value. Specifically, they recognize that this would provide enhanced capabilities for understanding disease biology, making disease diagnoses, and understanding the mechanisms of action of drug candidates (to name just a few).
[0024] With the help of modern single-cell technologies such as spectral flow cytometry or CITESeq, it is now possible to measure dozens to hundreds of proteins and thousands of RNA transcripts in thousands to millions of individual cells. Whether studying cells from healthy subjects, disease states, or animal models, there is usually a knowledge base related to cell lineage and function, which is often applied first to new data. Therefore, it is common practice to use protein or RNA expression to identify various cellular components, such as T cells, B cells, and additional subsets; for example, it is well established that T cells are lymphocytes that express the protein CD3. Gating is a commonly used method by which domain experts classify individual cells into various cell types, usually in a hierarchical manner. These cell types are often well understood from a lineage or functional perspective. This cell classification helps to extract additional insights from individual cell data; that is, new findings are best understood within the context of known cell types and subtypes. Cell typing typically involves measuring a few dimensions (usually protein markers) for each cell. The expression patterns of these biomarkers are well understood in the literature; for example, T cells are known to express CD3, while B cells are known to express CD19. Furthermore, the lineages of these cell types ( For exampleThe classification of cell types (including T cells and B cells as subsets of lymphocytes) has been well established. Therefore, it is suggested that this understanding of lineages and expression patterns is utilized, typically visualized using a two-dimensional graph, to identify cell types.
[0025] For example, a typical goroutine plot might show CD14 and CD45 expression used to identify monocytes and lymphocytes, followed by CD3 and CD19 expression in lymphocytes used to identify T cells and B cells, and then CD4 and CD8 expression in T cells used to identify CD4+ T cells and CD8+ T cells. In the plot, each point represents a cell, and the color indicates the density of cells in a specific region of the plot—red indicates many cells with data values close to each other, while blue indicates a sparse distribution.
[0026] Alternative methods (such as clustering) are also used to identify cell types, typically in more exploratory settings when cell lineages are not fully understood, or to further subdivide cell types after phylogenetic analysis.
[0027] The goal of defining cell types is to create context for additional data analysis. The number of biomarkers used to define cell types (typically protein biomarkers) is usually small (10 to 20 biomarkers). On the other hand, it is now possible to measure many more (hundreds to thousands) of additional protein and RNA biomarkers. Analyzing these additional biomarkers in the context of cell type and disease or other information related to the biosample is highly valuable. Therefore, it is common practice to quantify proteins, RNA, or other biomarkers in these cell types, especially those not yet used to define cell types, for further analysis. Visual analysis of single-cell data to complement quantification is also highly desirable to fully understand the heterogeneity of biomarker expression (and co-expression) between cell types. Given the high-dimensional nature of the data, visualization is often challenging. Dimensionality reduction methods (such as t-SNE and UMAP) are often applied to this high-dimensional single-cell data to visualize various cellular components and the expression of proteins or RNA or other biomarkers within these components. In general, these nonlinear methods attempt to learn the high-dimensional manifold of the data and provide a lower-dimensional (typically two-dimensional) approximation. They attempt to preserve the relationships (distances) between cells that are close to each other in high-dimensional space, i.e., neighborhood, at the cost of distorting the distances between distant cells, when reducing to a lower-dimensional space.
[0028] While dimensionality reduction provides a convenient way to visualize data in lower dimensions, the computation of dimensionality reduction is typically independent of the definition of cell types. Therefore, there is no guarantee that biologically significant cell types will be visually separated in a lower-dimensional space to easily facilitate the analysis of other biomarkers in that context. The inventors have recognized that cell types often overlap in lower-dimensional space, making it challenging to understand the differential expression of biomarkers within and between cell types.
[0029] The inventors have determined that this overlap or blurring of cell types in the projection space can be caused by two reasons. First, t-SNE or similar algorithms are often applied to dimensions beyond those involved in cell typing (protein and RNA expression). This is generally desirable because a goal of visual exploration is to determine whether, beyond cell typing, there are protein / RNA expressions that lead to the formation of cell subclusters. In other words, there is a natural competition between lineage markers used for cell type definition and other markers that may have similar expression in different cell types. Excluding non-lineage markers from dimensionality reduction can improve the separation between cell types, but the understanding of any cell subclusters formed by the expression of other markers will be limited. The second reason for the overlap is that cell lineage markers can vary significantly across different samples due to subject-to-subject variability or due to in vitro cell treatment. For example, in vitro activation of T cells will alter several T cell markers, but still provide sufficient resolution to define T cells and related subpopulations via ging. Variations in lineage markers between samples mean that the same cell type will appear in different regions of different samples in the projection space, making visual analysis more difficult. More importantly, since users are already aware that phylogenetic markers vary between samples, highlighting the projection of this difference at the expense of diluting patterns in other markers is generally not beneficial during analysis.
[0030] Based on this understanding, the inventors have conceived and put into practice a software and / or hardware facility (“the Facility”) for visualizing the intracellular co-expression of cellular components (such as proteins and RNA transcripts) using cell type-weighted dimensionality reduction.
[0031] Conventional dimensionality reduction methods evaluate two points ( For exampleThe method aims to determine whether cells are similar or dissimilar. This similarity assessment is first performed in a high-dimensional (“raw”) space. The method then seeks to preserve this similarity and dissimilarity between cells as much as possible in a low-dimensional (projected) space. For example, the t-SNE dimensionality reduction method computes the Euclidean distance in the high-dimensional space and converts it into a probabilistic similarity metric to generate a probability distribution. Then, through a minimization procedure, t-SNE attempts to preserve the probability distribution of the high-dimensional space in the low-dimensional space. t-SNE is further described in Visualizing Data using t-SNE, Laurens van der Maaten, Geoffrey Hinton, Journal of Machine Learning Research 9 (2008) 2579-2605, which is hereby incorporated in its entirety by reference.
[0032] While UMAP differs methodologically from t-SNE, it also relies on identifying which cells are close (similar) to each other and which are not. UMAP calculates Manhattan distance to determine points / cells that are closer to each other and connects them with higher probability to create a graph representation of the probabilistic data. UMAP then attempts to preserve the graph's structure in a low-dimensional space. Distance metrics ( For example The Manhattan distance (rather than the Euclidean distance) and other adjustable parameters defining the neighborhood size can be changed to obtain the optimal projection for different types of datasets. UMAP is further described by UMAP: Uniform Manifold Approximation and Projection, Leland McInnes, John Healy, Nathaniel Saul, Lukas, Großberger, Journal of Open Source Software, 3(29), 861, https: / / doi.org / 10.21105 / joss.00861, which is hereby incorporated in its entirety by reference.
[0033] Compared to conventional dimensionality reduction methods, the facility adjusts points / cells that are similar to each other, making those cells belonging to the same cell type more likely to be similar to each other.
[0034] In some implementations, the facility allows single-cell analyzer output data from a single subject's cell sample or "well" to undergo a process (such as gating or clustering) that assigns cell types to each cell in the sample based on the co-expression levels of combinations of characteristic components from different cell types. As an example, in some implementations, the facility operates to assess the co-expression of proteins in tumor-infiltrating cells extracted from lung cancer patients. In some implementations, samples from different subjects are pooled.
[0035] For each cell type belonging to a particular sample, the facility establishes a low-dimensional representation (such as a two-dimensional representation) of the cell type such that similarly assigned cell types are closer to each other than dissimilarly assigned cell types. In some implementations, this involves determining a hierarchy associated with the assigned cell types; generating an acyclic graph representing the hierarchy; determining the distance between each pair of assigned cell types in the graph, the distance being adjusted in some cases to reflect the frequency of the pair of cell types in the sample; determining a high-dimensional representation of each cell type based on the graph distance of each cell type to other assigned cell types; and applying dimensionality reduction techniques to these high-dimensional representations of each cell type to obtain a low-dimensional representation as an embedding.
[0036] The facility determines emphasis weights, which specify the degree to which cell types will be emphasized relative to the expression of cell components in the visualization generated by the facility, such as by receiving emphasis weights as user input. The facility generates a cell matrix, where each row represents a cell of the sample, and the cell matrix includes two sets of columns: (a) a variance-normalized version of the expression level detected in the cell for each cell component, and (b) a low-dimensional embedding of the cell type. The number of columns in the first and second sets varies depending on the implementation. In the cell matrix, the values in the columns of the first set are weighted relative to the values in the columns of the second set using emphasis weights. The facility subjects each row of the cell matrix to dimensionality reduction techniques to obtain visualization coordinates, which are used to draw a visual representation of the cells in a low-dimensional visualization space, such as a two-dimensional visualization space.
[0037] The inventors have used a neighborhood purity metric, reflecting the degree of co-localization of cells of the same cell type, as the basis for evaluating facility-generated visualizations. On a per-cell-type basis, the inventors observed a significant improvement in neighborhood purity in facility-generated visualizations compared to conventional visualizations that do not consider cell type, cell type lineage, or cell type hierarchy.
[0038] In some implementations, for a single cell type, the facility identifies one or more components that were not used to identify cells of that cell type, but whose expression levels direct specific cells to different subclusters within that cell type. This is achieved by selecting components and generating new visualizations in which the color of a visual indicator for each cell is based on the expression level of the selected component.
[0039] By operating in some or all of the manner described herein, the facility enables users to easily visualize the co-expression of components in a cell sample, including, for a single cell type, identifying one or more components that were not used to identify cells of that cell type, but whose expression levels direct specific cells to different subclusters within that cell type.
[0040] Furthermore, the facility enhances the functionality of computers or other hardware, such as by reducing the dynamic display area, processing, storage, and / or data transfer resources required to perform a specific task, thereby enabling that task to be performed by less powerful, smaller-capacity, and / or lower-cost hardware, and / or with lower latency, and / or reserving more of the saved resources for other tasks. For example, by generating more useful dimensionality-reduced visualizations than conventional methods, the facility avoids the processor cycles and memory footprint required to repeatedly perform conventional processes in pursuit of better and more useful visualizations.
[0041] Furthermore, for at least some of the domains and scenarios discussed in this paper, the processes described herein as automatically executed by computing systems are actually impossible to execute in the human brain for reasons including: the initial data, one or more intermediate states, and the final data are too large and / or poorly organized for human acquisition and processing, and / or belong to forms that the human brain cannot perceive and / or express; the data manipulation operations and / or sub-processes involved are too complex, and / or too different from typical human mental operations; the required response time is too short for human performance to meet; and so on.
[0042] Figure 1This is a block diagram illustrating some of the components of at least some of the computer systems and other devices typically incorporated into a facility that operates thereon. In various embodiments, these computer systems and other devices 100 may include server computer systems, virtual machines in cloud computing platforms or other configurations, desktop computer systems, laptop computer systems, netbooks, mobile phones, personal digital assistants, televisions, cameras, automotive computers, electronic media players, etc. In various embodiments, the computer system and device include zero or more of the following: a processor 101 for executing computer programs and / or training or applying machine learning models, such as a CPU, GPU, TPU, NNP, FPGA, or ASIC; a computer memory 102, such as RAM, SDRAM, ROM, PROM, etc., for storing programs and data, including facility and related data, operating system including kernel, and device drivers, when in use; a persistent storage device 103, such as a hard disk drive or flash drive, for persistently storing programs and data; a computer-readable media drive 104, such as a floppy disk, CD-ROM, or DVD drive, for reading programs and data stored on a computer-readable medium; and a network connection 105 for connecting the computer system to other computer systems to send and / or receive data, such as via the Internet or another network and its networking hardware, such as switches, routers, repeaters, cables and optical fibers, optical transmitters and receivers, radio transmitters and receivers, etc. Figure 1 The components shown and discussed above do not constitute the data signal itself. While the computer system configured as described above is typically used to support the operation of a facility, those skilled in the art will understand that the facility can be implemented using various types and configurations of devices with various components.
[0043] Figure 2 This is a flowchart illustrating a process performed by a facility in some embodiments to generate a visualization of intracellular co-expression of cellular components. In action 201, the facility receives analysis results directly or indirectly from a single-cell analyzer. The analysis results indicate the expression level detected in each of a plurality of cellular components for each cell in a cell sample (such as a sample obtained from a single human or other animal patient or other subject) processed by said analyzer. These analysis results may, for example, be stored in a data table.
[0044] Figure 3This is a table illustrating an example of the contents of a cell component expression level table, which in some embodiments is used by a facility to store the expression levels of multiple cell components detected by a single-cell analyzer in each of multiple individual cells of a cell sample. The cell component expression level table 300 consists of rows, such as rows 301-313, each corresponding to a different cell in the sample. Each row is divided into multiple columns. These columns include columns such as columns 351-355, each corresponding to a different cell component and containing a value indicating the expression level of said cell component in the cell corresponding to each particular row. For example, column 351 has content indicating the expression level of cell component Hu.CD152 in different cells. The value 0.183460072 at the intersection of column 351 and row 301 indicates the expression level of this cell component in the specific cell corresponding to row 301. Additional columns in the cell component expression level table 300 are discussed herein.
[0045] Although Figure 3 The content and organization of the tables illustrated in each table diagram discussed herein are intended to facilitate understanding by human readers. However, those skilled in the art will understand that the actual data structure used by the facility to store the information may differ from the illustrated tables. For example, the actual data structure may be organized in a different way; may contain more or less information than shown; may be compressed, encrypted, and / or indexed; may contain significantly more rows than shown, and so on. Furthermore, in some embodiments, the facility does not store the data shown in the table diagrams in tables, but rather in a semi-structured or unstructured data store such as JSON objects.
[0046] return Figure 2In action 202, the facility assigns cell types to each cell in the sample. For example, in various embodiments, the facility uses cell type identification techniques such as phylogenetics and clustering. Clustering techniques are described, for example, in Leiden: From Louvain to Leiden: guaranteeing well-connected communities, VA Traag, L. Waltman & N. J. van Eck, available at arxiv.org / abs / 1810.08473; and PARC: Ultrafast and accurate clustering of phenotypic data of millions of single cells, Shobana V Stassen, Dickson MD Siu, Kelvin CM Lee, Joshua WK Ho, Hayden KH So, Kevin K Tsia; Bioinformatics, Vol. 36, No. 9, May 2020, pp. 2778–2786, each of which is hereby incorporated in its entirety by reference.
[0047] return Figure 3 The cell component expression level table also includes a cell type column 356, in which the facility records the cell type identified for each cell in action 202. For example, the value at the intersection of column 356 and row 301 indicates that the cell corresponding to row 301 has been assigned the cell type "B cell".
[0048] return Figure 2 In action 203, the facility establishes a hierarchy that organizes all cell types assigned by the facility across cell samples in action 202. In establishing this hierarchy, the facility aims to create a structure in which cell types most similar in one or more aspects are close to each other within the hierarchy. In some embodiments, the hierarchy established by the facility reflects a known evolutionary lineage of a particular cell type, such that if a cell of a first type is considered to have evolved from a cell of a second type, then the first cell type is established as a child node of the second type within the hierarchy. For example, in some embodiments, as discussed below… Figure 4In the illustrated hierarchy 400, the facility establishes lymphocyte (Lymphs) 420 nodes as child nodes of single cell (Singlets) 410 nodes, reflecting the understanding that lymphocytes evolved from single cells. In some embodiments, step 203 involves automatically filtering pre-existing hierarchies of all cell types that might be assigned in action 202 to include only the cell types actually assigned for the current cell sample in step 202.
[0049] In some implementations, the facility enables users of the facility to determine the hierarchy used by the facility. In some such implementations, for example, the facility provides a user interface that shows the proposed hierarchy, in which users can manipulate the hierarchy by moving nodes representing different cell types to new positions within the hierarchy.
[0050] Figure 4 This is a hierarchy diagram illustrating an example hierarchy used by a facility in some embodiments. In hierarchy 400, child node cell types are indented below their parent node cell types. For example, in hierarchy 400, CD4+ CD45RA+ cell type 451 is a child node of CD4+ T cell cell type 450, which in turn is a child node of T cell cell type 440, which in turn is a child node of lymphocyte cell type 420, which in turn is a child node of single cell cell type 410. The cell types shown in the hierarchy are those cell types assigned to the cells in the sample, which are shown in... Figure 3 The expression levels of the cellular components shown are 300.
[0051] In various implementation schemes, facility use differs Figure 4 The table below shows various cell type hierarchies. Table 1 below illustrates the cell type hierarchies used by the facility in conjunction with flow cytometry cell gates in some embodiments.
[0052] Table 1
[0053]
[0054] Table 2 below shows the cell type hierarchy used by the facility in conjunction with CITESeq in some implementations.
[0055] Table 2
[0056]
[0057] Back Figure 2In action 204, the facility generates an acyclic graph representing the hierarchy established in action 203. In some implementations, this involves creating nodes representing each cell type that appears in the established hierarchy and connecting each child node to a node of its parent cell type via edges.
[0058] Figure 5 This is a schematic diagram illustrating an example acyclic hierarchy diagram generated by a facility in some implementations. Figure 500 consists of nodes connected by edges. Each node corresponds to a node from... Figure 4 One of the cell types in the cell type hierarchy shown. For example, it can be seen that a single cell node 510 corresponding to a single cell cell type 410 is connected to a lymphocyte node 520 corresponding to a lymphocyte cell type 420, which in turn is connected to a T cell node 540 corresponding to a T cell cell type 440, and so on.
[0059] Back Figure 2 In action 205, for each distinct pair of cell types assigned in action 202, the facility determines the distance between their nodes in the graph, said distance being weighted according to the frequency of each cell type in the cells of the cell sample for that pair. See below for reference. Figures 6 to 10 Discuss the details of action 205.
[0060] Figure 6 An example of a cell type adjacency matrix table used by a facility in some embodiments is shown, which is used to identify cell types whose nodes in a graph are directly connected. Cell type adjacency matrix table 600 consists of rows (such as rows 601-612), each row corresponding to one of the assigned cell types shown in column 650. Each row also contains additional columns 651-662, also corresponding to the assigned cell types. At each intersection of a row and substantial columns 651-662, the value indicates whether the pair of cell types represented by the row and column are directly connected—i.e., connected by a single edge. For example, the intersection of row 602 and column 651 indicates that a node of the lymphocyte cell type is directly connected to a node of the single cell type.
[0061] Figure 7 This is a table illustrating an example of a cell type frequency table used by a facility in some embodiments. The cell type frequency table stores the frequency of each assigned cell type in a sample—that is, the percentage of cells in the cell sample assigned to each of the assigned cell types. For example, line 701 indicates that the frequency of a single cell type being assigned to cells in the sample is 0.0144222—that is, 1.44222% of the cells in the sample are assigned to a single cell type.
[0062] Figure 8This is a diagram illustrating an example of the contents of a weighted cell type adjacency matrix table used by a facility in some embodiments. This weighted cell type adjacency matrix table is used to store the weights assigned to edges of directly connected pairs of nodes in a graph. The organization of the weighted cell type adjacency matrix table 800 corresponds to... Figure 6 The tissue shown is represented by cell type adjacency matrix table 600. The facility generates a weighted cell type adjacency matrix table 800 by replacing each value 1 in cell type adjacency matrix table 600 with a weight determined by the following formula: B[i,j] = sqrt(Fi + Fj) in Fi and Fj They are cell types i and j The frequency. For example, the value 0.687769308 at the intersection of column 851 and row 802 indicates the weight of combinations of lymphocyte cell types and individual cell types, which is determined by taking the frequency of these cell types (in... Figure 7 The values were obtained by taking the square root of the sum of 0.0144222 and 0.0449522, shown in rows 701 and 702 of the cell type frequency table 700.
[0063] The facility then proceeds to construct a weighted cell type distance table using the contents of the weighted cell type adjacency matrix table 800, which contains the weighted distances between each pair of different cell types in the graph, including those cell types that are directly and indirectly connected in the graph. The facility achieves this by finding the shortest path between nodes representing each pair of cell types and summing the weights shown in the weighted cell type adjacency matrix table 800 for each edge traversed along said path.
[0064] Figure 9This is a diagram illustrating an example of the contents of a weighted cell type distance table used by a facility in some embodiments. This table stores weighted distances between pairs of nodes representing corresponding cell type pairs in the graph. The organization of the weighted cell type distance table 900 is similar to that of the cell type adjacency matrix table 600 and the weighted cell type adjacency matrix table 800. The contents at each intersection of a row and a substantial column are determined in the manner discussed above: the shortest path between nodes corresponding to the cell types of said row and column is found, and the weights of the edges in said path are summed in the weighted cell type adjacency matrix table 800. For example, the facility obtains the value 1.716708936 at the intersection of row 801 of a single cell and row 901 of a lymphocyte by summing 0.687769308 at the intersection of column 855 of a T cell and row 802 of a lymphocyte. This value represents the weight of the edge between single cell node 510 and T cell 540 in the graph—that is, the edge between single cell node 510 and lymphocyte node 520, and the edge between lymphocyte node 520 and T cell node 540. In some embodiments, the facility performs variance normalization on the weighted cell type distances shown in the weighted cell type distance table 900.
[0065] Note that the CTD generated from the plot representation of cell types is in an arbitrary coordinate system independent of BD. Variance normalization places the coordinate systems of CTD and BD on an equal footing. The weighting of CTD and BD, as described below, is then directly related to their expected contributions. Variance normalization simply divides BD by the sum of the standard deviations of all BDs. The same operation is performed to normalize CTD; that is, CTD is divided by the sum of the standard deviations of all CTDs. Alternative forms exist for normalizing BD and CTD, or more broadly, for making their coordinates compatible. Note that all cells belonging to the same cell type will be set to the same CTD value. Since the number of BDs is typically much larger than the number of CTDs and varies between datasets, normalizing BD and CTD by their respective total variances ensures they are at similar scales before being combined to create an augmented data matrix. Adjustable weights w A convenient mechanism is provided to change the relative importance of BD compared to CTD in the final step of dimensionality reduction, which is discussed in the next step.
[0066] Figure 10 This is a table illustrating an example of the contents of a normalized weighted cell type distance table used by a facility in some implementations. This table stores variance-normalized weighted cell type distances. The organization of the normalized weighted cell type distance table 1000 is consistent with the above discussion. Figure 6 , Figure 8 and Figure 9 The organizations shown in the tables are similar. The values generated in this table are obtained by the facilities in the manner described above.
[0067] return Figure 2 In actions 206-209, the facility iterates through each cell type assigned in action 202. In action 207, the facility constructs a representation of the cell type, which consists of weighted distances between that cell type and all cell types assigned in action 202. For example, in some embodiments, this constructed representation is a concatenation of these distances in a consistent order among the constructed representations. In action 208, the facility creates a low-dimensional embedding of the cell type representation. In action 209, if there are still additional cell types to process, the facility continues with action 206 to process the next cell type; otherwise, the facility continues with action 210.
[0068] In some implementations, the facility uses techniques such as t-SNE; multidimensional scaling (“MDS”); or graph layout methods such as Graph Drawing by Force-DirectedPlacement, Fruchterman, Thomas MJ; Reingold, Edward M. (1991), Software:Practice and Experience, Wiley, 21 (11): 1129–1164, doi:10.1002 / spe.4380211102; algorithms for drawing general undirected graphs, Kamada, Tomihisa; Kawai, Satoru (1989), Information Processing Letters, Elsevier, 31 (1): 7–15, doi:10.1016 / 0020-0190(89)90102-6; and spring inserters with force-directed graph drawing algorithms, Kobourov, Stephen G. (2012). Available at arxiv.org / abs / 1201.3011, each of which is hereby incorporated in its entirety by reference.
[0069] Figure 11This is a table illustrating an example of the contents of a cell type low-dimensional embedding table used by a facility in some embodiments to store low-dimensional embeddings of cell types created by the facility in action 208. Each row of the cell type low-dimensional embedding table 1100 corresponds to a different cell type, as shown in column 1150. Each row further contains two values from columns 1151 and 1152, which together constitute an embedding—in other words, each embedding is a two-dimensional embedding. For example, row 1101 indicates that the facility has created a low-dimensional embedding (-1.701187264, 9.486135891) for a single cell type. Again, these embeddings are assigned by the facility in such a way that the distance between each pair of cell type embeddings in two-dimensional space reflects the level of similarity between these cell types.
[0070] return Figure 2 In action 210, the facility receives input selecting an emphasis weight, which specifies the degree to which cell types will be emphasized relative to the expression levels of cell components in the visualization generated by the facility. In some embodiments, this emphasis weight W is 0 to emphasize cell components while completely excluding cell types, 1 to emphasize cell types while completely excluding cell components, or an intermediate value to represent a different mixture of these two considerations.
[0071] In action 211, the facility generates a cell matrix, where each row represents a cell, and each row contains two sets of columns. The first set of columns corresponds to a variance-normalized version of the expression level detected in the cell for each cellular component, and the second set of columns corresponds to a low-dimensional embedding of the cell type. The first set of columns is weighted relative to the second set of columns based on or using emphasis weights. In some implementations, the facility uses the following formula to populate the cell matrix.
[0072]
[0073] N: Cell number M: Number of component dimensions L: The number of dimensions for cell type embedding, for example, 2 B N×M Variance calibration matrix for protein / RNA data C N×L Variance calibration matrix of cell type embedding w: Adjustable weight In some implementations, a grid of values is formed by a cell matrix generated by the facility in action 211, where each row represents each cell in the sample, wherein a first set of columns corresponds to the expression level of the component detected for each cell in the cells, and a second set of columns corresponds to a representation established for each of the cell types of the cells, the values in the first set of columns being weighted relative to the values in the second set of columns according to an emphasis weight.
[0074] In some implementations, a grid of values is formed by a cell matrix generated by the facility in action 211, where each row represents a cell in the sample, wherein a first set of columns corresponds to the expression level of the component detected for the cell, and a second set of columns corresponds to a representation established for the cell type of the cell, the values in the first set of columns being weighted relative to the values in the second set of columns according to an emphasis weight. Figure 3 The cell component expression level table 300 shown is an example of such a cell matrix, where each row in rows 301-313 represents one cell in the sample; columns 351-355 are a first set of columns corresponding to the component expression levels detected for the cells, and columns 357 and 358 are a second set of columns corresponding to the representation established for the cell type of the cells.
[0075] In action 212, the facility performs dimensionality reduction on each row of the cell matrix to obtain the visual coordinates of the cells corresponding to those rows. For example, to generate a two-dimensional visualization, the facility performs dimensionality reduction to produce a two-dimensional representation of each row. In various implementations, the facility employs a variety of dimensionality reduction techniques, including t-SNE or UMAP as described above. In action 213, the facility generates a visualization for the cell sample, wherein the visual representation of each cell is located based on the visual coordinates obtained for that cell in action 202, and the visual representation of the cell is colored based on the cell type assigned to the cell. Various visualizations generated by the facility in action 213 for the cell sample represented in the cell component expression level table 300 are shown in the discussion below. Figures 12 to 16 In Action 214, the facility enables the visualization generated in Action 213 to be presented, stored, and / or automatically analyzed. Following Action 214, the facility continues in Action 210 by receiving new inputs with different emphasis weights and repeating the process of generating visualizations using these new emphasis weights.
[0076] Those skilled in the art will understand that Figure 2 The actions shown, as well as those shown in each flowchart discussed below, can be changed in a variety of ways. For example, the order of the actions can be rearranged; some actions can be performed in parallel; the actions shown can be omitted or other actions can be included; the actions shown can be divided into sub-actions or multiple actions shown can be combined into a single action, etc.
[0077] Figures 12 to 16 A visualization generated by the facility from cell samples shown in Table 300, which shows the expression levels of cell components. Figure 12 This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0. In visualization 1200, each circle represents a cell in the cell sample, and its color corresponds to the cell type assigned to that cell. Additionally, rectangular text labels are placed near the center of the cell location for each cell type. For example, the label "B cell" is located near the center of the blue cell cluster 1201 representing a sample cell assigned cell type B. It can be seen in this visualization that, with W=0 selected to emphasize cell components while completely excluding cell types, many cell types are mixed together. For example, the areas with reference numerals 1202, 1203, and 1204 contain circles of three different mixed colors, corresponding to cells assigned cell types lymphocytes, CD4+ T cells, and CD4+CD45RA+ cells. This mixing may interfere with the usefulness of this visualization for some purposes.
[0078] Although Figure 12 Each of the display diagrams discussed below illustrates a display whose format, organization, information density, etc., is best suited to certain types of display devices. However, those skilled in the art will understand that the actual display presented by the facility may differ from those shown, as they may be optimized for specific other display devices, or may omit shown visual elements, include visual elements not shown, or have been reorganized, reformatted, re-visualized, or shown at different magnification levels, etc.
[0079] Figure 13 This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.25. In visualization 1300, region 1301 of B cells is... Figure 12 The B cells 1201 shown are clustered slightly more tightly. Regions 1302, 1303, and 1304, representing lymphocytes, CD4+ T cells, and CD4+CD45RA+ cells respectively, are now clearly distinguishable, and... Figure 12 The mixed formations at locations labeled 1202, 1203, and 1204 in the attached figure provide a contrast. Therefore, the greater emphasis on cell type achieved by choosing W=0.25 instead of W=0 helps to see and visually grasp the location of cells within a specific cell type, while preserving the proximity between cells of different cell types but with similar component expression levels.
[0080] Figure 14 This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.5. Figure 15This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.75. As the emphasis weight value increases, the level of separation between cell clusters of the same cell type continues to increase.
[0081] Figure 16 This is a visualization showing a sample cell sample generated by the facility using an emphasis weight value of 0.99. In visualization 1600, cells of different cell types are clustered very tightly and far from other groups, making it difficult to visualize the similarity of cellular components between groups and cell types.
[0082] Figure 17 This is a flowchart illustrating a process performed by a facility in some implementations to generate alternative visualizations, where cellular visual representations use color to indicate the expression levels of specific components. This is useful, for example, when seeking to understand the reasons behind the span of a large set of circles for a particular cell type, as well as significant distances between certain circles, and differential distances to circles of other cell types. For example, in... Figure 14 In this study, group 1410, consisting of CD8+CD45RA+ cells, spans a large area and includes two distinct leaflets, 1411 and 1412. Using this process can help elucidate the relationships between cells in the two leaflets and with other cell types.
[0083] In action 1701, the facility receives input selecting cellular components (such as cellular component Hu.CD27). In action 1705, the facility generates a visualization in which the visual representation of each cell is located based on its visual coordinates, and the visual representation of each cell is colored based on the expression level of the cellular component selected in action 1701. In action 1703, the facility causes the visualization generated in action 1702 to be presented, stored, and / or automatically analyzed. After action 1703, this process ends.
[0084] Figure 18 It is a visualization that shows according to Figure 17 The visualization shown is generated by the process. Visualization 1800 contains... Figure 14 The same circles are shown in the visualization 1400, but they are based on... Figure 17 The process was colored differently. Specifically, the color reflected different expression levels of the cellular component Hu.CD27. The correspondence between color and expression level is shown in Figure 1850, where curve 1880 shows the frequency of different expression levels of this component in all cells of the sample at each of several different expression levels 1860 1870. For example, blue was assigned to a lower expression level of approximately 3 1861 (in 10... 0 With 10 1Between), and yellow-green was assigned to approximately 85% of the expression level in 1962 (i.e., in 10). 1 With 10 2 (between). By examining the circles in group 1810, it can be seen that those circles in leaf 1811 have a color corresponding to the higher expression level 1962 of Hu.CD27 (i.e., yellow-green), while the circles in leaf 1812 correspond to the lower expression level 1861 (i.e., blue).
[0085] The various embodiments described above can be combined to provide further embodiments. All U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patent applications, and non-patent publications mentioned in this specification and / or listed in the application data sheets are incorporated herein by reference in their entirety. If necessary, aspects of the embodiments may be modified to incorporate the concepts of various patents, applications, and publications to provide additional embodiments.
[0086] Based on the detailed description above, these and other changes can be made to the embodiments. Generally, the terminology used in the appended claims should not be construed as limiting the claims to the specific embodiments disclosed in the specification and claims, but should be interpreted to include all possible embodiments and the full scope of their equivalents. Therefore, the claims are not limited by this disclosure.
Claims
1. A method for performing actions on a cell sample in a computing system, the method comprising: Obtain analysis results indicating the detected expression level of each of multiple cellular components for each cell in the sample; Obtain the cell type of each cell in the sample, wherein the cell type belongs to multiple cell types based on the expression level of cell components in each cell; For each pair of cell types among the multiple cell types, obtain the similarity level between each pair of cell types; For each of the multiple cell types, a first representation of each cell type is established based on the acquired similarity level between each cell type and each other cell type; Obtain emphasis weights, which specify the degree to which cell type will be emphasized relative to the expression level of cell components when determining visualization coordinates for each cell of the sample; Generate a cell matrix comprising a grid of values, wherein each row represents a cell in the sample, wherein a first set of columns corresponds to the expression level of a component detected for the cell, and a second set of columns corresponds to the first representation established for the cell type of the cell, wherein the values in the first set of columns are weighted relative to the values in the second set of columns according to an acquired emphasis weight; as well as Dimensionality reduction is performed on the rows of the generated cell matrix to obtain the visual coordinates of each cell in the sample.
2. The method of claim 1, further comprising constructing a visualization image, the visualization image containing, for each cell of the sample, a visual indication of the cell, the visual indication appearing at a spatial location specified by the visualization coordinates obtained for the cell.
3. The method of claim 1, wherein in the constructed visualization image, each visual indicator of a cell is displayed in a color corresponding to the cell type of the cell.
4. The method of claim 1, wherein in the constructed visualization image, each visual indicator of a cell is displayed with a color corresponding to the expression level indicated by the cellular component distinguished for the cell.
5. The method of claim 4, further comprising receiving an input specifying the distinguished cellular components.
6. The method of any one of claims 2 to 5, further comprising rendering the constructed visualization image.
7. The method of any one of claims 2 to 6, further comprising persistently storing the constructed visualization image.
8. The method of any one of claims 2 to 7, further comprising invoking automatic analysis for the constructed visualization image.
9. The method of any one of claims 1 to 8, wherein for each of the plurality of cell types, establishing a first representation of the cell type comprises: For each of the multiple cell types, obtain the distance in the adjacency graph between the multiple cell types, which reflects the hierarchy established for the multiple cell types; For each of the multiple cell types, a second representation of the cell type is constructed by concatenating the distance values obtained for the cell type based on all cell types of the multiple cell types; The second representation of the cell type of the multiple cell types is embedded into the embedding space to obtain the first representation of the cell type of the multiple cell types.
10. The method of claim 9, wherein the embedding is performed using a process selected from: t-distributed random neighborhood embedding (t-SNE); Multidimensional scaling (MDS); Force-oriented layout; The Kamada algorithm for drawing general undirected graphs; and Kobourov Spring Embedder and Force Direction Map Drawing Algorithm.
11. The method of any one of claims 9 and 10, further comprising: Receive input specifying the hierarchy to be established for the various cell types; The adjacency graph is constructed according to the hierarchy, wherein each of the multiple cell types is represented by a node, and the nodes are directly or indirectly connected by edges; and For each of the multiple cell types, the distance is determined by calculating the minimum number of edges between a pair of nodes representing the cell type.
12. The method of any one of claims 9 to 11, further comprising: For each of the cell types, determine the frequency of that cell type in the sample; as well as The value of the distance obtained is determined by weighting the acquired distance according to the determined frequency, based on all cell types for the multiple cell types.
13. The method of any one of claims 1 to 12, further comprising receiving input specifying the acquired emphasis weights.
14. The method of any one of claims 1 to 13, wherein the dimensionality reduction is performed using a process selected from: t-distributed random neighborhood embedding (t-SNE); and Unified manifold approximation and projection (UMAP).
15. One or more computer-readable media instances, all having content configured to cause a computing system to perform the method of any one of claims 1 to 14.
16. One or more memories that jointly store a cell matrix data structure, said data structure comprising: Multiple first entries, each first entry corresponding to a cell among multiple cells constituting a cell sample, each first entry including: A first group of one or more values, wherein the one or more values collectively represent the expression level of a cellular component detected in the cell corresponding to the first entry; and A second set of one or more values, collectively comprising a representation of a cell type determined for the cell, wherein the cell type representations are assigned such that the distance between a pair of cell type representations represents the level of dissimilarity between the pair of cell types. The second group of one or more values is weighted relative to the first group of one or more values according to an emphasis weight. The content of the data structure is made available to create a visualization of the cell sample, the visualization including: for each of the plurality of cells, a visual indication of the cell placed at a spatial location determined based on the content of the first entry corresponding to the cell.
17. The one or more memories of claim 16, wherein each of the plurality of first entries further comprises: The coordinates determined by the visual indication of the cell corresponding to the first entry, based on the first set of values and the second set of values of the first entry.
18. One or more memories as claimed in any one of claims 16 and 17, wherein the data structure further comprises: Data representing a visualization of the cell sample, the visualization comprising: for each of the plurality of cells in the cell sample, a visual indication of the cell placed at a spatial location determined based on the content of a first entry corresponding to the cell, the visual indication being colored according to the cell type determined for the cell.
19. One or more memories as claimed in any one of claims 16 and 17, wherein the data structure further comprises: Data representing a visualization of the cell sample, the visualization comprising: for each of the plurality of cells in the cell sample, a visual indicator of the cell placed at a spatial location determined based on the content of a first entry corresponding to the cell, the visual indicator being colored according to the expression levels of distinguished cellular components detected in the cell.
20. One or more memories as claimed in any one of claims 16 to 19, wherein the data structure further comprises: Multiple second entries, each second entry corresponding to one of multiple cell types determined for the multiple cells of the sample, each second entry including: The cell types are represented such that the distance between a pair of cell types represents the level of dissimilarity between the pair of cell types.