Method for molecular marking of cell mass and application thereof

By combining molecular marking of cell clumps and single-cell sequencing technology, the problem of difficulty in obtaining cell interaction information and single-cell gene expression profiles in the prior art is solved, and high-throughput and high-resolution cell interaction analysis is achieved, reducing experimental costs and complexity.

CN120099140APending Publication Date: 2025-06-06TSINGHUA UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510288407.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

It is difficult to obtain cell interaction information and single-cell gene expression profiles at the same time in the prior art. The acquisition of cell interaction information is biased or incomplete, and it is difficult to study the gene expression characteristics of cell interaction at the single-cell level. It is also high cost and complex in operation, making it difficult to meet the needs of large-scale and high-throughput cell interaction research.

Method used

By molecularly labeling the cell clumps, giving them a unique combination tag, tracing the cell clumps to which a single cell belongs, directly determining the interaction between cells, and combining single-cell sequencing technology to achieve high-resolution analysis of cell interactions within the cell clumps.

Benefits of technology

It realizes high-throughput and high-resolution revealing the cell interaction network, accurately determines the gene expression characteristics of the interacting cells, and directly analyzes the interaction relationship without complex calculations, reducing experimental costs and complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120099140A_ABST
    Figure CN120099140A_ABST
Patent Text Reader

Abstract

The invention relates to a method for molecular marking of a cell mass and application thereof, and belongs to the technical field of cell interaction. Specifically, the method comprises the step of carrying out combined label labeling treatment on a cell block mass. The method is suitable for labeling the cell masses of various tissue types, the cell masses are endowed with unique combination labels through molecular markers, the cell masses to which single cells belong are traced, the interaction between the cells is directly determined, the resolution and accuracy of interaction analysis between the cells are improved, and meanwhile, the risk of misclassification in the experiment process is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of cell interaction, and in particular, to a method for molecularly labeling cell aggregates and an application thereof. Background Art

[0002] Cell-cell interaction (CCI) plays a key role in maintaining tissue homeostasis, regulating immune responses and other biological processes. In-depth research on cell-cell interactions is of great significance for understanding the complexity of biological systems, revealing disease mechanisms and developing new treatments.

[0003] Currently, the methods used to study cell interactions include: single-cell sequencing technology (scRNA-seq), spatial transcriptomics technology (ST), proximity labeling-based cell interaction technology (Proximity labeling), cell cluster deconvolution method, spatial nuclei indexing technology (Spatial nuclei indexing) and tissue fragment sequencing (fragment-seq), etc. However, the shortcomings of these methods are:

[0004] 1) Unable to obtain cell interaction information and single-cell gene expression profiles simultaneously: Traditional single-cell sequencing loses cell interaction and spatial information. Spatial transcriptome technology often has insufficient resolution or complex operation, making it difficult to simultaneously meet the needs of analyzing cell interactions and single-cell gene expression;

[0005] 2) The acquisition of cell interaction information is biased or incomplete: the proximity labeling method is limited to pre-selected cell types, and the cluster deconvolution method relies on computational inference, which is inaccurate and cannot obtain true interaction information;

[0006] 3) It is difficult to study the gene expression characteristics of cell interactions at the single-cell level: Existing methods cannot directly obtain the gene expression profiles of interacting cells, which is not conducive to studying the gene expression characteristics of interacting cells and cannot deeply analyze the impact of interactions on cell functions;

[0007] 4) High cost, complex operation, and difficulty in large-scale application make it difficult to meet the needs of large-scale, high-throughput cell interaction research. Summary of the invention

[0008] The present application aims to solve at least one of the technical problems in the related art to a certain extent. To this end, the present application proposes a cell interaction analysis method based on molecular labeling of cell clusters, which is applicable to a variety of tissue types, can reveal the cell interaction network with high throughput and high resolution, accurately determine the gene expression characteristics of interacting cells, and directly analyze the interaction relationship without complex calculations.

[0009] In a first aspect of the present application, the present application proposes a method for molecularly labeling a cell mass. According to an embodiment of the present application, the method includes: performing a combined labeling process on the cell mass.

[0010] The above method is applicable to the labeling of cell clusters of various tissue types. It can give cell clusters unique combination labels through molecular markers, trace the cell clusters to which single cells belong, directly determine the interactions between cells, improve the resolution and accuracy of cell-to-cell interaction analysis, and reduce the risk of misclassification during the experiment.

[0011] In the second aspect of the present application, the present application proposes a sequencing method. According to an embodiment of the present application, the method comprises: performing molecular labeling on a cell mass using the method of the first aspect; performing dissociation on the cell mass after molecular labeling; and performing sequencing on a single cell obtained after dissociation.

[0012] This method can directly capture cell interactions by molecularly labeling cell clusters and combining them with single-cell sequencing technology, achieving high-resolution analysis of cell interactions within cell clusters and revealing the gene expression characteristics of cell interactions without the need for computational deconvolution, thereby reducing the complexity and errors of data processing.

[0013] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0015] Figure 1 Schematic diagram of the cell mass molecular labeling technology process provided for some examples of this application;

[0016] Figure 2 Schematic diagrams of oligonucleotide structures provided for some examples of this application;

[0017] Figure 3 Schematic diagram of the connection method of oligonucleotide, first molecular tag, second molecular tag and third molecular tag provided for some examples of this application;

[0018] Figure 4Schematic diagrams of cell labeling and cell interactions provided for some examples of this application; wherein a is a schematic diagram of the CCI-seq technical process; be is a box plot of indicators in each droplet after data filtering; b is the total number of combined indexes; c is the number of top 5 indexes; d is the ratio of the first and second ranked indexes; e is the ratio of the top 5 indexes;

[0019] Figure 5 Schematic diagram of cell counting results based on RNA annotation (y-axis, as true value) and species-specific Cell-ID annotation (x-axis) provided for some examples of this application;

[0020] Figure 6 Schematic diagram of the relative abundance results of cells assigned to cell clusters of specified sizes provided for some examples of this application;

[0021] Figure 7 Schematic diagram of the cell type composition results within the cell clusters provided for some examples of this application;

[0022] Figure 8 Schematic diagram of the main cell types and their spatial arrangement in the mouse kidney provided for some examples of this application;

[0023] Fig. 9 Schematic diagram of the visualization results of Uniform Manifold Approximation and Projection (UMAP) of single-cell RNA data of mouse kidney CCI-seq provided for some examples of this application;

[0024] Fig.10 Schematic diagram of the cell-cell interaction network of mouse kidney provided for some examples of this application; wherein a is a schematic diagram of the relative abundance results of cells assigned to cell clusters of specified sizes; b is a schematic diagram of the cell-cell interaction network results of mouse kidney;

[0025] Fig.11 Schematic diagram of the interaction strength between cell types provided for some examples of this application; wherein the top bar graph shows the number of cells of each cell type;

[0026] Fig.12 Schematic diagram of the enrichment and depletion results of the interaction strength between cell types provided for some examples of this application; wherein the top bar graph shows the number of cells of each cell type;

[0027] Fig.13Schematic diagram of mouse kidney single molecule fluorescence in situ hybridization (asmFISH) results provided for some examples of this application; wherein, a is an asmFISH image of the glomerulus, and DAPI counterstaining shows the cell-to-cell interaction between podocytes and monocytes; Nphs2 (yellow), podocytes; Lyz2 (red), monocytes; dotted lines indicate cell boundaries, and red arrows indicate the interaction between podocytes and monocytes; b is an asmFISH image of the kidney, and DAPI counterstaining shows the cell-to-cell interaction between proximal straight tubule cells (PSTs) and B cells. Atp11a (yellow), PSTs; Cd79a (red), B cells. The dotted lines indicate cell boundaries, and the red arrows indicate the interaction between PSTs and B cells;

[0028] Fig.14 Schematic diagram of the main cell types and their spatial arrangement in the mouse intestine provided for some examples of this application;

[0029] Fig.15 Schematic diagram of the cell-cell interaction network results of the mouse intestine provided for some examples of this application; cells (nodes) are colored by cell type (left) and experiment (right), and gray edges indicate detected cell-cell interactions; the dashed box highlights Lgr5 + Interactions within ISCs;

[0030] Fig.16 Schematic diagram of the relative abundance results of cells assigned to cell clusters of different sizes provided for some examples of this application;

[0031] Fig.17 Schematic diagram of the interaction strength results between different cell types provided for some examples of this application; wherein the top bar graph shows the number of cells of each cell type;

[0032] Fig.18 Schematic diagram of the enrichment and depletion results of the interaction strength between different cell types provided for some examples of this application; wherein the top bar graph shows the number of cells of each cell type;

[0033] Fig.19 Schematic diagram of the interaction results between TA cells and other cell types provided for some examples of this application; wherein a is the interacting cell type; b is the spatial localization result;

[0034] Fig. 20 Schematic diagram of mouse small intestine single molecule fluorescence in situ hybridization (asmFISH) results provided for some examples of this application; the left picture is an asmFISH image of intestinal crypts with DAPI counterstaining; Lgr5 (green): Lgr5 +Stem cells; Mki67 (yellow): proliferation marker; Sorbs2 (red): crypt bottom TA; Rbp7 (cyan): crypt top TA; dotted lines indicate crypt and cell boundaries; right: changes in fluorescence signal intensity from the crypt bottom to the crypt top area;

[0035] Fig.21 Schematic diagram of the interaction results between mouse intestinal TA cells and other cells provided for some examples of this application; wherein, a is the result of coloring the cells by marker scores, and the black arrow indicates the inferred spatial distribution of TA cells from the bottom to the top of the crypt; b is the result of coloring the cells by pseudo-time; c is the result of correlation analysis between the pseudo-time of TA cells and the marker scores; the dotted line indicates the fitted linear regression curve. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In the present application, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. In the description of the present application, unless otherwise specified, "multiple" refers to two or more than two.

[0038] For nucleotides, the term "identity" is used to describe or compare the degree of nucleotide similarity of two or more nucleotide sequences. The percentage of "sequence homology" between a first sequence and a second sequence can be calculated by dividing [the number of nucleotides in the first sequence that are identical to the nucleotides at the corresponding positions in the second sequence] minus [the total number of nucleotides in the first sequence] and then multiplying by [100%], where each deletion, insertion, substitution or addition of a nucleotide in the second nucleotide sequence - relative to the first nucleotide sequence - is considered to be a difference at a single nucleotide (position). Alternatively, the degree of sequence identity between two or more nucleotide sequences can be calculated using known computer algorithms for sequence alignment, such as NCBI Blast v2.0, using standard settings. Some other techniques, computer algorithms and setups for determining the degree of sequence identity are described, for example, in WO 04 / 037999, EP 0 967 284, EP 1 085 089, WO 00 / 55318, WO 00 / 78972, WO 98 / 49185 and GB 2357768-A.

[0039] It should be noted that the "combined molecular signature" described in this application is the same as the "combined index".

[0040] It should be noted that the "Cell-ID" described in this application is the most abundant or top-ranked combined index in the predetermined cell.

[0041] This application is based on the inventor's discovery and understanding of the following facts and problems:

[0042] Traditional single-cell sequencing technology usually needs to be performed in a cell suspension state, which means that the spatial location information of cells in the tissue and the interaction relationship between cells will inevitably be lost during the experiment. Although the interaction between cells can be inferred to a certain extent by performing bioinformatics analysis on single-cell transcriptome data, such as analyzing ligand-receptor relationships, these inferences are only indirect analyses based on gene expression and cannot directly capture the true cell interaction relationship at the single-cell level, nor can they determine the gene expression characteristics of interacting cells. These analyses are only interactions between cell populations, and cannot determine the specific interacting objects of individual cells.

[0043] The development of spatial transcriptomics technology has made up for the shortcomings of traditional single-cell sequencing to a certain extent. It attempts to obtain the gene expression information of cells while retaining the spatial information of cells, so as to more truly reflect the cell status in tissues. However, the current commercial spatial transcriptomics technology, such as the platform of 10x Genomics, is usually unable to achieve single-cell resolution research. This means that during the analysis process, a spatial point may contain signals from multiple cells, making it difficult to accurately identify the gene expression information of a single cell. Although technical processes such as Slide-seq, Seq-scope and Stereo-seq can achieve single-cell resolution and can simultaneously analyze spatial position and gene expression information, they face high technical barriers, complex operations, high application costs, and difficulty in accurately obtaining cell boundary information. In addition, these technologies have other limitations, such as limited gene detection, and some genes with low gene expression levels are difficult to detect; at the same time, due to the limitations of the technology itself, cell boundaries are difficult to determine, which affects the accurate analysis of cell interactions; in addition, these technologies often require specialized equipment, which further limits their large-scale application.

[0044] Cell interaction techniques based on proximity labeling, such as labeling cells near "bait" cells through enzymes or photocatalysis, followed by single-cell sequencing. Although this method can be used to identify other cells that interact with specific cell types, it requires the pre-selection of specific "bait" cell types, which limits the scope of the analysis and cannot generate large-scale cell-cell interaction networks unbiasedly among all cell types in the tissue, and it is difficult to discover new interactions. In addition, these methods are unable to obtain gene expression information of "bait" cells, making it difficult to study the impact of interactions on cell function. This method can only analyze preset cell types and is prone to missing important interaction information.

[0045] Cell cluster deconvolution methods, such as PIC-seq, ClumpSeq, and CIM-Seq, sequence partially dissociated cell clusters and then perform computational deconvolution in combination with single-cell sequencing data to infer the cell types in the clusters and further determine cell interactions. However, these methods rely on complex computational inferences and are prone to inaccurate deconvolution. Since the gene expression profiles of different cell types may have similarities, during the deconvolution process, it is easy to misjudge signals that do not belong to the same cell as the same cell, resulting in erroneous interaction inferences. At the same time, since this method does not directly measure the gene expression of single cells, it is difficult to obtain accurate gene expression profiles of interacting cells, which is not conducive to studying the gene expression characteristics of cell interactions.

[0046] In addition, spatial nuclear indexing technologies, such as Slide-tags, distinguish spatial locations by marking individual cell nuclei within tissue sections using spatial oligonucleotide molecular tags (or barcodes). However, these methods face the challenge of low cell nucleus recovery, resulting in sparse tissue sampling and the inability to capture data on the vast majority of cell-cell interactions in tissues. Moreover, these methods only detect RNA molecules in the cell nucleus and ignore the large amount of mRNA information in the cytoplasm, which limits their accurate analysis of cellular gene expression profiles.

[0047] There are other methods, such as tissue fragment sequencing, which can characterize the single-cell transcriptome in spatially different tissue microenvironments, but its analysis objects are limited to large tissue fragments (200-450μm), each of which contains tens of thousands of cells. It lacks the spatial resolution required for detailed interaction studies and cannot provide interaction information at the single-cell level.

[0048] To this end, the present application proposes a method for molecularly labeling cell aggregates and its application, which has the following beneficial technical effects:

[0049] Traditional scRNA-seq technology needs to be performed in a cell suspension state, which will lose the spatial position and interaction relationship of cells in the tissue. This application retains the physical contact information between cells by marking at the cell cluster level, so that the interaction between cells can be directly determined.

[0050] Commercial ST technologies (such as the 10x Genomics platform) are difficult to achieve single-cell resolution, while high-resolution ST technologies (such as Slide-seq, Seq-scope, Stereo-seq) are complex to operate, expensive, and difficult to accurately obtain cell boundary information. This application uses a combined index tagging method to achieve analysis of cell interactions at single-cell resolution, avoiding high costs and complex operations.

[0051] Proximity labeling technology requires the pre-selection of specific "bait" cell types, which limits the scope of analysis and cannot obtain gene expression information of interacting cells. This application does not require pre-selection of cell types, can capture the interactions of all cell types in the tissue unbiasedly, and simultaneously obtain the gene expression profiles of interacting cells.

[0052] Deconvolution methods that eliminate cell clumps (such as PIC-seq, ClumpSeq, CIM-Seq) rely on complex computational deconvolution and are prone to inaccurate interaction inferences. This application directly labels cell clumps to obtain gene expression information and direct interaction relationships for each cell without relying on computational inferences.

[0053] Spatial nuclear indexing technology faces the challenges of low cell nucleus recovery rate and sparse sampling, and only detects RNA molecules in the cell nucleus, ignoring a large amount of mRNA information in the cytoplasm. This application can capture the interaction information of most cells in the tissue and detect the RNA information of all cells by marking the cell membrane.

[0054] Fragment-seq is limited to analyzing large tissue fragments (200-450 μm) and lacks the spatial resolution required for detailed interaction studies. This application can directly identify cellular interactions at the single-cell level, providing more detailed interaction information.

[0055] Specifically, the technical solution of this application is as follows:

[0056] On the one hand, the present application proposes a method for molecular labeling of cell agglomerates, referring to Figure 1 , the method includes: performing combined labeling on the aforementioned cell clumps. By labeling the cell clumps, it is possible to effectively distinguish directly interacting cells and avoid the loss of cell interaction information during single cell separation. By limiting the number of cells in the cell clumps, the quality and spatial resolution of the sequencing data are improved while ensuring the true capture of cell interactions, so that subsequent data analysis can more accurately identify and resolve the interaction network between cells and avoid increased difficulty in resolution and information loss due to excessively large clumps.

[0057] In some examples of the present application, the aforementioned cell mass includes: 2 to 30 cells. It is understandable that the aforementioned cell mass may include: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 cells.

[0058] In some preferred examples of the present application, the aforementioned cell cluster includes: 2-20 cells, thereby ensuring that the cell interactions are truly captured while avoiding errors caused by excessive clustering.

[0059] In some examples of the present application, the aforementioned cell aggregates are pre-anchored with oligonucleotides.

[0060] In some examples of the present application, the aforementioned oligonucleotide is modified with a chemical molecule, and the aforementioned oligonucleotide is anchored on the cell membrane of the cell in the aforementioned cell cluster through the chemical molecule.

[0061] In some examples of this application, reference Figure 2The oligonucleotide comprises: an anchor sequence, a co-anchor sequence, an upper oligonucleotide sequence and a lower oligonucleotide sequence, wherein at least a portion of the 5' end of the anchor sequence is complementary to at least a portion of the co-anchor sequence, at least a portion of the 5' end of the lower oligonucleotide sequence is complementary to at least a portion of the 3' end of the anchor sequence, and at least a portion of the 3' end of the upper oligonucleotide sequence is complementary to at least a portion of the 3' end of the lower oligonucleotide sequence.

[0062] In some examples of the present application, the aforementioned chemical molecule is located at the 5' end of the aforementioned anchor sequence.

[0063] In some examples of the present application, the aforementioned chemical molecule is located at the 3' end of the aforementioned co-anchoring sequence.

[0064] In some examples of the present application, the aforementioned chemical molecule is optionally cholesterol or fatty acid.

[0065] It is understandable that anchoring the oligonucleotide to the cell membrane by chemical molecule modification is only an exemplary method, and those skilled in the art can also achieve the anchoring of the oligonucleotide by other methods, such as the interaction between biotin and avidin.

[0066] In some preferred examples of the present application, the aforementioned oligonucleotide is modified with cholesterol, and the aforementioned oligonucleotide is anchored to the cell membrane of the cells in the aforementioned cell cluster through cholesterol.

[0067] In some examples of the present application, the aforementioned anchor sequence has a nucleotide sequence as shown in SEQ ID NO: 1 or a nucleotide sequence having 90% similarity thereto.

[0068] GTAACGATCCAGCTGTCACTACACGTCTGAACTCCAGTCAC (SEQ ID NO: 1).

[0069] In some examples of the present application, the aforementioned co-anchor sequence has a nucleotide sequence as shown in SEQ ID NO: 2 or a nucleotide sequence having 90% similarity thereto.

[0070] AGTGACAGCTGGATCGTTAC (SEQ ID NO: 2).

[0071] In some examples of the present application, the aforementioned oligonucleotide sequence has a nucleotide sequence as shown in SEQ ID NO: 3 or a nucleotide sequence having 90% similarity thereto.

[0072] TGACTTGAGATCGGAAGAGC (SEQ ID NO: 3).

[0073] Optionally, the aforementioned oligonucleotide sequence has a nucleotide sequence as shown in SEQ ID NO: 4 or a nucleotide sequence having 90% similarity thereto.

[0074] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 4).

[0075] In some examples of the present application, the aforementioned combined labeling treatment is performed in the following manner: a predetermined plurality of cell clusters are respectively assigned to the first micropore group, and the first molecular label in each first micropore is subjected to the first labeling treatment with the cell cluster; the first labeling treatment products in the first micropore group are merged and respectively assigned to the second micropore group, and the second molecular label in each second micropore is subjected to the second labeling treatment product; the second labeling treatment products in the second micropore group are merged and respectively assigned to the third micropore group, and the third molecular label in each third micropore is subjected to the third labeling treatment product. Through multiple rounds of allocation-split-pooling method, each cell cluster has a unique combined label, so as to accurately determine the cell source.

[0076] The molecular labeling of the present application is achieved through three rounds of split-pool molecular labeling, relying on the split-pool strategy to give the sample a unique combination of molecular labels (combination index). This method gradually adds different molecular label combinations to each cell cluster through multiple rounds of physical partitioning and random mixing, so that each cell cluster has a unique combination index.

[0077] It is understood that washing is required to remove unbound oligonucleotides after each round of labeling. Washing unbound oligonucleotides can remove excess labeled molecules, reduce background noise, improve labeling specificity and efficiency, avoid cross-reactions, and thus ensure the accuracy and credibility of experimental results.

[0078] The present application does not specifically limit the first microwell group, the second microwell group and the third microwell group. For example, the first microwell group, the second microwell group and the third microwell group can be selected from PCR plates commonly used in biological experiments.

[0079] It can be understood that the aforementioned labeling process can be achieved by connecting the first molecular tag, the second molecular tag and the third molecular tag to the oligonucleotide.

[0080] In some examples of the present application, the first molecular tag and the second molecular tag are double-stranded nucleotide sequences, and both ends of the double-stranded nucleotide sequence have sticky ends; the third molecular tag is selected from single-stranded nucleotide sequences; wherein the oligonucleotide is connected to the first molecular tag via a sticky end, the second molecular tag is connected to the first molecular tag via a sticky end, and the third molecular tag is connected to the sticky end of the second molecular tag.

[0081] The aforementioned oligonucleotide, the first molecular tag, the second molecular tag, and the third molecular tag are connected in the following manner: Figure 3 shown.

[0082] In some examples of the present application, the aforementioned first molecular tag has a nucleotide sequence as shown in SEQ ID NO: 5 or having 90% similarity thereto (first molecular tag A chain).

[0083] In some examples of the present application, the aforementioned first molecular tag has a nucleotide sequence as shown in SEQ ID NO: 6 or a nucleotide sequence having 90% similarity thereto (first molecular tag B chain).

[0084] TN1N1N1N1N1N1N1N1N1N1N1N1N1N1N1N1TCAACAG (SEQ ID NO: 5).

[0085] CAAGTCAAN2N2N2N2N2N2N2N2N2N2N2N2N2N2N2N2 (SEQ ID NO: 6).

[0086] Among them, N1 and N2 are one of A, T, C, and G, and N1 and N2 are complementary pairs.

[0087] Based on the above sequence, molecular tags are added to the cell clusters in each microwell in the first microwell group. Since N1 and N2 are random sequences, the cell clusters in the same microwell will have the same first molecular tag sequence, while the cell clusters in different microwells will have different first molecular tag sequences.

[0088] In some examples of the present application, the aforementioned second molecular tag has a nucleotide sequence as shown in SEQ ID NO: 7 or having 90% similarity thereto (second molecular tag A chain).

[0089] In some examples of the present application, the aforementioned second molecular tag has a nucleotide sequence as shown in SEQ ID NO: 8 or having 90% similarity thereto (second molecular tag B chain).

[0090] TN3N3N3N3N3N3N3N3N3N3N3N3N3N3N3N3GTTCAGT (SEQ ID NO: 7).

[0091] AGTTGTCAN4N4N4N4N4N4N4N4N4N4N4N4N4N4N4N4 (SEQ ID NO: 8).

[0092] Among them, N3 and N4 are one of A, T, C, and G, and N3 and N4 are complementary pairs.

[0093] Based on the above sequence, molecular tags are added to the cell clusters in each microwell in the second microwell group. Since N1 and N2 are random sequences, the cell clusters in the same microwell will have the same second molecular tag sequence, while the cell clusters in different microwells will have different second molecular tag sequences.

[0094] In some examples of the present application, the aforementioned third molecular tag has a nucleotide sequence as shown in SEQ ID NO: 9 or a nucleotide sequence having 90% similarity thereto.

[0095] CAAGTCANNNNNNNNGCTTTAAGGCCGGTCCTAGCAA (SEQ ID NO: 9).

[0096] Wherein, N is one of A, T, C, and G.

[0097] Based on the above sequence, a molecular tag is added to the cell clusters in each microwell in the third microwell group. Since X1 and Y1 are random sequences, the cell clusters in the same microwell will have the same third molecular tag sequence, while the cell clusters in different microwells will have different third molecular tag sequences.

[0098] Through three rounds of split-pooling, each cell cluster has a unique combination label.

[0099] It should be noted that the 5' end of any sequence of the first molecular tag, the second molecular tag and the third molecular tag of the present application is phosphorylated to facilitate the ligation reaction.

[0100] In some examples of the present application, the first labeling treatment, the second labeling treatment, and the third labeling treatment are independently performed at room temperature and a rotation speed of 12 to 17 rpm for 3 to 8 minutes. The labeling treatment conditions are used to ensure efficient and accurate labeling of cell aggregates.

[0101] In some examples of the present application, the first labeling process, the second labeling process, and the third labeling process are respectively and independently sticky end connection. The connection product obtained by connecting the first molecular tag, the second molecular tag, and the third molecular tag through the labeling process is a combined index.

[0102] In this application, "Combinatorial Index" refers to a unique nucleotide sequence tag assigned to each cell cluster or single cell by combining multiple molecular tags during the molecular labeling and sequencing of cell clusters. This index is used to identify and track single cells from the same cell cluster in sequencing data, ensure the accuracy of data attribution, and improve the quality and reliability of sequencing data through denoising.

[0103] In some examples of the present application, the first labeling treatment, the second labeling treatment, and the third labeling treatment are independently performed in the presence of T4 DNA ligase.

[0104] In some examples of the present application, the aforementioned cell agglomerates are obtained by mechanically cutting or enzymatically digesting tissues.

[0105] It should be noted that the cell agglomerates of the present application can be tissues with a fixed spatial structure, such as kidney, small intestine and tumor tissue, and are also applicable to tissues without a fixed morphological structure, such as blood, lymph nodes, etc.

[0106] In some examples of the present application, the enzyme used for digestion treatment includes at least one of trypsin, dispase II, collagenase type IV, and deoxyribonuclease I.

[0107] In some examples of the present application, the aforementioned digestion treatment is carried out at 37° C. for 15 to 25 minutes. Optionally, it is 15 minutes, 16 minutes, 17 minutes, 18 minutes, 19 minutes, 20 minutes, 21 minutes, 22 minutes, 23 minutes, 24 minutes or 25 minutes. In some preferred examples of the present application, the aforementioned digestion treatment is carried out at 37° C. for 20 minutes.

[0108] For example, 10 mM EDTA-PBS is used to incubate the cell pellet on ice for 15 minutes, and the supernatant is collected after vigorous shaking with cold PBS. For example, for mouse kidney tissue, a 10 mM EDTA-PBS containing Dispase II, Collagenase IV, DNase I, and CaCl 2 Incubate at 37 °C for 20 min.

[0109] In another aspect, the present application proposes a sequencing method, which comprises: using any of the aforementioned methods to molecularly label cell aggregates; dissociating the molecularly labeled cell aggregates; and sequencing the single cells obtained after the dissociation. The sequencing method based on cell aggregate labeling can comprehensively and accurately analyze cell interactions and ensure the traceability of cell sources.

[0110] In some examples of the present application, the aforementioned dissociation treatment is performed under the action of trypsin. The aforementioned can be selected from TrypLE TM The conditions for the aforementioned dissociation treatment can be selected from shaking at 800 rpm for 5 minutes at 37°C.

[0111] It is understandable that in order to ensure that the cells used for sequencing are single cells, the cells after the dissociation treatment can be further filtered using a 35 μm cell filter.

[0112] In some examples of the present application, the aforementioned sequencing processing is performed using a single-cell sequencing platform, such as the 10xGenomics single-cell sequencing platform.

[0113] The raw sequencing data after sequencing can be processed using the CellRanger tool to generate a gene-cell matrix.

[0114] In some examples of the present application, the same combined index in the sequencing data is an indication that the cells are from the same cell cluster.

[0115] For the description of the combined index, please refer to the above and will not be described in detail here. The "same combined index" mentioned in this application means that in the sequencing data, the combined index sequences of different single cells are consistent, indicating that these single cells are derived from the same cell mass. Specifically, in the molecular labeling and sequencing process of the cell mass, each cell mass will be marked by multiple rounds of molecular tags so that each cell mass has a unique combined index. During sequencing, if the combined indexes of multiple single cells are exactly the same, it means that these single cells originally belonged to the same cell mass. This matching is used for cell tracing and single cell interaction data analysis to ensure the information integrity at the cell mass level and reduce the attribution errors caused by sequencing errors.

[0116] In some examples of the present application, the aforementioned method further includes: denoising the sequencing data. The denoising step is a conventional step in sequencing data processing, such as deduplication of the combined index data obtained by sequencing; extracting cell barcodes and removing adapter sequences, etc. Among them, deduplication can be performed by commonly used biological tools, such as seqkit (v.2.5.1); extracting cell barcodes can be achieved by the umi_tools (v1.1.4) tool, and removing adapter sequences can be achieved by the cutadapt (v.4.4) tool.

[0117] In some examples of the present application, the denoised data is subjected to quality control processing of the low-quality combined index to obtain a high-quality combined index. At least one of the following criteria is an indication of a high-quality combined index: 1) the number of indexes in valid droplets is higher than the upper quartile of the number of indexes in empty droplets; 2) the ratio of the first-level index (Rank 1) in valid droplets is higher than the upper quartile of the ratio of the first-level index of empty droplets; 3) the ratio of the number of first-level indexes to the number of second-level indexes (Rank 2) is greater than 2.

[0118] The aforementioned droplet refers to a tiny liquid sphere formed by microfluidics technology, which is usually used to encapsulate cells, molecules or other biological samples separately. Each droplet usually contains a single cell or a single molecule, which serves as an independent reaction unit for subsequent biological analysis. In single-cell RNA sequencing, genomics research and other high-throughput experiments, droplets serve as independent reaction spaces that can effectively isolate different samples and avoid cross-contamination, thereby achieving efficient and accurate analysis.

[0119] The aforementioned valid droplet is a droplet containing qualified gene expression data.

[0120] The aforementioned empty droplet is a droplet that does not contain qualified gene expression data.

[0121] The aforementioned first-level index is the most frequent index.

[0122] In some examples of the present application, for any two cell types, the interaction strength is calculated as the product of the cell numbers of different cell type pairs or the ratio of the number of combinations of the same cell type pairs to the total number of combinations within each cell cluster, and then these ratios are summarized across all clusters and normalized by the total cell number.

[0123] In some examples of this application, the interaction strength between any two cells is calculated as follows:

[0124]

[0125] Among them, Intensity AB represents the interaction strength between cell types A and B; N A and N B Respectively represent the total number of cell types A and B in all cell aggregates; n iA and n iB Respectively represent the total number of cell types A and B in cell cluster i; n i represents the total number of cells in the ith cell cluster; n A represents the total number of cells of cell type A; K represents the total number of cell aggregates.

[0126] The present application is illustrated below by way of examples, but this should not be construed as limiting the scope of the subject matter of the present application to the following examples. All technologies implemented based on the above content of the present application belong to the scope of the present application. The compounds or reagents used in the following examples can be purchased from commercial sources or prepared by conventional methods known to those skilled in the art; the experimental instruments used can be purchased from commercial sources.

[0127] In the following examples, the single-cell interaction sequencing method based on cell cluster combined index markers is abbreviated as CCI-seq.

[0128] Example 1: Cell type differentiation and cell interaction identification based on CCI-seq

[0129] 1. Purpose of the experiment

[0130] This example aims to verify the specificity of CCI-seq technology in cell labeling and its ability to accurately identify cell-cell interactions. By mixing cells from different species and using species-specific markers, it is possible to evaluate whether CCI-seq technology can distinguish different types of cells and correctly identify which cells have been in the same cell mass, thereby inferring the interactions between them.

[0131] 2. Experimental Materials

[0132] Human cell line: HEK293T cells;

[0133] Mouse cell line: NIH / 3T3 cells;

[0134] Culture medium: Dulbecco's Mdified Eagle's Medium (DMEM), containing 10% fetal bovine serum (FBS);

[0135] Cell detachment reagent: 0.25% Trypsin-EDTA;

[0136] Cell counter: C-Chip disposable blood cell counter;

[0137] Microplate: AggreWellTM400 6-well microplate;

[0138] CMO molecular tags: human-specific and mouse-specific CMO molecular tags;

[0139] Ligase: T4 DNA ligase;

[0140] 10x Genmics Kit: 10X Genmics 3'V3.1 Kit.

[0141] 3. Experimental Procedure

[0142] 3.1 Cell culture

[0143] HEK293T and NIH / 3T3 cells were cultured separately.

[0144] The cells were cultured at 37°C, 5% C 2 The cells were cultured under the same conditions until an appropriate cell density was reached.

[0145] The cells were washed with PBS, digested with trypsin and the digestion was terminated, and the cell suspension was collected.

[0146] Count cells using a hemocytometer.

[0147] 3.2 Three-dimensional spheroid culture

[0148] Cells were cultured using AggreWellTM400 microplates, with 42,000 cells seeded per well (approximately 6 cells per microwell);

[0149] At 37°C, 5% C 2 The cells were cultured in an incubator for 24 hours to form spheres of about 20 cells in size;

[0150] The cell pellets were collected and single cells were removed using a 10 μm filter.

[0151] 3.3 Cell labeling

[0152] Human and mouse cell spheres were collected separately and labeled with human- and mouse-specific CM molecular tags, respectively;

[0153] Anchoring the cell membrane: connecting the cholesterol-modified oligonucleotide (CM) to the cell membrane as an "anchor" for subsequent molecular tag connection;

[0154] Combined index labeling: Use three rounds of split-pool molecular label connection to label each cell cluster with a unique combined index;

[0155] First round: Disperse cell clumps into 96-well plates and add specific tags to each well;

[0156] Second round: After mixing, the cells are spread into a second 96-well plate and a second round of specific tags are added;

[0157] Round 3: After mixing the cells again, they are dispersed into a third 96-well plate and a third round of specific tags are added to achieve combined index labeling;

[0158] After three rounds of labeling, each cell cluster has a unique combined index.

[0159] 3.4 Cell mixing and dissociation

[0160] Mix the labeled human and mouse cell spheres;

[0161] The cell clumps were dissociated into single-cell suspensions using enzymatic dissociation.

[0162] 3.5 Single-cell sequencing

[0163] Single cells were sequenced using the 10x Genmics 3'kit;

[0164] The mRNA library and Cell-ID library were constructed separately.

[0165] 4. Data Analysis

[0166] 4.1 Data Preprocessing

[0167] CellRanger software was used to align sequencing reads to the human (hg38) and mouse (mm10) genomes;

[0168] Generate gene expression matrix and cell index data;

[0169] Low-quality cells and genes were filtered using Scanpy.

[0170] 4.2 Cell type differentiation

[0171] Distinguish human cells from mouse cells based on the UMI counts of human and mouse genomes;

[0172] UMI count thresholds were set to exclude low-quality cells and to label cells with both high human and mouse UMI counts as multisomes.

[0173] 4.3 Combined Index Analysis

[0174] Use the seqkit tool to remove duplicate index reads;

[0175] The umi_tls tool was used to extract cell molecular labels;

[0176] Use the cutadapt tool to remove adapter sequences and normalize index reads;

[0177] Determine the Cell-ID of each cell, i.e. the most abundant combination index;

[0178] Define the criteria for high-quality Cell-IDs, such as the number of indexes, the proportion of rank 1 indexes, and the ratio of the number of rank 1 and rank 2 indexes;

[0179] Filter out low-quality Cell-IDs and cell clumps;

[0180] Only cells that met the quality criteria were retained for subsequent analysis.

[0181] 4.4 Cell-cell interaction analysis

[0182] Cells with the same Cell-ID were defined as coming from the same cell cluster, i.e., interacting cells;

[0183] The cell types in each cell cluster were counted and the interactions between cell types were calculated.

[0184] 5. Experimental Results

[0185] This example evaluates CCI-seq using 3D spheroids of human HEK293T and mouse NIH / 3T3 cells cultured to a size of approximately 20 cells. These spheroids were labeled with human- and mouse-specific CMO barcodes ( Figure 4 a), then these cell clumps were mixed and subjected to three rounds of pooling and barcode ligation, followed by complete dissociation and single-cell RNA sequencing. Multiple signals detected in single-cell RNA sequencing were excluded, and 3,826 human cells and 3,643 mouse cells were obtained. In the single-cell RNA sequencing data obtained by CCI-seq analysis, the median number of RNA unique molecular identifiers (UMIs) for human cells and mouse cells was 12,062 and 11,066, respectively, and the number of genes per cell was 4,133 and 3,454, respectively; these values ​​are comparable to the results of standard droplet single-cell RNA sequencing experiments, indicating that the cell labeling process of CCI-seq does not adversely affect the acquisition of single-cell RNA sequencing data.

[0186] Multiple combined indexes were detected in each droplet. Among them, Cell-ID accounted for 29.3% of the total index signal, and its abundance was 43.8 times higher than the second-ranked index, indicating that the labeling process was highly specific. Further data filtering was performed to retain only cells with high-quality Cell-IDs and cell clumps of no more than 20 cells. Ultimately, 2,203 human cells and 1,839 mouse cells were obtained, which had high-quality Cell-IDs, accounting for 48.6% of all indexes, and their abundance was 86.3 times higher than the second-ranked index ( Figure 4 Of these, 99.7% of human cells and 93.8% of mouse cells had the correct species-specific Cell-ID tag ( Figure 5 ), reflecting the high specificity of cell markers.

[0187] Finally, approximately 95% (3,831 / 4,042) of the cells were distributed into clusters of more than two cells, providing information on cell-cell interactions ( Figure 6 Only about 0.3% (2 / 682) of the cell clumps contained human and mouse cells ( Figure 7 ), reflecting low cross-contamination between cell clumps. These results demonstrate that CCI-seq enables specific cell labeling and accurate identification of cell-cell interactions.

[0188] The above experimental results show that CCI-seq technology can achieve highly specific cell labeling and accurately identify cell-cell interactions; and verify the effectiveness of CCI-seq technology in distinguishing cell types and identifying cell interactions in complex cell mixtures.

[0189] Example 2: Application of CCI-seq technology in mouse kidney tissue

[0190] 1. Purpose of the experiment

[0191] This example aims to use CCI-seq technology to verify the ability of this technology to reveal the network of cell-cell interactions and discover new cell interactions in mouse kidneys with clear tissue structure and cell spatial arrangement characteristics.

[0192] 2. Experimental Materials

[0193] Experimental animals: 6-8 week old C57BL / 6J male mice;

[0194] Dissociation solution: PBS solution containing Dispase II, Collagenase IV and DNase I;

[0195] Red blood cell lysis buffer;

[0196] Cell strainer: 70μm cell strainer;

[0197] 10x Genomics kit: for single-cell sequencing.

[0198] 3. Experimental Procedure

[0199] (1) Kidney tissue preparation

[0200] Kidneys were removed from mice and placed in pre-chilled PBS;

[0201] Use scissors to cut the kidney tissue into pieces of approximately 1 mm in size.

[0202] (2) Tissue dissociation

[0203] Transfer the kidney tissue fragments to the EP tube containing the dissociation solution;

[0204] Grind the tissue repeatedly with scissors on ice for about 2 minutes;

[0205] Add more dissociation solution and incubate in a shaker at 37°C for 20 minutes;

[0206] The tissue suspension was filtered through a 70 μm cell strainer and repeatedly rinsed with PBS;

[0207] The filtrate was collected and centrifuged at 1000 rpm for 3 min at 4 °C;

[0208] Remove the supernatant, add red blood cell lysis buffer, and incubate at room temperature for 5 minutes to remove red blood cells.

[0209] (3) Cell clump collection

[0210] Resuspend the cell pellet using a wide-bore pipette tip for subsequent experiments.

[0211] (4) Cell labeling and sequencing

[0212] Using a similar process as in Example 1, the cell membrane was anchored and combined with index labeling;

[0213] Single-cell sequencing was performed using 10x Genomics 3'kit to construct mRNA and Cell-ID libraries.

[0214] 4. Data Analysis

[0215] (1) Data preprocessing

[0216] Sequencing reads were aligned to the mouse genome (mm10) using CellRanger software;

[0217] Generate gene expression matrix and cell index data;

[0218] Low-quality cells and genes were filtered using the Scanpy package.

[0219] (2) Cell type identification

[0220] Unsupervised clustering of 11,519 cells with high-quality gene expression profiles;

[0221] Based on the expression of known cell marker genes, 17 different cell types were identified, including:

[0222] Three types of collecting duct cells: collecting duct intercalated cells (CD-ICs), collecting duct principal cells (CD-PCs), and collecting duct transitional cells (CD-Trans);

[0223] Four glomerular cell types: podocytes, endothelial cells, mesangial cells, and pericytes;

[0224] There are five types of renal tubular cells: proximal straight tubule cells (PSTs), proximal convoluted tubule cells (PCTs), ascending limb cells of the loop of Henle (ALHs), descending limb cells of the loop of Henle (DLHs), and distal tubule cells (DCTs);

[0225] Three types of immune cells: T cells, B cells, and monocytes.

[0226] (3) Cell interaction analysis

[0227] The frequency of co-occurrence between any two cell types in cell clusters was calculated and normalized to obtain the "interaction strength";

[0228] The statistical significance of the strength of cell interactions was assessed using a permutation test;

[0229] Cells with the same Cell-ID were defined as being from the same cell cluster, ie, interacting cells.

[0230] (4) Spatial positioning verification

[0231] The spatial positional relationship of the new cell interactions discovered by CCI-seq was verified by asmFISH (amplification-based single molecule fluorescent in situhybridization) technology.

[0232] 5. Experimental Results

[0233] This example uses CCI-seq to analyze mouse kidney, which has a well-defined tissue anatomy and known spatial arrangement of cell types ( Figure 8 After data quality control and filtering, 11,519 high-quality gene expression profiles were obtained from three mouse kidney biological replicates ( Fig. 9 ) and Cell-ID data. Of these cells, 49.8% (5,731 / 11,519) were from a single cell cluster and thus had information on physical cell-cell interactions ( Fig.10 a and b). Unsupervised clustering of the gene expression profiles of 11,519 cells identified 17 distinct cell types based on the expression of classical marker genes36,37 Fig. 9 These cell types include three collecting duct cell types, four glomerular cell types, five tubular cell types, and three immune cell types ( Fig. 9 ).

[0234] This example further quantifies cell-cell "interaction strength," which reflects how often any two cell types appear together in a cell clump, normalized to account for their overall abundance and the total number of possible pairings. Using this approach, cell-cell interactions consistent with known anatomical structures in the kidney were identified. For example, interactions between collecting duct intercalated cells (CD-ICs), collecting duct principal cells (CD-PCs), and collecting duct transition cells (CD-Trans) were detected ( Fig.11 ), which are known to regulate ion transport and influence urine concentration and pH in the collecting ducts36,38. Interactions between podocytes, endothelial cells, and mesangial cells have also been detected ( Fig.11 ), these cells are thought to contribute to the structural integrity and functional capacity of the glomerulus.

[0235] This example also calculates the statistical significance of the interaction strength ( Fig.12 ), and observed that cell-cell interactions in the collecting ducts or glomeruli were highly enriched. In particular, several enriched interactions between immune cells and non-immune cells were also found, including previously unknown interactions between monocytes and podocytes, and interactions between B cells and proximal rectal tubule cells (PSTs) ( Fig.12 By using single-molecule fluorescence in situ hybridization (asmFISH) experiments, we indeed observed the direct proximity of these cell types in vivo ( Fig.13 a and b).

[0236] The above experimental results show that CCI-seq technology successfully detected known cell interactions in mouse kidneys and discovered new cell interactions; this study confirmed that CCI-seq technology can be used to reveal cell interaction networks in complex tissues, providing a new perspective for studying kidney function and disease; the experimental results show that CCI-seq technology can be used to study the microenvironment and intercellular communication of the kidney.

[0237] Overall, the application in mouse kidneys further validates that CCI-seq technology can accurately reveal cell-to-cell interactions in complex tissues and discover new, physiologically significant cell interactions. Through combined indexing of cell clusters and single-cell sequencing, this technology can provide high-resolution cell interaction information.

[0238] Example 3: Application of CCI-seq technology in mouse small intestine tissue

[0239] 1. Experimental purpose: This example aims to use CCI-seq technology to study the mouse small intestine with a characteristic crypt-villus structure, verify the ability of this technology to analyze cell-cell interactions in complex tissues, and reveal new cell subtypes and their spatial distribution.

[0240] 2. Experimental Materials

[0241] Experimental animals: C57BL / 6J male mice (6-8 weeks old);

[0242] Cell dissociation reagents: Dispase II, Collagenase IV, DNase I and CaCl 2 Mixed liquid;

[0243] Cell strainer: 70μm cell strainer;

[0244] 10x Genomics kit: for single-cell sequencing.

[0245] 3. Experimental Procedure

[0246] 3.1 Small Intestine Tissue Preparation

[0247] The small intestine was removed from the mouse and placed in pre-chilled PBS;

[0248] Use scissors to cut the small intestinal tissue into small pieces.

[0249] 3.2 Tissue dissociation

[0250] Transfer the small intestinal tissue fragments into the EP tube containing the dissociation solution;

[0251] Grind the tissue repeatedly on ice using scissors for about 2 minutes;

[0252] Add more dissociation solution and incubate at 37°C in a shaker for 20 min;

[0253] The tissue suspension was filtered through a 70 μm cell strainer and repeatedly rinsed with PBS;

[0254] The filtrate was collected and centrifuged at 1000 rpm for 3 min at 4 °C;

[0255] Remove the supernatant and resuspend the cell pellet using a wide-mouth pipette tip for subsequent experiments.

[0256] 3.3 Cell labeling and sequencing

[0257] Using a similar process as for the human-mouse hybrid experiment, the cell membrane was anchored and combined with index labeling;

[0258] Single-cell sequencing was performed using 10x Genomics 3'kit to construct mRNA and Cell-ID libraries.

[0259] 4. Data Analysis

[0260] 4.1 Data Preprocessing

[0261] Sequencing reads were aligned to the mouse genome (mm10) using CellRanger software;

[0262] Generate gene expression matrix and cell index data;

[0263] Use Scanpy package to filter low-quality cells and genes;

[0264] SCALEX was used to correct the differences between different batches of data.

[0265] 4.2 Cell type identification

[0266] Unsupervised clustering of 6,412 cells with high-quality gene expression profiles;

[0267] Based on the expression of known cell marker genes, eight different cell types were identified, including: Lgr5+ intestinal stem cells (ISCs), transitional proliferating cells (TA), Paneth cells, goblet cells, intestinal epithelial cells, enteroendocrine cells (EECs), tuft cells, and B cells.

[0268] 4.3 Cell interaction analysis

[0269] The frequency of co-occurrence between any two cell types in cell clusters was calculated and normalized to obtain the "interaction strength";

[0270] The statistical significance of the strength of cell interactions was assessed using a permutation test;

[0271] Cells with the same Cell-ID were defined as being from the same cell cluster, ie, interacting cells.

[0272] 4.4 Analysis of spatial subtypes of intestinal epithelial cells and goblet cells

[0273] Calculate the “landmark score” of enterocytes and goblet cells to determine their spatial distribution along the crypt-villus axis;

[0274] According to the landmark score, intestinal epithelial cells were divided into three spatial subtypes: villus base, villus middle and villus apex, and goblet cells were divided into two subtypes: crypt goblet cells and villus goblet cells.

[0275] 4.5 Analysis of spatial subtypes of transitional proliferating cells (TA)

[0276] According to the interacting cell types of TA cells, TA cells are divided into two spatial subtypes: crypt top TA cells and crypt bottom TA cells.

[0277] 4.6 Spatial Positioning Verification

[0278] The spatial distribution and interactions of cells discovered by CCI-seq were verified by asmFISH technology.

[0279] 5. Experimental results:

[0280] To further evaluate CCI-seq, this example further analyzed the mouse small intestine, which formed a clear crypt-villus structure ( Fig.14 In brief, crypts contain Lgr5+ intestinal stem cells (ISCs), their progeny (transition amplifying cells (TA)), and Paneth cells. Intestinal epithelial cells are mainly distributed in the villi, while other secretory cells, including goblet cells (mucus secretion), enteroendocrine cells (EECs; hormone secretion), and hair cells (chemosensation and immune response), are distributed along the crypt-villus axis ( Fig.14 ).

[0281] In this example, 6,412 high-quality gene expression profiles ( Fig.15 ) and Cell-ID data for single cells. Notably, 56.6% (3,631 / 6,412) of the cells had physical interactions ( Fig.16 ). By correcting for batch effects using SCALEX45 and performing unsupervised clustering of the gene expression profiles of these cells, eight different cell types were identified based on the expression of classical marker genes. These cell types include Lgr5+ISCs, TA cells, Paneth cells, goblet cells, intestinal epithelial cells, EECs, hair cells, and B cells ( Fig.15 ).

[0282] By quantifying the strength and preference of interactions between all cell types, we observed preferential interactions between different spatial subtypes of enterocytes (e.g., enterocytes at the villus tip interact with enterocytes at the villus tip and mid-villus, whereas enterocytes at the villus base interact with enterocytes at the villus base and mid-villus) ( Fig.17 , Fig.18 ). Similarly, villus goblet cells prefer to interact with enterocytes located in the villi, while crypt goblet cells prefer to interact with TA cells located in the crypts ( Fig.17 , Fig.18 In addition, a well-known interaction between Lgr5+ISCs and Paneth cells was identified ( Fig.17 , Fig.18 ). These findings support the ability of CCI-seq to detect physiologically relevant cell-cell physical interactions in tissues.

[0283] TA cells are derived from Lgr5+ intestinal stem cells (ISCs) and are characterized by rapid cell division. TA cells are essential for ensuring the continuous and efficient renewal and regeneration of the intestinal epithelium. It has been established that TA cells are located throughout the crypts and their extensions can reach the villi ( Fig.14 ). However, landmark scoring in previous studies did not include TA cells. Therefore, it is unclear whether TA cells have specific spatial subtypes with different functions. TA cells frequently interact with enterocytes at the base of the villi and cells in the crypts, such as Lgr5+ intestinal stem cells, Paneth cells, and crypt goblet cells ( Fig.17 ). Therefore, TA cells were divided into two spatial subtypes: crypt-apical TA and crypt-basal TA, based on their spatial localization with respect to interacting cells ( Fig.19 a and b).

[0284] Subsequently, we examined the differentially expressed genes (DEGs) of these two TA subtypes and showed that crypt-bottom TAs expressed high levels of Lgr5+ intestinal stem cell marker genes, including Sorbs2 and Olfm4. In contrast, crypt-apical TAs showed high expression of genes associated with mature intestinal epithelial cells, such as Rbp7 and Smim24. Using asmFISH, we verified that Sorbs2-expressing TA cells were confined to the crypt base, while Rbp7-expressing TA cells were confined to the crypt apex, confirming that CCI-seq accurately identified these spatially distinct TA subtypes ( Fig. 20 ).

[0285] Following the method used to distinguish spatial subtypes of intestinal epithelial cells, the top 10 DEGs of crypt-bottom TA and crypt-apical TA were defined as spatial landmark genes, and landmark scores for individual TA cells were calculated. These scores were highly correlated with the differentiation pseudo-temporal scores of each TA cell inferred from single-cell trajectory analysis (Pearson correlation coefficient = 0.74, Fig.21 ac). Finally, gene ontology (GO) enrichment analysis was performed, and the results showed that DEGs of TA cells at the bottom of the crypt were enriched in terms related to the Wnt signaling pathway, indicating that Wnt ligands at the base of the crypt have an effect on them. In contrast, TA cells at the top of the crypt showed higher metabolic activity and energy production, reflecting that they are similar to mature intestinal epithelial cells with active ATP absorption and metabolic processes.

[0286] In summary, the application in the mouse small intestine further verifies that CCI-seq technology can accurately reveal cell-to-cell interactions in tissues with complex structures and discover new, physiologically significant cell subtypes. Through combined index labeling of cell clusters and single-cell sequencing, this technology can provide high-resolution cell interaction information and spatial distribution information of cell subtypes.

[0287] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0288] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A method for molecular labeling of cell aggregates, characterized in that: include: The cell aggregates are subjected to combined labeling treatment.

2. The method according to claim 1, characterized in that The cell mass includes: 2-30 cells; Preferably, the cell mass comprises: 2-20 cells.

3. The method according to claim 2, characterized in that The cell pellets are pre-anchored with oligonucleotides.

4. The method according to claim 3, characterized in that The oligonucleotide is modified with a chemical molecule, and the oligonucleotide is anchored on the cell membrane of the cell in the cell cluster through the chemical molecule; Optionally, the oligonucleotide comprises: an anchor sequence, a co-anchor sequence, an upper oligonucleotide sequence and a lower oligonucleotide sequence, wherein at least a portion of the sequence at the 5' end of the anchor sequence is complementary to at least a portion of the sequence at the co-anchor sequence, at least a portion of the sequence at the 5' end of the lower oligonucleotide sequence is complementary to at least a portion of the sequence at the 3' end of the anchor sequence, and at least a portion of the sequence at the 3' end of the upper oligonucleotide sequence is complementary to at least a portion of the sequence at the 3' end of the lower oligonucleotide sequence; Optionally, the chemical molecule is located at the 5' end of the anchor sequence; Optionally, the chemical molecule is located at the 3' end of the co-anchor sequence; Optionally, the chemical molecule comprises: cholesterol or fatty acid.

5. The method according to claim 4, characterized in that The anchor sequence has a nucleotide sequence as shown in SEQ ID NO: 1 or a nucleotide sequence having 90% similarity thereto; Optionally, the co-anchor sequence has a nucleotide sequence as shown in SEQ ID NO: 2 or having 90% similarity thereto; Optionally, the oligonucleotide sequence has a nucleotide sequence as shown in SEQ ID NO: 3 or a nucleotide sequence having 90% similarity thereto; Optionally, the oligonucleotide sequence has a nucleotide sequence as shown in SEQ ID NO: 4 or a nucleotide sequence having 90% similarity thereto.

6. The method according to any one of claims 1 to 5, characterized in that: The combined label marking process is performed in the following manner: Distributing a predetermined plurality of cell clusters to first microwell groups respectively, and performing a first labeling process on the first molecular tag and the cell cluster in each first microwell; The first labeling products in the first microwell group are combined and respectively distributed to the second microwell group, and the second molecular tag in each second microwell and the first labeling products are subjected to a second labeling treatment; The second labeling products in the second microwell group are combined and respectively distributed to the third microwell group, and the third molecular tag in each third microwell and the second labeling products are subjected to a third labeling treatment.

7. The method according to claim 6, characterized in that The first molecular tag and the second molecular tag are double-stranded nucleotide sequences, both ends of the double-stranded nucleotide sequences have sticky ends; the third molecular tag is selected from single-stranded nucleotide sequences; The oligonucleotide is connected to the first molecular tag via a sticky end, the second molecular tag is connected to the first molecular tag via a sticky end, and the third molecular tag is connected to the sticky end of the second molecular tag.

8. The method according to claim 7, characterized in that The first molecular tag has a nucleotide sequence as shown in SEQ ID NO: 5 or a nucleotide sequence having 90% similarity thereto; The first molecular tag has a nucleotide sequence as shown in SEQ ID NO: 6 or a nucleotide sequence having 90% similarity thereto; Optionally, the second molecular tag has a nucleotide sequence as shown in SEQ ID NO: 7 or having 90% similarity thereto; The second molecular tag has a nucleotide sequence as shown in SEQ ID NO: 8 or a nucleotide sequence having 90% similarity thereto; Optionally, the third molecular tag has a nucleotide sequence as shown in SEQ ID NO: 9 or a nucleotide sequence having 90% similarity thereto.

9. The method according to claim 6, characterized in that The first labeling treatment, the second labeling treatment, and the third labeling treatment are independently performed at 16° C.-37° C. and a rotation speed of 12-17 rpm for 3-8 minutes; Optionally, the first labeling treatment, the second labeling treatment, and the third labeling treatment are independently connected by sticky ends; Optionally, the first labeling treatment, the second labeling treatment, and the third labeling treatment are each independently performed in the presence of T4 DNA ligase.

10. The method according to claim 1 or 2, characterized in that: The cell mass is obtained by mechanically cutting or enzymatically digesting the tissue; Optionally, the enzyme comprises at least one of trypsin, dispase II, collagenase type IV, and deoxyribonuclease I; Optionally, the digestion treatment is carried out at 37° C. for 15 to 25 min, preferably 20 min.

11. A sequencing method, characterized in that: include: The method according to any one of claims 1 to 10 is used to perform molecular labeling on the cell aggregates; Dissociate the cell aggregates that have been treated with molecular markers; The single cells obtained after dissociation are sequenced.

12. The sequencing method according to claim 11, characterized in that The dissociation treatment is carried out under the action of trypsin.

13. The sequencing method according to claim 11, characterized in that: The sequencing process is performed using a single-cell sequencing platform.

14. The sequencing method according to claim 13, characterized in that: The same combined index in the sequencing data indicates that the cells are from the same cell cluster.

15. The sequencing method according to claim 14, characterized in that: Further comprising performing quality control processing of the low-quality combined index on the sequencing processed data; optionally, the combined index having at least one of the following criteria is indicative of a high-quality combined index: 1) The number of indexes in valid droplets is higher than the upper quartile of the number of indexes in empty droplets; 2) the first-level index ratio of valid droplets is higher than the upper quartile of the first-level index ratio of empty droplets; 3) The ratio of the number of first-level indexes to the number of second-level indexes is greater than 2.

Citation Information

Patent Citations

  • GABA B receptors

    GB2357768A

  • Lepidopteran GABA-gated chloride channels

    WO1998049185A1

  • Methods and reagents for modulating cholesterol levels

    WO2000055318A2

  • Regulation with binding cassette transporter protein abc1

    WO2000078972A2