Marker combination, detection reagent and cell definition method

By combining biomarkers and detection reagents on the Xenium platform, the challenges of cell segmentation and type definition in cardiac spatial transcriptomics have been solved, enabling precise identification and subtype definition of cardiomyocytes and improving the reliability of cardiac research.

CN121931233APending Publication Date: 2026-04-28SHANGHAI CHILDRENS MEDICAL CENT AFFILIATED TO SHANGHAI JIAOTONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511828081.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing spatial transcriptomics techniques struggle to achieve precise cell segmentation and cell type definition in cardiac tissue, particularly exhibiting errors in subcellular resolution and cell type identification of cardiomyocytes.

Method used

A combination of biomarkers, including Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Myom2, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Myh7, and Corin genes, combined with detection reagents and the Xenium platform, enables subcellular resolution transcript coordinate localization and cell type definition through fluorescent probe hybridization and rolling circle amplification.

Benefits of technology

It enables accurate identification and subtype definition of cell types in cardiac tissue, providing a reliable basis for cardiac research and improving the accuracy of cell segmentation and the sensitivity and specificity of cell type definition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121931233A_ABST
    Figure CN121931233A_ABST
Patent Text Reader

Abstract

The invention discloses a marker combination, a detection reagent and a cell definition method. The marker combination comprises one or more genes selected from the following genes: Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Myom2, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Myh7 and Corin. The invention further discloses a method for detecting the biomarker combination. The marker combination, the detection reagent and the cell definition method provided by the invention can effectively and sensitively identify the cell type and the cell subtype of the heart based on Xenium, realize accurate cell segmentation and cell type definition in heart space transcriptomics research, provide a reliable basis for heart research, and have wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to a combination of biomarkers, detection reagents, and a method for cell segmentation and definition based on cardiac spatial transcriptomics. Background Technology

[0002] Spatial transcriptomics represents a major breakthrough following single-cell transcriptomics (scRNA-seq), aiming to address the core issue of spatial information loss in traditional transcriptomics analysis. The function of biological tissues is highly dependent on the spatial arrangement of cells (e.g., the conduction system of the heart, the tumor microenvironment), but single-cell techniques, requiring tissue dissection, cannot preserve positional information. In 2016, Joakim Lundeberg's team first proposed spatial transcriptomics technology for in-situ RNA capture, which was named Technology of the Year by *Nature Methods* in 2020, marking its emergence as a key tool in life science research. 1 .

[0003] Based on the principles, spatial transcriptomics technology is divided into two categories. 2 :

[0004] 1. Imaging-based technologies (such as MERFISH, seqFISH+)

[0005] Principle: In situ imaging of RNA via multiple rounds of fluorescence in situ hybridization (FISH) achieves subcellular resolution (~0.5 μm). Limitations: Low throughput (typically detecting hundreds to thousands of genes), requires pre-defined targets, and long imaging time (hours to days).

[0006] 2. Sequencing-based technologies (such as Visium, Slide-seq)

[0007] Principle: mRNA is captured using spatial barcode probes, combined with NGS sequencing, supporting whole transcriptome analysis. Limitations: Limited resolution (Visium dot diameter 55 μm, covering multiple cells), making it difficult to accurately locate single cells.

[0008] Xenium, an in-situ spatial analysis platform from 10x Genomics, combines the advantages of imaging and targeted sequencing, resolving the core contradiction of existing technologies—the balance between high resolution and high throughput. Utilizing the multi-round padlock probe hybridization (FISH) principle, it achieves subcellular resolution and targeted detection, and can simultaneously detect 300–500 genes (expandable to 5000+), covering key pathway genes. 3 .

[0009] Currently, the core challenge of spatial transcriptomics lies in accurately defining cell boundaries (i.e., "cell definition"), which is fundamental to understanding cell function, interactions, and the tissue microenvironment. Other technologies (such as Visium and Slide-seq) struggle to achieve single-cell segmentation due to insufficient resolution or signal crosstalk, while Xenium technology offers a solution to this challenge with its subcellular resolution and high throughput.

[0010] The heart is a highly spatially structured organ, and its function depends on precise cell localization and interactions. Therefore, defining cardiomyocyte cell types is particularly important for cardiac spatial transcriptomics research. Traditional spatial transcriptomics, due to its inability to segment cells, cannot accurately define cell types, making it difficult to apply to the heart. For example, Purkinje fibers (10-20 μm in diameter) are masked by mixed sites in Visium, and in MERFISH, they are misidentified as fibroblasts due to the lack of membrane staining. Furthermore, mature cardiomyocytes often exhibit a multinucleated state, which can easily lead to errors when treated as multicellular cells. Accurately segmenting and identifying different cardiomyocyte types has become a key focus and challenge in this field. 4 .

[0011] Applying Xenium technology to cardiac spatial transcriptomics research holds promise for achieving precise cell segmentation and cell type definition, providing a reliable foundation for cardiac research. However, current research is limited, lacking specific cell segmentation and cell definition targets and protocols applicable to the heart.

[0012] References

[0013] 1. Williams CG, Lee HJ, Asatsuma T, Vento-Tormo R, Haque A. Anintroduction to spatial transcriptomics for biomedical research. Genome Med.2022 Jun 27;14(1):68. doi: 10.1186 / s13073-022-01075-1. PMID: 35761361; PMCID:PMC9238181.

[0014] 2. Cheng M, Jiang Y, Xu J, Mentis AA, Wang S, Zheng H, Sahu SK, LiuL, Xu X. Spatially resolved transcriptomics: a comprehensive review of their technological advances, applications, and challenges. J Genet Genomics. 2023 Sep;50(9):625-640. doi: 10.1016 / j.jgg.2023.03.011. Epub 2023 Mar 27. PMID: 36990426.

[0015] 3. Marco Salas S, Kuemmerle LB, Mattsson-Langseth C, et al. Optimizing Xenium In Situ data utility by quality assessment and best-practice analysis workflows. Nat Methods. Published online March 13, 2025. doi:10.1038 / s41592-025-02617-2

[0016] 4. Nguyen Q, Tung LW, Lin B, Sivakumar R, Sar F, Singhera G, Wang Y, Parker J, Le Bihan S, Singh A, M V Rossi F, Collins C, Bashir J, Laksman Z. Spatial Transcriptomics in Human Cardiac Tissue. Int J Mol Sci. 2025 Jan 24;26(3):995. doi: 10.3390 / ijms26030995. PMID: 39940764; PMCID: PMC11817049.。 Summary of the Invention

[0017] To address the lack of existing technologies for cell segmentation and target definition based on the Xenium platform for cardiac applications, this invention provides a biomarker combination, a detection reagent, and a cell definition method. The biomarker combination provided by this invention can serve as a target for cell segmentation and cell definition in cardiac applications. The detection reagent enables effective and specific detection of the target. Furthermore, a method for defining the segmented cells has been developed based on the biomarker combination and / or the detection reagent. The biomarker combination, detection reagent, and cell definition method provided by this invention enable effective and sensitive identification of cardiac cell types and subtypes, achieving precise cell segmentation and cell type definition in cardiac spatial transcriptomics research based on Xenium technology, thus providing a reliable foundation for cardiac research.

[0018] The present invention solves the above-mentioned technical problems through the following technical means:

[0019] A first aspect of the present invention provides a biomarker combination comprising one or more genes selected from the following: Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Myom2, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Myh7, and Corin.

[0020] In some embodiments of the present invention, the combination of markers includes one or more subsets of markers selected from the following:

[0021] (a) Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn and Myom2;

[0022] (b) Aurkb, Pole, Ect2, and Ccna2;

[0023] (c) Nppa, Nppb, and Xirp2;

[0024] (d)Myh7; and,

[0025] (e) Corin;

[0026] In some preferred embodiments of the present invention, the marker combination includes Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Myom2, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Myh7, and Corin.

[0027] A second aspect of the present invention provides a detection reagent for detecting the expression level of a combination of biomarkers as described in the first aspect.

[0028] In some embodiments of the present invention, the detection reagent comprises primers, probes and / or antibodies; and / or, the expression level is a protein expression level and / or an mRNA transcription level.

[0029] In some preferred embodiments of the present invention, the probe is a circularizable DNA probe, and each probe contains a specific sequence at both ends that is complementary to the mRNA of the target marker.

[0030] In some preferred embodiments of the present invention, the detection reagent comprises one or more of the following probes:

[0031]

[0032]

[0033]

[0034]

[0035] The ENSEMBAL number is based on the genome sequence of the Norwegian rat (Rattus norvegicus), with the genome assembly version number mRatBN7.2.

[0036] In some embodiments of the present invention, the detection reagent comprises all of the following probes:

[0037]

[0038]

[0039]

[0040] ;

[0041] The ENSEMBAL number is based on the genome sequence of the Norwegian rat (Rattus norvegicus), with the genome assembly version number mRatBN7.2.

[0042] A third aspect of the present invention provides a method for defining cardiac cells, comprising the step of defining cardiac cells at the expression levels of a combination of markers as described in the first aspect.

[0043] In some embodiments of the present invention, the method is preferably based on the Xemium platform.

[0044] In some preferred embodiments of the present invention, the method includes:

[0045] S1: Perform cell segmentation on the test sample on the Xemium platform to obtain a single-cell mask containing the cytoplasmic boundary;

[0046] S2: Summarize transcripts that fall into the same mask and obtain the expression levels of each marker in the marker combination as described in the first aspect;

[0047] S3: Based on the expression level of each marker in the marker combination, define the cells in the sample to be tested to determine the cell type and / or cell state.

[0048] In some preferred embodiments of the present invention, in S3, the cell type is determined by the expression level of a marker, the marker being as defined in (a) of the marker combination described in the first aspect; the cell type being epicardial cells, smooth muscle cells, pericytes, immune cells, endothelial cells, fibroblasts, and / or cardiomyocytes.

[0049] In some preferred embodiments of the present invention, in S3, the cell state is determined by the expression level of a marker, the cell state being a proliferating cell or a cardiomyocyte subclass, and the marker being (b)-(e) selected from the marker combinations described in the first aspect; the cardiomyocyte subclass is preferably hypertrophic cardiomyocytes, immature cardiomyocytes, and / or mature cardiomyocytes.

[0050] In some preferred embodiments of the present invention, in step S3, when determining the cell type, the expression level of the marker is compared with a cell type definition standard, wherein the cell type definition standard is:

[0051] First layer: If the cell expresses Upk3b≥3, it is identified as an epicardial cell; otherwise, proceed to the next layer.

[0052] Second layer: If the cell expresses Myh11 ≥ 4, it is identified as a smooth muscle cell; otherwise, proceed to the next layer.

[0053] Third layer: If the cell expression test shows Kcnj8≥4, it is determined to be a pericyte; otherwise, proceed to the next layer.

[0054] Fourth layer: If the cell expresses Ptprc≥4, it is identified as an immune cell; otherwise, it proceeds to the next layer.

[0055] Fifth layer: If cells express Cdh5 and / or Dcn ≥ 5, they are determined to be either endothelial cells or fibroblasts, where CdhDH5 (GeneE) > Dcn (GeneF) indicates an endothelial cell; otherwise, they are fibroblasts. Otherwise, proceed to the next layer; and,

[0056] Sixth layer: If Myom2(GeneG)≥3, then it is a cardiomyocyte; otherwise, it is an undefined cell.

[0057] In some preferred embodiments of the present invention, in step S3, when determining the cell state, the expression level of the marker is compared with a cell state definition standard, wherein the cell state definition standard is:

[0058] If cells express Aurkb + Pole + Ect2 + Ccna2 > 5, they are classified as proliferating cells; and / or,

[0059] If the cell is defined as a cardiomyocyte, and the cell is determined to be a cardiomyocyte of the corresponding cell state when it meets the following conditions: when expressing Nppa + Nppb + Xirp2>20, it is determined to be a hypertrophic cardiomyocyte; or, if the cell expresses Myh7>20, it is determined to be an immature cardiomyocyte; or, if the cell expresses Corin>5, it is determined to be a mature cardiomyocyte.

[0060] In some preferred embodiments of the present invention, in S1, the cell segmentation is performed using Xeniumv2.0 software to obtain the single-cell mask.

[0061] In some embodiments of the present invention, before S1, in situ hybridization is performed on the test sample using probes defined in the detection reagents described in the two aspects, so that each probe is specifically linked to the mRNA molecule of the target marker to form a loop, and then rolling circle amplification is performed to obtain a gene-specific barcode copy. After multiple rounds of fluorescent probe hybridization, imaging and elution, optical barcode decoding is performed to obtain transcript coordinates and gene identity at subcellular resolution.

[0062] In some embodiments of the present invention, step S1 includes cell nucleus identification and staining of the sample to be tested, and image acquisition for cell segmentation.

[0063] In some preferred embodiments of the present invention, DAPI is used as a nuclear marker in cell nuclear identification and staining, and / or ATP1A1 and / or E-cadherin are used as membrane markers, and / or 18S ribosomal RNA is used as a cytoplasmic marker, and / or αSMA and / or vimentin are used as cell phenotypic markers.

[0064] A fourth aspect of the present invention provides a system for defining myocardial tissue cells in the Xemium platform, the system comprising the following modules:

[0065] An input module is used to input single-cell data of the sample to be tested, the single-cell data including expression level detection values ​​of the combination of markers as described in the first aspect;

[0066] An analysis module is used to analyze the test sample data to obtain analysis results; wherein, when the single-cell data meets the judgment condition, the analysis result is output as "satisfied"; when the test sample data does not meet the judgment condition, the analysis result is output as "not satisfied"; and / or,

[0067] The judgment module determines the cell type and / or cell state based on the single-cell data and outputs the judgment result. When the analysis result is "satisfied", the judgment result is the corresponding cell type and / or cell state. When the analysis result is "not satisfied", the analysis module performs the next level of judgment until the output analysis result is "satisfied".

[0068] In some embodiments of the present invention, the determination criteria are the cell type definition criteria and / or the cell state definition criteria as defined in the method described in the third aspect.

[0069] The fifth aspect of the present invention provides a readable medium storing a program that, when executed by a processor, enables the system functions as described in the fourth aspect.

[0070] In this invention, “ENSEMBAL number” and “Ensembl ID” can be used interchangeably, both referring to the number in the Ensembl database corresponding to the rat reference genome Rattus_norvegicus.mRatBN7.2.110 (7th generation assembly of the Brown Norway strain (mRatBN7.0), whose gene annotation version is derived from Ensembl release 110).

[0071] In this invention, the term "cell segmentation" refers to the step in Xenium spatial in situ analysis where, using morphological multi-channel images (including but not limited to nuclear dyes, cell membrane / cytoplasmic dyes, RNA probe signals, and / or intracellular protein markers) obtained from the same tissue slice as input, and employing artificial intelligence or classical image algorithms, the cell population in the image is divided into individual, independent, closed regions, and corresponding cell masks are generated. The masks are used to define the spatial boundaries for subsequent transcript counting, thereby achieving single-cell-level gene expression quantification.

[0072] In this invention, the term "cell calling" or "cell assignment" refers to the process of statistically analyzing the number of transcripts, genes, and / or morphological characteristics in each mask region after cell segmentation and obtaining the cell mask, and assigning a unique cell identifier (cell ID) to the mask that meets the "effective cell" standard based on a preset threshold or machine interpretation model; masks that do not meet the standard are discarded or marked as fragments / background, thereby ensuring that each row of the downstream expression matrix represents a real and complete single cell.

[0073] In this invention, the term "machine definition" refers to defining cell types by defining a machine learning model based on the cell category definition rules.

[0074] The technical problem that this invention seeks to solve is:

[0075] 1. Segmentation of myocardial tissue cells:

[0076] ① It can accurately identify cell boundaries;

[0077] ② Subcellular level gene expression, such as the distinction between the cell nucleus and the cell membrane;

[0078] ③ Algorithm identification for multinucleated cells.

[0079] 2. Cell definition:

[0080] ① Identify different cell types in cardiac tissue: including epicardial cells, endothelial cells, fibroblasts, pericytes, smooth muscle cells, cardiomyocytes, immune cells, etc.

[0081] ② Identify different cell subtypes: such as proliferating cells of different cell types;

[0082] ③ Highly efficient and specific definition results.

[0083] 3. Stable, sensitive, and effective gene and hybridization probes

[0084] ① Expression on a single cell, in appropriate amounts;

[0085] ②Reduce the mutual interference between different probes;

[0086] ③ High-specificity hybridization and in situ presentation.

[0087] Invention point 3 provides a marker gene for invention point 2 for cell definition; the combination of invention points 1 and 2 can provide a solution for cardiac Xenium research.

[0088] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.

[0089] The reagents and raw materials used in this invention are all commercially available.

[0090] The significant advantages of this invention are as follows: It provides a combination of biomarkers, a detection reagent, and a cell definition method. The biomarker combination can serve as a target for cell segmentation and cell definition in the heart, and the detection reagent enables effective and specific detection of the target. Furthermore, a method for defining the segmented cells is developed based on the biomarker combination and / or the detection reagent. Verification has shown that the biomarker combination, detection reagent, and cell definition method provided by this invention can effectively and sensitively identify cardiac cell types and subtypes in Xenium, enabling precise cell segmentation and cell type definition in cardiac spatial transcriptomics studies based on Xenium technology, thus providing a reliable foundation for cardiac research. Attached Figure Description

[0091] Figure 1 This refers to the Xenium myocardial tissue cell segmentation technique.

[0092] Figure 2 A comparison of Visum spatial transcriptome cell segmentation and Xenium myocardial tissue multimodal cell segmentation.

[0093] Figure 3 Xenium myocardial tissue cells were divided into single-core and dual-core myocardial cells.

[0094] Figure 4 In situ images and sensitivity and specificity statistics of myocardial tissue-specific biomarkers and definition methods.

[0095] Figure 5 In situ images and sensitivity and specificity statistics of myocardial tissue-specific biomarkers and definition methods.

[0096] Figure 6 The algorithm flow for defining major cell categories.

[0097] Figure 7 The algorithm flow for defining cell subclasses.

[0098] Figure 8Evaluation of cell category definitions: The first row shows single-cell Umap images of fibroblasts (Fib), endothelial cells (EC), pericytes (Peri), and immune cells (Immune); the second row shows HE staining of corresponding tissues and expression of relevant cell marker genes; the third row shows the cell definition results and their spatial relationship with gene expression; the fourth row shows single-cell Umap images of smooth muscle cells (SMC), epicardium cells (Epi), cardiomyocytes (CM), and proliferating cells (Proliferating cells); the fifth row shows HE staining of corresponding tissues; the sixth row shows the expression of relevant cell marker genes; and the seventh row shows the cell definition results and their spatial relationship with gene expression.

[0099] Figure 9 Evaluation of the results of the cardiomyocyte subclass definition.

[0100] Figure 10 A and B are used to evaluate each cell defined by the algorithm, calculating the defined sensitivity and specificity: Figure 10 A is a cell that can be identified as a fibroblast based on the expression of Dcn. If it is defined as a fibroblast according to the patented algorithm, it is correct. If it is defined as an endothelial cell, the specificity decreases. If it is an undefined cell, the sensitivity decreases. Figure 10 B represents the quantitative analysis of sensitivity and specificity for each cell type.

[0101] Figure 11 This image shows a view of the heart and blood vessels and the expression of Xenium marker genes. The yellow circle represents the endothelial layer of the blood vessel, where endothelial cell Cdh5 (orange) is the main gene expressed. The green circle represents the middle layer of the blood vessel, where smooth muscle marker gene Myh11 (green) is the main gene expressed. The pink circle represents the outer layer of the blood vessel, where fibroblast marker Dcn (purple) is the main gene expressed.

[0102] Figure 12 This document presents the cell definitions, sensitivity, and specificity statistics for Xenium's existing myocardial tissue cell definition scheme. Detailed Implementation

[0103] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.

[0104] The specific biomarkers involved in this application and their corresponding numbers in the database are as follows:

[0105] Upk3b(Ensembl ID: ENSRNOG00000023686)、Myh11(ENSRNOG00000057880)、Kcnj8(Ensembl ID: ENSRNOG00000013463)PrcEpr ID:ENSRNOG00000000655)、Cdh5(Ensembl ID: ENSRNOG00000013324)、Dcn(Ensembl ID:ENSRNOG00000004554)、Myom2 (Ensembl ID: ENSRNOG00000011754)、Aurkb(Ensembl ID:ENSRNOG00000005659)、Pole(Ensembl ID: ENSRNOG00000037449)、Ect2(Ensembl ID:ENSRNOG00000024365)、Ccna2(Ensembl ID: ENSRNOG00000015423)、Nppa(Ensembl ID:ENSRNOG00000008176)、Nppb ID:(Ensembl ID: ENSRNOG00000008141)、Xirp2(Ensembl ID:ENSRNOG00000034258)、Myh7(EnSRNOG00000018997)、Corin(Ensembl ID:ENSRNOG00000002302)、myh7b(Ensembl ID: ENSRNOG00000018997)、Tnni1(EnSRNOG00000009073)、Fgf12:Ensembl ID:(Ensembl ID: ENSRNOG00000001931)、Acta2(Ensembl ID:ENSRNOG00000058039)、Pdgfrb(EnSRNOG00000018461)、Flt1(Ensembl1) ID:ENSRNOG00000000940)、Col1a2(Ensembl ID: ENSRNOG00000011292)、Wt1(Ensembl ID:ENSRNOG00000013074)、Myl2(Ensembl ID: ENSRNOG00000030848)、Top2a(Ensembl ID:ENSRNOG00000053047)、Mki67(Ensembl ID:ENSRNOG00000028137), Ccnb1 (Ensembl ID: ENSRNOG00000058539) and Xirp2 (Ensembl ID: ENSRNOG00000034258).

[0106] Example 1: Segmentation of Myocardial Tissue Cells

[0107] This invention provides a multimodal cardiomyocyte segmentation technique that uses nuclear markers (DAPI), membrane markers (ATP1A1, E-cadherin), cytoplasmic markers (18S ribosomal RNA), and protein staining (αSMA / vimentin) to label the cell nucleus, cell membrane, and cytoplasm respectively, in order to distinguish cell boundaries and achieve the purpose of cell segmentation. Figure 1 ).

[0108] Materials: Staining kits (10×Genomics, catalog numbers PN-2000991 and PN-1000661) contain four channels for identifying right-ventricular cardiomyocytes (RV cells): membrane markers (ATP1A1, E-cadherin), cytoplasmic markers (18S ribosomal RNA), and protein staining (αSMA / Vimentin), DAPI staining solution (catalog number PN2000762), and autofluorescence quenching mixture (catalog number PN2000753).

[0109] Specific steps: The mixture was incubated with Xenium slides containing heart sections at room temperature for 30 minutes. The slides were then treated with Xenium autofluorescence quenching solution and incubated in the dark at room temperature for 10 minutes. Next, the slides were incubated with 500 μL of Xenium nuclear staining buffer in the dark at room temperature for 1 minute. Cell segmentation analysis was performed using Xenium v2.0 software. The segmentation process consisted of three main steps: 1) Cell boundary demarcation, including multinucleated cells such as cardiomyocytes, by membrane staining; 2) Cell segmentation by nuclear staining, protein staining, and RNA staining; 3) Extending the nuclear boundary by 5 μm until it reached the boundary of another cell.

[0110] Results Display:

[0111] 1) Visium spatial transcriptomics uses a spatial lattice method to identify tissue regions for sequencing. Currently, the highest resolution is 2 μm, but there are gaps between the lattice points, which cannot simulate the actual cell morphology, cannot accurately locate to the subcellular level, and can lead to abnormal cell proportions (larger cells correspond to more regions with higher proportions). This can produce more errors in practical applications. Figure 2 (the upper half).

[0112] 2) Xenium multimodal cell segmentation can accurately identify the cell membrane, thus depicting the true boundaries of the cell with a resolution of 0.2 μm. It can also accurately locate gene expression in subcellular organelles (such as the nucleus and cell membrane). Figure 2 The lower half (expression of nuclear genes).

[0113] 3) Multimodal cell segmentation can clearly identify the boundaries of cardiomyocytes, and can also identify and segment multi-core myoblasts based on the continuity of the cell membrane. Figure 3 ).

[0114] Example 2: Design of Specific Genes and Probes

[0115] Xenium sequencing requires probe hybridization to locate genes, depending on the experimental objective.

[0116] The 10x official recommended organ gene sets contain 300 genes per set, with each gene typically containing 6-8 probes. However, there is currently no gene set applicable to the heart. Using genes from other organs to define the heart is less effective (for example, the myocardial cell gene set lacks corresponding marker genes).

[0117] The 5k gene set officially released by 10x contains 5,000 genes, each with 1-2 probes, used to define the heart. However, gene expression is low, leading to severe channel congestion and inaccurate results.

[0118] To accurately define and validate cell types, this invention designed a gene set including the genes Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Myom2, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Myh7, and Corin, and independently designed 6-8 probes for each gene. Each probe contains two 18-22 nucleotide two-hybrid regions containing gene-specific barcodes, which can independently bind to complementary RNA targets after incubation at 50°C for 16 hours. Based on gene expression, cell definition methods can be applied to define tissue cells.

[0119] The probes designed in this invention are shown in Table 1:

[0120] Table 1 Probe sequence information

[0121]

[0122]

[0123]

[0124]

[0125] The ENSEMBAL number is based on the genome sequence of the Norwegian rat (Rattus norvegicus), with the genome assembly version number mRatBN7.2.

[0126] Probe application steps: The Xenium platform (10x Genomics) employs a four-step process for targeted RNA detection: (1) probe hybridization; (2) post-hybridization washing; (3) ligation reaction; and (4) enzymatic amplification. The platform uses a custom 300-fold probe array to process FF tissue sections, followed by removal of unbound probes for 30 minutes at room temperature using Xenium post-hybridization washing buffer (PN2000395). After the free probes are removed, ligation reagents are used to seal the junctions between probe regions, forming circular DNA structures (PN2000391, 2000397, and 2000398, incubated at 37°C for 2 hours). These ligation products are then subjected to rolling circle amplification using the master amplification mixture (PN2000392 and 2000399, 30°C / 120 minutes), generating hundreds of gene-specific barcode copies during amplification. Xenium analysis software version 3.0.0.15 is used to decode the subcellular localization of RNA targets. The instrument first undergoes 10-20 minutes of initialization and self-test. The Xenium consumables include sample washing buffer (PN3001198-001200), probe removal buffer (PN3001201), and two sample cartridges, which must be assembled into the instrument. The slides are then scanned for one hour. Autofluorescence is assessed by selecting the yellow channel, and the panoramic scan image is checked to ensure no sample residue. The slides are divided into grid cells, each representing a field of view (FOV). FOVs containing tissue are selected, while all blank FOVs are excluded. Once the consumables are confirmed to be loaded, the run is initiated. During imaging, punctate structures with localized fluorescence intensity are identified as potential RNA spots. Each gene in the panel exhibits a unique fluorescence pattern in different imaging channels; spots matching specific patterns are then decoded and labeled based on gene ID. After the run, the slides are cleaned and unloaded. All FOV images are stitched together and the data exported. The onboard analysis workflow generates a quality score for each detected transcript, reflecting the confidence difference between signal detection and decoding. Data for TMA1-TMA4 were obtained using instrument software version 1.1.2.4 and analysis version xenium-1.1.0.2, while data for TMA5 were obtained using instrument software version 3.0.2.0 and analysis version xenium-3.0.0.15.

[0127] result:

[0128] 1) Gene set and corresponding probe establishment: The gene set designed according to the present invention can be applied to cell definition, and the probe can be applied to Xenium in situ sequencing.

[0129] 2) Gene expression specificity in single cells: All genes were designed based on the team's previous single-cell nuclear sequencing results, and the specificity of gene expression in single cells was evident. Figure 8 In the first and fourth rows, the major cell categories are expressed only in their respective cell types, and not in other cell types.

[0130] 3) High-specificity gene labeling and in situ presentation in xenium. For example, using vascular labeling... Figure 11 It can be seen that different vascular layers are composed of different cell types, and different cell types specifically express their marker genes, which are all derived from the designed gene set.

[0131] Example 3: Screening and definition of myocardial tissue-specific markers

[0132] For the segmented cells, further cell definition is required based on the specific expression levels of the genes. During the development of this invention, multiple gene sets were designed as specific biomarker combinations, and corresponding probe libraries were designed (specific probe sequences are shown in Tables 1 and 2; probe application steps are as described in Example 2). Corresponding definition rules were also established, and cells were defined based on the specific expression levels of these genes. The optimal biomarker combination and corresponding definition method were obtained through screening.

[0133] Table 2 Probe sequence information

[0134]

[0135]

[0136]

[0137] The ENSEMBAL number is based on the genome sequence of the Norwegian rat (Rattus norvegicus), with the genome assembly version number mRatBN7.2.

[0138] 1. Define Method 1 (First Modification)

[0139] 1.1 Specific biomarker combination 1

[0140] Including: Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Upk3b, Myom2, Myh7b, Aurkb, Pole, Ect2, Ccna2, Nppa, Tnni1, and Fgf12.

[0141] 1.2 Cell Definition

[0142] Inside the core:

[0143] First, define smooth muscle cells as those with Myh11 nuclear expression ≥4;

[0144] Then, pericytes are defined as cells with Kcnj8 nuclear expression ≥4;

[0145] Then, immunity is defined as: Ptprc nuclear expression ≥4;

[0146] Then define endothelium as: Cdh5 nuclear expression ≥5;

[0147] Then it forms fibers: Dcn core ≥5;

[0148] Inside the core + outside the core:

[0149] Define epicardium: Upk3b≥3;

[0150] Then define cardiomyocytes as: Myom2≥3 or Myh7b≥2;

[0151] Finally, the other cells are defined as undefined cells.

[0152] 1.3 Rules for Defining Cell Subpopulations

[0153] Proliferating cells: For each cell type, if the sum of Aurkb (GeneH), Pole (GeneI), Ect2 (GeneJ), and Ccna2 (GeneK) is independently measured, and the sum exceeds 6, the cell type is labeled as a proliferating cell. The same method can be used to define mature, hypertrophic, and immature cardiomyocyte subsets.

[0154] Hypertrophic cardiomyocytes: Total Nppa expression level >20;

[0155] Immature cardiomyocytes: Tnni1 expression level >10;

[0156] Mature cardiomyocytes: Fgf12 (GeneP) expression level >5.

[0157] 2. Define Method 2 (Second Revision)

[0158] 2.1 Specific biomarker combination 2

[0159] Including: Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Upk3b, Myom2, Myh7b, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Tnni1, Myh7, Fgf12, and Corin.

[0160] 2.2 Cell Definition

[0161] First, define smooth muscle cells as those with ≥4% MYH11 expression in the nucleus.

[0162] Then, pericytes are defined as cells with KCNJ8 nuclear expression ≥4;

[0163] Then, immunity is defined as: Ptprc nuclear expression ≥4;

[0164] Endothelial or fibroblastic: >5 CDH5 or DCN nuclei;

[0165] First, define the endothelium: CDH5 ≥ DCN;

[0166] Then it forms fibers: CDH5 < DCN;

[0167] Inside the core + outside the core:

[0168] Define epicardium: Upk3b≥3;

[0169] Then define cardiomyocytes as: MYOM2≥3 or myh7b≥2;

[0170] Finally, the remaining cells are defined as undefined cells.

[0171] 2.3 Rules for Defining Cell Subpopulations

[0172] Proliferating cells: For each cell type, if the sum of Aurkb (GeneH), Pole (GeneI), Ect2 (GeneJ), and Ccna2 (GeneK) is independently measured, and the sum exceeds 8, the cell type is labeled as a proliferating cell. The same method can be used to define mature, hypertrophic, and immature cardiomyocyte subsets.

[0173] Hypertrophic cardiomyocytes: Total expression levels of Nppa, Nppb, and Xirp2 >50;

[0174] Immature cardiomyocytes: Tnni1 and Myh7 expression levels >10;

[0175] Mature cardiomyocytes: Fgf12 and Corin (GeneP) expression levels >5.

[0176] 3. Define Method 3

[0177] 3.1 Specific biomarker combination 3 (i.e., control biomarker)

[0178] Acta2, Pdgfrb, Ptprc, Flt1, Col1a2, Wt1, Myl2, Tnni1, Top2a, Mki67, Ccnb1, Aurkb, Nppb, Xirp2, Myh7, Myh7b and Corin.

[0179] 3.2 Cell Definition

[0180] First, define smooth muscle cells as those expressing ≥4 sites of Acta2 in the nucleus.

[0181] Then, pericytes are defined as cells with ≥4 nuclear sites of Pdgfrb expression.

[0182] Then, immunity is defined as: Ptprc nuclear expression ≥ 4 sites.

[0183] Endothelial or fibroblastic: Flt1 or Col1a2 nuclei with more than 5

[0184] First, define the endothelium: Flt1 ≥ Col1a2

[0185] Then it forms fibers: Flt1 is less than Col1a2

[0186] Inside the core + outside the core:

[0187] Define epicardium: Wt1 ≥ 3

[0188] Then define cardiomyocytes as: Myl2 ≥ 3 or Tnni1 ≥ 2.

[0189] Finally, other cells are defined as undefined cells.

[0190] 3.3 Rules for Defining Cell Subpopulations

[0191] Proliferating cells: For each cell type, if the sum of Top2a (GeneH), Mki67 (GeneI), Ccnb1 (GeneJ), and Aurkb (GeneK) is independently measured and exceeds 8, the cell type is labeled as a proliferating cell. The same method can be used to define mature, hypertrophic, and immature cardiomyocyte subsets.

[0192] Hypertrophic cardiomyocytes: Nppb and Xirp2 total expression levels >50;

[0193] Immature cardiomyocytes: Myh7 or Myh7b expression level >10;

[0194] Mature cardiomyocytes: Corin (GeneP) expression level >5

[0195] 4. Finalized combination of specific biomarkers and cell definition method:

[0196] 4.1 The specific biomarker combination includes: Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Myom2, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Myh7, and Corin.

[0197] 4.2 The specific rules for defining major cell categories are as follows ( Figure 6 ):

[0198] Step 1: If the cell expresses Upk3b (GeneA) ≥ 3, it is identified as an epicardial cell; otherwise, it proceeds to the next layer.

[0199] Step 2: If the cell expresses Myh11 (GeneB) ≥ 4, it is identified as a smooth muscle cell; otherwise, proceed to the next layer.

[0200] Step 3: If the cell expresses Kcnj8 (GeneC) ≥ 4, it is determined to be a pericyte; otherwise, proceed to the next layer.

[0201] Step 4: If the cell expresses Ptprc(GeneD) ≥ 4, it is identified as an immune cell; otherwise, it proceeds to the next layer.

[0202] Step 5: If the cell expresses Cdh5 (GeneE) or Dcn (GeneF) ≥ 5, it is determined to be an endothelial cell or a fibroblast; otherwise, it proceeds to the next layer: where Cdh5 (GeneE) > Dcn (GeneF), it is an endothelial cell; otherwise, it is a fibroblast.

[0203] Step 6: If Myom2(GeneG)≥3, then it is a cardiomyocyte; otherwise, it is an undefined cell.

[0204] 4.3 Rules for defining cell subsets ( Figure 7 )

[0205] Proliferating cells: For each cell type, if the sum of Aurkb (GeneH), Pole (GeneI), Ect2 (GeneJ), and Ccna2 (GeneK) is independently measured, and the sum exceeds 5, the cell type is labeled as a proliferating cell. The same method can be used to define mature, hypertrophic, and immature cardiomyocyte subsets.

[0206] Hypertrophic cardiomyocytes: Total expression level of Nppa+Nppb+Xirp2 (GeneL-N, respectively) >20;

[0207] Immature cardiomyocytes: Myh7 (GeneO) expression level >20;

[0208] Mature cardiomyocytes: Corin (GeneP) expression level >5.

[0209] The samples were cell-defined using the specific marker combinations shown in Examples 1-4 (corresponding probes are shown in Tables 1 and 2; cell segmentation steps for probe application are as shown in Example 2) and cell definition methods. To clarify the effectiveness of each scheme, the sensitivity and specificity of each scheme were tested. Specifically, the samples (n=10) were randomly selected 20x magnification fields of view of the cells segmented in Example 1. Cells of various types in the field of view were summarized, and the machine-defined cells were manually defined according to the definition method. The results were compared with the machine-defined cells. If the machine definition was incorrect, the specificity decreased; if the machine defined an undefined cell, the sensitivity decreased. Specificity = 1 - number of incorrectly defined cells / total number of cells of that cell type; Sensitivity = 1 - number of undefined cells of that cell type / total number of cells of that cell type.

[0210] The following is the specific R language source code used in this invention to define machine learning models (machine-defined):

[0211] library(future)

[0212] if (future::supportsMulticore()) {future::plan(future::multicore,workers = 6)} else {future::plan(future::multisession, workers = 6)}

[0213] options(future.globals.maxSize = 10000 * 1024^2)

[0214] library(Seurat)

[0215] library(data.table)

[0216] library(dplyr)

[0217] library(tidyr)

[0218] options(scipen = 999)

[0219] expr=readRDS('.. / combin.data.RDS')

[0220] DefaultAssay(expr)="Xenium"

[0221] #P1714-2

[0222] sample='P1714-3'

[0223] id=4

[0224] dat=fread(paste0(' / bhpublic / datas02 / luolf / xenium / E20240481-01-03 / RNA / ',sample,' / outs / transcripts.csv.gz'))

[0225] dat=as.data.frame(dat)

[0226] cells3=data.frame(cell_id=colnames(expr),expr$orig.ident)

[0227] dat2=dat[,c(2,4)]

[0228] dat2=dat[dat$cell_id != 'UNASSIGNED',]

[0229] dat2$cell_id=paste0(dat2$cell_id,'_',id)

[0230] henei=dat2[dat2$overlaps_nucleus==1,c(2,4)]

[0231] #hewai=dat2[dat2$overlaps_nucleus==0,c(2,4)]

[0232] summary_data <- henei %>%

[0233] group_by(cell_id, feature_name) %>%

[0234] summarise(gene_count = n(), .groups = 'drop')

[0235] summary_data=as.data.frame(summary_data)

[0236] summary_data2=merge(summary_data,cells3,by='cell_id')

[0237] summary_data2=summary_data2[,-4]

[0238] genes_info <- list(Myh11 = 4, Kcnj8 = 4, Ptprc = 4, Cdh5 = 5, Dcn =5)

[0239] celltype=c('Smooth muscle cells','Pericytes','Immune cells','Endothelial cells','Fibroblasts')

[0240] gene_classification <- data.frame(cell_id = character(), celltype =character(), stringsAsFactors = FALSE)

[0241] # Create a collection to track the assigned cell_ids

[0242] assigned_cell_ids <- character()

[0243] # Loop through each gene and its corresponding selection criteria

[0244] for (i in 1:length(names(genes_info)) ){

[0245] # Get the current gene's screening criteria

[0246] gene_name <- names(genes_info)[i]

[0247] threshold <- genes_info[[gene_name]]

[0248] # Screening current genes

[0249] gene_data <- subset(summary_data2, feature_name == gene_name & gene_count >= threshold)

[0250] # Ensure that the cell_id selected by the current gene screening is not duplicated.

[0251] valid_cells <- setdiff(gene_data$cell_id, assigned_cell_ids) # Remove the assigned cell_ids

[0252] # Update the assigned cell_ids

[0253] assigned_cell_ids <- c(assigned_cell_ids, valid_cells)

[0254] # Add the valid cell_ids to gene_classification

[0255] gene_classification <- rbind(gene_classification, data.frame(cell_id = valid_cells, celltype = celltype[i]))}

[0256] table(gene_classification$celltype)

[0257] dat3 = dat2[!(dat2$cell_id %in% gene_classification$cell_id), c(2, 4)]

[0258] summary_data3 <- dat3 %>%

[0259] group_by(cell_id, feature_name) %>%

[0260] summarise(gene_count = n(),.groups = 'drop')

[0261] summary_data3 = as.data.frame(summary_data3)

[0262] summary_data4 = merge(summary_data3, cells3, by = 'cell_id')

[0263] summary_data4 = summary_data4[, -4]

[0264] summary_data5=rbind(summary_data2,summary_data4)

[0265] write.table(summary_data5,paste0(sample,'_genecount2.csv'),sep = ',',col.names = T,row.names = F,quote = F)

[0266] genes_info <- list(Upk3b = 3, Myom2 = 3)

[0267] Myh7b =2

[0268] celltype=c('Epicardial cells','Cardiomyocytes')

[0269] gene_classification2 <- data.frame(cell_id = character(), celltype =character(), stringsAsFactors = FALSE)

[0270] # Create a collection to track the assigned cell_ids

[0271] assigned_cell_ids <- character()

[0272] # Loop through each gene and its corresponding selection criteria

[0273] for (i in 1:length(names(genes_info)) ){

[0274] # Get the current gene's screening criteria

[0275] gene_name <- names(genes_info)[i]

[0276] threshold <- genes_info[[gene_name]]

[0277] # Screening current genes

[0278] gene_data <- subset(summary_data4, feature_name == gene_name & gene_count >= threshold)

[0279] # Ensure that the cell_id selected by the current gene screening is not duplicated.

[0280] `valid_cells <- setdiff(gene_data$cell_id, assigned_cell_ids)` # Removes assigned cell_ids.

[0281] # Update the assigned cell_id

[0282] assigned_cell_ids <- c(assigned_cell_ids, valid_cells)

[0283] # Add a valid cell_id to gene_classification

[0284] gene_classification2 <- rbind(gene_classification2, data.frame(cell_id = valid_cells, celltype=celltype[i]))

[0285] }

[0286] # Loop through each gene and its corresponding selection criteria

[0287] #for (i in 1:length(names(genes_info)) ){

[0288] # Get the current gene's screening criteria

[0289] i=1

[0290] gene_name <- names(genes_info)[i]

[0291] threshold <- genes_info[[gene_name]]

[0292] # Screening current genes

[0293] gene_data <- subset(summary_data4, feature_name == gene_name & gene_count >= threshold)

[0294] # Ensure that the cell_id selected by the current gene screening is not duplicated.

[0295] `valid_cells <- setdiff(gene_data$cell_id, assigned_cell_ids)` # Removes assigned cell_ids.

[0296] # Update the assigned cell_id

[0297] assigned_cell_ids <- c(assigned_cell_ids, valid_cells)

[0298] # Add a valid cell_id to gene_classification

[0299] gene_classification2 <- rbind(gene_classification2, data.frame(cell_id = valid_cells, celltype=celltype[i]))

[0300] #}

[0301] i=2

[0302] # Screening current genes

[0303] gene_data <- subset(summary_data4, (feature_name == 'Myom2' & gene_count >= 3) | (feature_name == 'Myh7b' & gene_count >= 2))

[0304] # Ensure that the cell_id selected by the current gene screening is not duplicated.

[0305] `valid_cells <- setdiff(gene_data$cell_id, assigned_cell_ids)` # Removes assigned cell_ids.

[0306] # Update the assigned cell_id

[0307] assigned_cell_ids <- c(assigned_cell_ids, valid_cells)

[0308] # Add valid cell_ids to gene_classification

[0309] gene_classification2 <- rbind(gene_classification2, data.frame(cell_id = valid_cells, celltype=celltype[i]))

[0310] gene_classification3=rbind(gene_classification,gene_classification2)

[0311] unassigned=data.frame(cell_id=setdiff(unique(summary_data5$cell_id),gene_classification3$cell_id),celltype='Unassign')

[0312] gene_classification3 <- rbind(gene_classification3,unassigned)

[0313] print(table(gene_classification3$celltype))

[0314] write.table(gene_classification3,paste0(sample,'_celltype2.csv'),sep= ',',col.names = T,row.names = F,quote = F)

[0315] 12.20##

[0316] genes_info <- list(

[0317] Myh11 = 4,

[0318] Kcnj8 = 4,

[0319] Ptprc = 4 )

[0321] celltype=c('Smooth muscle cells','Pericytes','Immune cells')

[0322] gene_classification <- data.frame(cell_id = character(), celltype =character(), stringsAsFactors = FALSE)

[0323] assigned_cell_ids <- character()

[0324] for (i in 1:length(names(genes_info)) ){gene_name <- names(genes_info)[i] threshold <- genes_info[[gene_name]] gene_data <- subset(summary_data2, feature_name == gene_name & gene_count >= threshold)

[0325] valid_cells <- setdiff(gene_data$cell_id, assigned_cell_ids) # Remove the assigned cell_id

[0326] assigned_cell_ids <- c(assigned_cell_ids, valid_cells)

[0327] gene_classification <- rbind(gene_classification, data.frame(cell_id= valid_cells, celltype=celltype[i]))

[0328] }

[0329] # Screen the remaining cells in the nucleus for endothelial and fibroblast classification

[0330] remaining_cells <- summary_data2[!(summary_data2$cell_id %in%assigned_cell_ids), ]

[0331] # Handling the merging logic of CDH5 and DCN

[0332] cdh5_dcn_data <- subset(remaining_cells, (feature_name == 'Cdh5' &gene_count >= 5) | (feature_name == 'Dcn' & gene_count >= 5))

[0333] # Calculate the total expression of CDH5 and DCN

[0334] cdh5_dcn_data_wide <- cdh5_dcn_data %>%

[0335] spread(key = feature_name, value = gene_count, fill = 0)

[0336] cdh5_dcn_data_wide$total_count <- rowSums(cdh5_dcn_data_wide[, c('Cdh5', 'Dcn')], na.rm = TRUE)

[0337] # Classified as endothelial or fibroblast based on CDH5 and DCN expression levels.

[0338] cdh5_dcn_data_wide$celltype <- ifelse(cdh5_dcn_data_wide$Cdh5 >=cdh5_dcn_data_wide$total_count / 2, 'Endothelial cells', 'Fibroblasts')

[0339] # Add the classification results to the original data

[0340] gene_classification <- rbind(gene_classification, data.frame(cell_id= cdh5_dcn_data_wide$cell_id, celltype = cdh5_dcn_data_wide$celltype))

[0341] table(gene_classification$celltype)

[0342] # Data processing and screening outside the nucleus

[0343] dat3=dat2[!(dat2$cell_id %in% gene_classification$cell_id),c(2,4)]

[0344] summary_data3 <- dat3 %>%

[0345] group_by(cell_id, feature_name) %>%

[0346] summarise(gene_count = n(),.groups = 'drop')

[0347] summary_data3=as.data.frame(summary_data3)

[0348] summary_data4=merge(summary_data3,cells3,by='cell_id')

[0349] summary_data4=summary_data4[,-4]

[0350] summary_data5=rbind(summary_data2,summary_data4)

[0351] gene_classification2 <- data.frame(cell_id = character(), celltype =character(), stringsAsFactors = FALSE)

[0352] assigned_cell_ids2 <- character()

[0353] gene_name <- 'Upk3b'

[0354] threshold <- 3

[0355] # Screening current genes

[0356] gene_data <- subset(summary_data4, feature_name == gene_name & gene_count >= threshold)

[0357] # Ensure that the cell_id selected by the current gene screening is not duplicated.

[0358] `valid_cells <- setdiff(gene_data$cell_id, assigned_cell_ids)` # Removes assigned cell_ids.

[0359] # Update the assigned cell_id

[0360] assigned_cell_ids <- c(assigned_cell_ids, valid_cells)

[0361] # Add a valid cell_id to gene_classification

[0362] gene_classification2 <- rbind(gene_classification2, data.frame(cell_id = valid_cells, celltype='Epicardial cells'))

[0363] gene_data <- subset(summary_data4, (feature_name == 'Myom2' & gene_count >= 3) | (feature_name == 'Myh7b' & gene_count >= 2))

[0364] # Ensure that the cell_id selected by the current gene screening is not duplicated.

[0365] valid_cells <- setdiff(gene_data$cell_id, assigned_cell_ids) # Remove the assigned cell_ids

[0366] # Update the assigned cell_ids

[0367] assigned_cell_ids <- c(assigned_cell_ids, valid_cells)

[0368] # Add the valid cell_ids to gene_classification

[0369] gene_classification2 <- rbind(gene_classification2, data.frame(cell_id = valid_cells, celltype='Cardiomyocytes'))

[0370] gene_classification3=rbind(gene_classification,gene_classification2)

[0371] unassigned=data.frame(cell_id=setdiff(unique(summary_data5$cell_id),gene_classification3$cell_id),celltype='Unassign')

[0372] gene_classification3 <- rbind(gene_classification3,unassigned)

[0373] print(table(gene_classification3$celltype))

[0374] write.table(gene_classification3,paste0(sample,'_celltype2.csv'),sep= ',',col.names = T,row.names = F,quote = F)

[0375] The results are as Figure 4 and Figure 5As shown, the sensitivity and specificity of the final definition method determined after optimization are significantly improved. Both the control biomarker combinations and the biomarker combination provided by this invention can perform xenium-based cell definition, but the sensitivity and specificity are superior when using the biomarker combination and its corresponding definition method provided by this invention.

[0376] Example 3: Performance Verification of the Myocardial Tissue Definition Algorithm

[0377] 1. The specific rules for defining major cell categories are as follows ( Figure 6 ):

[0378] Step 1: If the cell expresses Upk3b (GeneA) ≥ 3, it is identified as an epicardial cell; otherwise, it proceeds to the next layer.

[0379] Step 2: If the cell expresses Myh11 (GeneB) ≥ 4, it is identified as a smooth muscle cell; otherwise, proceed to the next layer.

[0380] Step 3: If the cell expresses Kcnj8 (GeneC) ≥ 4, it is determined to be a pericyte; otherwise, proceed to the next layer.

[0381] Step 4: If the cell expresses Ptprc(GeneD) ≥ 4, it is identified as an immune cell; otherwise, it proceeds to the next layer.

[0382] Step 5: If the cell expresses Cdh5 (GeneE) or Dcn (GeneF) ≥ 5, it is determined to be an endothelial cell or a fibroblast; otherwise, it proceeds to the next layer: where Cdh5 (GeneE) > Dcn (GeneF), it is an endothelial cell; otherwise, it is a fibroblast.

[0383] Step 6: If Myom2(GeneG)≥3, then it is a cardiomyocyte; otherwise, it is an undefined cell.

[0384] 2. Rules for defining cell subsets ( Figure 7 )

[0385] Proliferating cells: For each cell type, if the sum of Aurkb (GeneH), Pole (GeneI), Ect2 (GeneJ), and Ccna2 (GeneK) is independently measured, and the sum exceeds 5, the cell type is labeled as a proliferating cell. The same method can be used to define mature, hypertrophic, and immature cardiomyocyte subsets.

[0386] Hypertrophic cardiomyocytes: Total expression level of Nppa+Nppb+Xirp2 (GeneL-N, respectively) >20;

[0387] Immature cardiomyocytes: Myh7 (GeneO) expression level >20;

[0388] Mature cardiomyocytes: Corin (GeneP) expression level >5.

[0389] Results Display:

[0390] 1) Based on the cell types defined by the algorithm, visualize them in space and associate them with the expression of specific marker genes.

[0391] The cell type-related results of the final determined cell definition method are as follows: Figure 8 As shown: Figure 8 The first row shows the expression of marker genes for fibroblasts, endothelial cells, pericytes, and immune cells in single cells, which can demonstrate the effectiveness of marker genes. Figure 8 The second row shows the expression of marker genes in Xenium. It can be seen that the marker genes of different cell populations are fixedly expressed in a certain cell. The Xenium cell type can be determined based on the threshold. Figure 8 The third row shows the expression of marker genes after defining cell types, to observe whether the marker genes are expressed in the corresponding cell types; Figure 8 The fourth row shows the expression of marker genes for smooth muscle, epicardium, myocardium, and proliferating cells in single cells; Figure 8 The fifth row shows the HE staining pattern corresponding to Xenium; Figure 8 The sixth row shows the expression of marker genes in Xenium; Figure 8 The seventh row shows the cell definition results and their spatial relationship with gene expression. Figure 8 The first to seventh rows are used to observe the specificity of marker gene expression and the accuracy of defining cell types.

[0392] 2) The definition results of cell subclasses are as follows: Figure 9 As shown, the expression of marker genes for cardiomyocyte subsets and cell definitions in different groups show that the marker genes corresponding to the subsets are specifically expressed in the defined subsets.

[0393] 3) Cell definition

[0394] Five 20× fields of view are randomly selected. Cell types are manually determined based on the expression levels of marker genes in the cells according to the final determined cell definition method. These cell types are then compared with those determined by the cell definition algorithm (machine definition). If the machine fails to identify a cell type, the sensitivity decreases; if it identifies the wrong cell type, the specificity decreases. Figure 10 The sensitivity and specificity of the machine algorithm are calculated in this way, such as A). Figure 10 As shown in B.

[0395] Example 5

[0396] Compared with existing Xenium cell segmentation and cell definition schemes, the scheme designed in this patent solves the following problems:

[0397] 1. The segmentation of heart cells is not accurate enough:

[0398] Existing solutions

[0399] ① Defining cell types based on nuclear expansion patterns is not a good fit for cell boundaries because different cells have different sizes;

[0400] ② Conventional multimodal staining is insufficient to fully identify multinucleated cells in the myocardium, leading to bias;

[0401] ③ Other spatial transcriptomics studies (such as VisiumHD and Stereoseq) use a lattice-like approach for spatial localization, which makes it difficult to locate subcellular structures, such as the cell membrane and cytoplasm, in practical applications.

[0402] The proposed scheme uses conventional multimodal staining for the heart, which can identify cardiomyocyte course, cell boundaries, subcellular structures, and perform multi-core myocardial cell segmentation, thus better reflecting the real-world situation than existing schemes.

[0403] 2. Existing solutions all define cell populations based on dimensionality reduction clustering and gene definition.

[0404] The existing approach is a standard two-step methodology for spatial transcriptome cell type determination:

[0405] Step 0. Control and Standardization

[0406] Input: Gene × spot / Cell UMI matrix

[0407] Filtering threshold:

[0408] n_features: 200-500 (Visiu usually 300–6,000)

[0409] n_counts: 500-0,000 (adaptive to 1–99th percentile)

[0410] mt%:<20%

[0411] Background removal: SoupX or Cellbender

[0412] Duplication detection: Scrublet / DoubletFinder (5-10% duplication rate)

[0413] Normalization:

[0414] SCTransform(vst.flavor="glmGamPoi") or LogNormalize(scale.factor=1e4) +HVG=3000 + ScaleData

[0415] Step 1. Dimensionality Reduction and Clustering

[0416] (1) Batch integration (multiple slices)

[0417] Harmony (theta=2, lambda=1, max.iter.cluster=20)

[0418] (2) PCA

[0419] HVG=3000; nPCs=30-50 (40 is commonly used)

[0420] (3) Neighbor Map

[0421] k=30, metric="euclidean"

[0422] (4) Clustering

[0423] Leiden(resolution=0.8; range 0.4–1.2)

[0424] (5) UMAP

[0425] n_neighbors=30, min_dist=0.3, spread=1.0

[0426] Step 2. Define cell types based on classic marker genes (specify parameters)

[0427] (1) Differential gene screening (DEG)

[0428] Method: Wilcoxon

[0429] log2FC ≥ 0.25 (strictly down to 0.5)

[0430] FDR <0.05

[0431] pct_in - pct_out ≥ 0.20

[0432] (2) Marker gene scoring and assignment (hard threshold)

[0433] marker library: PanglaoDB, CellMarker, Azimuth, Tabula

[0434] Scoring method:

[0435] Seurat: AddModuleScore

[0436] Scanpy: tl.score_genes / AUCell

[0437] Assignment rules:

[0438] Top 1 score z ≥ 1.0

[0439] Top1 - Top2 ≥ 0.1 - 0.2

[0440] The positive rate of core marker genes is ≥60%.

[0441] Otherwise, mark it as Mixed / Undetermined

[0442] (3) Core markers of cells

[0443] - Epithelium: EPCAM, KRT8, KRT18, KRT19

[0444] -Fiber formation: COL1A1, COL1A2, DCN, LUM

[0445] - Inner lining: PECAM1, VWF, KDR

[0446] - Smooth muscle: ACTA2, MYH11, RGS5

[0447] -Immunity: PTPRC

[0448] -T: CD3E, CD4: IL7R; CD8: CD8A; Treg: FOXP3, IL2RA

[0449] -B: MS4A1, CD79A

[0450] Mac: LYZ, CSF1R, CD68

[0451] -Neurons: RBFOX3, SNAP25

[0452] - Lesser protrusion: MBP, MOBP

[0453] -Astral collagen: GFAP, AQP4

[0454] (4) Hybrid spot processing (specific to spatial data)

[0455] - cell2location:

[0456] N_cells_per_location_prior=10

[0457] detection_alpha=20

[0458] max_epochs=30000

[0459] lr=0.002

[0460] A percentage ≥0.15–0.20% is considered the dominant cell type.

[0461] - SPOTlight: nmf_k=50, theta=0.1

[0462] - BayesSpace (Visium): q=15, gamma=3

[0463] (5) Report confidence level

[0464] -Top1 / Top2 score difference

[0465] - Expression ratio of core marker genes

[0466] - Moran's I space (p<0.05)

[0467] Reproducible key parameters:

[0468] -n_features: 300–6000

[0469] -mt%:<20%

[0470] -HVG: 3000

[0471] -PCA: nPCs=40

[0472] -KNN: k=30

[0473] -Leiden: resolution=0.8

[0474] -UMAP: min_dist=0.3, n_neighbors=30

[0475] -DEG: Wilcoxon; log2FC≥0.25; FDR<0.05

[0476] - Annotation thresholds: z≥1.0; Top1–Top2≥0.1; Core tags≥60%

[0477] - ell2location: priority=10; alpha=20; epochs=30000; lr=0.002

[0478] By comparing the above methods in the prior art ( Figure 12 ) and the cell definition method of this application ( Figure 4 The accuracy and sensitivity of the “final” (as shown in the figure) demonstrate that the scheme definition of the cell involved in this application is more sensitive and accurate.

[0479] As mentioned above, the gene marker combination, detection reagents, and cell definition method provided by this invention are applicable to the heart, contain genes validated by single-cell nuclear data, and each gene is individually designed with 6-8 probes, which can be used to specifically and sensitively define heart cells.

Claims

1. A combination of markers, characterized in that, The biomarker combination includes one or more genes selected from the following: Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Myom2, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Myh7, and Corin.

2. The marker combination as described in claim 1, characterized in that, The combination of markers includes one or more subsets of markers selected from the following: (a) Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn and Myom2; (b) Aurkb, Pole, Ect2, and Ccna2; (c) Nppa, Nppb, and Xirp2; (d)Myh7; and, (e) Corin; Preferably, the combination of markers includes Upk3b, Myh11, Kcnj8, Ptprc, Cdh5, Dcn, Myom2, Aurkb, Pole, Ect2, Ccna2, Nppa, Nppb, Xirp2, Myh7, and Corin.

3. A detection reagent, characterized in that, The detection reagent is used to detect the expression level of the biomarker combination as described in claim 1 or 2; Preferably, the detection reagent comprises primers, probes, and / or antibodies; and / or, the expression level is a protein expression level and / or an mRNA transcription level; More preferably, the probe is a circularizable DNA probe, with each end of the probe containing a specific sequence complementary to the mRNA of the target marker.

4. The detection reagent as described in claim 3, characterized in that, The detection reagent contains one or more of the following probes, for example, all of the following probes: 、 、 、 ; The ENSEMBAL number is based on the genome sequence of the Norwegian rat (Rattus norvegicus), with the genome assembly version number mRatBN7.

2.

5. A method for defining cells, characterized in that, The method includes a step of defining cardiac cells at the expression levels of the combination of markers as described in claim 1 or 2; the method is preferably performed on the Xemium platform. Preferably, the method includes: S1: Perform cell segmentation on the test sample on the Xemium platform to obtain a single-cell mask containing the cytoplasmic boundary; S2: Summarize transcripts that fall into the same mask and obtain the expression level of each marker in the marker combination as described in claim 1 or 2; S3: Based on the expression level of each marker in the marker combination, define the cells in the sample to be tested to determine the cell type and / or cell state.

6. The method as described in claim 5, characterized in that, In S3, cell type is determined by the expression level of a marker, wherein the marker is as defined in (a) of the marker combination as described in claim 2; the cell type is epicardial cells, smooth muscle cells, pericytes, immune cells, endothelial cells, fibroblasts, and / or cardiomyocytes; and / or, In step S3, the cell state is determined by the expression level of a marker, wherein the cell state is a proliferating cell or a cardiomyocyte subclass, and the marker is selected from (b)-(e) as defined in the marker combination as described in claim 2; the cardiomyocyte subclass is preferably hypertrophic cardiomyocytes, immature cardiomyocytes and / or mature cardiomyocytes.

7. The method as described in claim 6, characterized in that, In step S3, when determining the cell type, the expression level of the marker is compared with the cell type definition criteria, which are: First layer: If the cell expresses Upk3b≥3, it is identified as an epicardial cell; otherwise, proceed to the next layer. Second layer: If the cell expresses Myh11 ≥ 4, it is identified as a smooth muscle cell; otherwise, proceed to the next layer. Third layer: If the cell expression test shows Kcnj8≥4, it is determined to be a pericyte; otherwise, proceed to the next layer. Fourth layer: If the cell expresses Ptprc≥4, it is identified as an immune cell; otherwise, it proceeds to the next layer. Fifth layer: If cells express Cdh5 and / or Dcn ≥ 5, they are determined to be either endothelial cells or fibroblasts, where CdhDH5 (GeneE) > Dcn (GeneF), indicating endothelial cells; otherwise, they are fibroblasts; otherwise, proceed to the next layer; and, Sixth layer: If Myom2(GeneG) ≥ 3, then it is a cardiomyocyte; otherwise, it is an undefined cell; and / or, When determining cell state, the expression levels of the markers are compared with cell state definition criteria, which are: If cells express Aurkb + Pole + Ect2 + Ccna2 > 5, they are classified as proliferating cells; and / or, If the cell is defined as a cardiomyocyte, and the cell is determined to be a cardiomyocyte of the corresponding cell state when it meets the following conditions: when expressing Nppa + Nppb + Xirp2>20, it is determined to be a hypertrophic cardiomyocyte; or, if the cell expresses Myh7>20, it is determined to be an immature cardiomyocyte; or, if the cell expresses Corin>5, it is determined to be a mature cardiomyocyte.

8. The method according to any one of claims 5-7, characterized in that, In S1, the cell segmentation is performed using Xeniumv2.0 software to obtain the single-cell mask.

9. The method according to any one of claims 5-7, characterized in that, The step S1 includes in situ hybridization of the test sample with probes defined in the detection reagent as described in claim 3 or 4, whereby each probe is specifically linked to the mRNA molecule of the target marker to form a loop, followed by rolling circle amplification to obtain a gene-specific barcode copy. After multiple rounds of fluorescent probe hybridization, imaging, and elution, the optical barcode is decoded to obtain subcellular resolution transcript coordinates and gene identity; and / or, The step S1 includes identifying and staining the cell nuclei of the sample to be tested, and acquiring images for cell segmentation; preferably, DAPI is used as a nuclear marker, and / or, ATP1A1 and / or E-cadherin are used as membrane markers, and / or, 18S ribosomal RNA is used as a cytoplasmic marker, and / or, αSMA and / or vimentin are used as cell phenotypic markers.

10. A system for defining myocardial tissue cells in the Xemium platform, characterized in that, The system includes the following modules: An input module is used to input single-cell data of the sample to be tested, wherein the single-cell data includes expression level detection values ​​of the combination of markers as described in claim 1 or 2. An analysis module is used to analyze the test sample data to obtain analysis results; wherein, when the single-cell data meets the judgment condition, the analysis result is output as "satisfied"; when the test sample data does not meet the judgment condition, the analysis result is output as "not satisfied"; and / or, The judgment module determines the cell type and / or cell state based on the single-cell data and outputs the judgment result; wherein, when the analysis result is "satisfied", the judgment result is the corresponding cell type and / or cell state; when the analysis result is "not satisfied", the analysis module performs the next level of judgment until the output analysis result is "satisfied"; Preferably, the determination criteria are the cell type definition criteria and / or the cell state definition criteria as defined in the method of claim 7.

11. A readable medium, characterized in that, The readable medium stores a program that, when executed by a processor, enables the functionality of the system as described in claim 10.