Cell characteristic typing method and related equipment

By clustering and feature extraction of spatial transcriptome data of multiple cells, characteristic gene sets are constructed and enriched analysis, the problem of low classification accuracy of subtypes of heterogeneous disease cells is solved, and precise classification and annotation of disease cells is achieved.

CN120199338APending Publication Date: 2025-06-24BGI RES SOUTHWEST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311786839.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When subtype classification of disease cells with highly heterogeneous disease, the required number of spatial transcriptome data is huge, requiring complex data processing and analysis methods and sufficient computing resources, and it is difficult to ensure high accuracy of subtype classification.

Method used

A cell characteristic typing method is proposed. By obtaining spatial transcriptome data of multiple cells, clustering analysis is performed to determine the target cell, feature extraction is performed to obtain the target gene program and characteristic gene, construct a characteristic gene set, and determine the molecular subtype characteristics of the target gene program through enrichment analysis.

Benefits of technology

By analyzing spatial transcriptome data, the inference of molecular subtype characteristics of genetic programs of disease cells with high heterogeneity is achieved, which reduces the difficulty of data analysis and improves the accuracy and reliability of classification and annotation of disease cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199338A_ABST
    Figure CN120199338A_ABST
Patent Text Reader

Abstract

The invention relates to the field of biological medicine, and provides a cell characteristic typing method and related equipment, the method comprises the following steps: carrying out clustering analysis on a plurality of cells according to cell data of the plurality of cells, determining a plurality of target cells, and selecting target data corresponding to the plurality of target cells; performing feature extraction on the target data to obtain a target gene program in the target data, determining feature genes in the target gene program, and constructing a feature gene set of the target gene program by using the feature genes; and determining the enrichment degree of each reference gene in the reference gene set in the feature gene set, and determining the molecular subtype characteristics of the target gene program corresponding to the feature gene set according to the enrichment degree. By means of the method, the accuracy and reliability of classification and annotation of disease cells can be improved while the difficulty of data analysis of the spatial transcriptome data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technologies, and particularly to a method for cell feature classification and related equipment. Background Art

[0002] Spatial transcriptomics technology combines tissue section staining and transcriptome sequencing methods, and can obtain spatial transcriptome data containing the spatial position information and gene expression information of cells. The cell subtype classification can be determined by analyzing the spatial transcriptome data, so as to determine the corresponding disease types of cells to achieve the diagnosis, treatment or control of diseases.

[0003] However, when classifying subtypes of disease cells with high heterogeneity, the amount of spatial transcriptome data required is huge, and complex data processing and analysis methods as well as sufficient computing resources are needed, and it is difficult to ensure a high accuracy rate of subtype classification. Summary of the Invention

[0004] In view of the above, it is necessary to propose a method for cell feature classification and related equipment, which can solve the problems of low accuracy rate of subtype classification obtained based on complex spatial transcriptome data and high resource consumption.

[0005] An embodiment of this application provides a method for cell feature classification. The method includes: obtaining cell data of multiple cells, where the cell data includes the spatial transcriptome data of the multiple cells; performing clustering analysis on the multiple cells according to the cell data, determining multiple target cells from the multiple cells, and selecting target data corresponding to the multiple target cells from the cell data; extracting features from the target data to obtain a target gene program in the target data, and determining characteristic genes in the target gene program, and constructing a characteristic gene set of the target gene program by using the characteristic genes; determining the enrichment degree of each reference gene in a preset reference gene set in the characteristic gene set, and determining the molecular subtype characteristics of the target gene program corresponding to the characteristic gene set according to the enrichment degree.

[0006] In one embodiment, before performing clustering analysis on the multiple cells according to the cell data, the method further includes: preprocessing the cell data, and the preprocessing includes one or more of the following processing methods: data denoising, normalization processing, and data correction processing.

[0007] In one embodiment, the spatial transcriptome data includes a gene expression matrix for each of the multiple cells; the clustering analysis of the multiple cells based on the cell data includes: clustering and grouping the multiple cells according to a first similarity between the gene expression matrices to obtain multiple groups of cells; determining, as first cells, the cells of the first category from the multiple groups of cells according to a second similarity between each group of cells and the cells of a preset first category; determining the gene expression level of a first gene in the first cells, and taking the first cells corresponding to the gene expression levels greater than a preset threshold as the target cells, where the first gene includes genes of a preset second category.

[0008] In one embodiment, the target data includes the target spatial transcriptome data of each target cell; the feature extraction of the target data to obtain the target gene programs in the target data includes: performing a first matrix decomposition operation on the target spatial transcriptome data based on a preset algorithm to determine the functional subtype to which each target cell belongs, and determining the number of the functional subtypes corresponding to the multiple target cells, where the preset algorithm includes a consensus-based non-negative matrix factorization algorithm; based on the number of the functional subtypes, using the preset algorithm to perform a second matrix decomposition operation on the target spatial transcriptome data to obtain a first score for each gene program of each target cell and a first weight for each second gene in each gene program, where each gene program includes multiple second genes; determining a second score for each gene program of each target cell based on the first score, and taking, as the first gene programs, the gene programs that meet a preset first requirement among the multiple gene programs according to the second score; determining a second weight for each first gene program based on the first weight, and taking, as the second gene programs, the first gene programs that meet a preset second requirement according to the second weight; clustering and grouping the second gene programs to obtain multiple groups of second gene programs, aggregating each group of second gene programs into a target gene program, and determining the correspondence between the target gene program and the second gene programs.

[0009] In one embodiment, the performing a first matrix decomposition operation on the target spatial transcriptome data based on a preset algorithm to determine the functional subtype to which each target cell belongs includes: obtaining a gene pattern matrix of each target spatial transcriptome data based on the first matrix decomposition operation, and determining the functional subtype to which each target cell belongs according to the gene pattern matrix, where the gene pattern matrix is used to indicate the expression pattern of the genes of the target cell.

[0010] In one embodiment, based on the number of the functional subtypes, a second matrix decomposition operation is performed on the target spatial transcriptome data using the preset algorithm to obtain a first score of each gene program of each target cell and a first weight of each second gene in each gene program, including: obtaining a consensus matrix and a gene matrix of the gene expression program of each target cell based on the second matrix decomposition operation; determining the first score of the gene program according to the consensus matrix, where the first score is used to indicate the relative strength degree of the gene program; taking the genes in the gene program as the second genes, and determining the weight of each second gene in the corresponding gene program as the first weight according to the gene matrix.

[0011] In one embodiment, determining a second score of each gene program of each target cell based on the first score, and according to the second score, taking the gene program that meets a preset first requirement in the gene program as a first gene program, including: determining a third weight of each target cell in a target category, where the target category represents the category to which the target cell belongs; determining the second score of each gene program in all cells of the target category according to the product of the first score and the third weight; determining the average value of the second scores of all gene programs, and taking the gene programs with the second score greater than or equal to the average value as the first gene programs.

[0012] In one embodiment, determining the characteristic genes in the target gene program includes: determining a first occurrence frequency of each third gene in each target gene program, and determining a fourth gene from the third genes based on the first occurrence frequency, including: taking the first number of third genes corresponding to the first number of the first occurrence frequencies sorted in the front according to the order from large to small as the fourth genes; determining a second occurrence frequency of each fourth gene in all target gene programs, and determining the characteristic genes in the target gene program according to the second occurrence frequency, including: taking the second number of fourth genes corresponding to the second number of the second occurrence frequencies sorted in the front according to the order from large to small as the characteristic genes.

[0013] An embodiment of the present application provides a cell feature typing device, which includes: a data acquisition module for acquiring cell data of multiple cells, where the cell data includes the spatial transcriptome data of the multiple cells; a clustering analysis module for performing clustering analysis on the multiple cells according to the cell data, determining multiple target cells from the multiple cells, and selecting target data corresponding to the multiple target cells from the cell data; a feature extraction module for extracting features from the target data to obtain a target gene program in the target data, determining feature genes in the target gene program, and constructing a feature gene set of the target gene program using the feature genes; a feature classification module for determining the enrichment degree of each reference gene in a preset reference gene set in the feature gene set, and determining the molecular subtype characteristics of the target gene program corresponding to the feature gene set according to the enrichment degree.

[0014] An embodiment of the present application provides an electronic device, which includes a processor and a memory, and the processor is used to implement the cell feature typing method when executing a computer program stored in the memory.

[0015] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to implement the cell feature typing method when executed by a processor.

[0016] In summary, for the cell feature typing method described in the present application, by acquiring the spatial transcriptome data of multiple cells and performing clustering analysis on the multiple cells according to the spatial transcriptome data, target cells of a preset category are preliminarily determined from the multiple cells; by extracting features from the target data of the target cells, a target gene program of the target cells and feature genes in the target gene program are obtained, so as to obtain a feature gene set that can be used to indicate the gene characteristics of the target cells; by determining the enrichment degree of genes in the reference gene set in the feature gene set, the molecular subtype characteristics of the target gene program are determined, so as to determine the type of the target cells corresponding to the target gene program. The present application can infer the molecular subtype characteristics of the gene programs of disease cells with high heterogeneity by performing data analysis on the spatial transcriptome data, while reducing the difficulty of performing data analysis on the spatial transcriptome data and improving the accuracy and reliability of classifying and annotating disease cells. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a structural diagram of an electronic device provided by an embodiment of the present application.

[0018] Figure 2 is a flowchart of a cell feature typing method provided by an embodiment of the present application.

[0019] Figure 3 It is a flowchart of a method for determining multiple target cells provided by an embodiment of the present application.

[0020] Figure 4 It is a flowchart of a method for determining a target gene program provided by an embodiment of the present application.

[0021] Figure 5 It is a flowchart of a method for determining characteristic genes in a target gene program provided by an embodiment of the present application.

[0022] Figure 6 It is an example diagram of a KEGG pathway enrichment bubble chart of MP provided by an embodiment of the present application.

[0023] Figure 7 It is an example diagram of a bar chart of MP enrichment pathway summary provided by an embodiment of the present application.

[0024] Figure 8 It is an example diagram of a visualization heat map of GP expression correlation included in MP provided by an embodiment of the present application.

[0025] Figure 9 It is an example diagram of a Venn diagram of MP characteristic genes provided by an embodiment of the present application.

[0026] Figure 10 It is an example diagram of a scatter plot of characteristic genes included in pathway enrichment genes provided by an embodiment of the present application.

[0027] Figure 11 It is a flowchart of a refined process of step S42 provided by an embodiment of the present application.

[0028] Figure 12 It is a flowchart of a refined process of step S43 provided by an embodiment of the present application.

[0029] Figure 13 It is a flowchart of a cell characteristic typing method provided by another embodiment of the present application.

[0030] Figure 14 It is a flowchart of a method for determining a characteristic gene set provided by another embodiment of the present application.

[0031] Figure 15 It is a structural diagram of a cell characteristic typing device provided by an embodiment of the present application. Detailed implementation manners

[0032] In order to be able to more clearly understand the above objects, features and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terms used in the specification of this application are only for the purpose of describing embodiments in one embodiment, and are not intended to limit this application.

[0034] It should be noted that "at least one" in this application means one or more, and "a plurality" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0035] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0036] In one embodiment, the spatial transcriptomics technology combines tissue section staining and transcriptome sequencing methods, and can obtain spatial transcriptome data containing the spatial location information and gene expression information of cells. By analyzing the spatial transcriptome data, the subtype classification of cells can be determined, so as to determine the disease types corresponding to the cells to achieve the diagnosis, treatment or control of diseases.

[0037] However, when classifying subtypes of disease cells with high heterogeneity, the amount of spatial transcriptome data required is huge, complex data processing and analysis methods and sufficient computing resources are needed, and it is difficult to ensure a high accuracy rate of subtype classification.

[0038] For example, malignant tumors are a type of disease with high heterogeneity, and the subtype classification of their tumor cells involves the analysis of multiple samples. However, the integrated annotation analysis of a sufficient number of malignant tumor sample data is challenging, which requires considering sample batch problems and tumor consistency problems between samples, which directly affects the reliability of tumor cell classification and annotation.

[0039] To solve the above problems, an embodiment of the present application provides a method for cell feature typing, which obtains the spatial transcriptome data of multiple cells, performs clustering analysis on the multiple cells according to the spatial transcriptome data, and thus preliminarily determines target cells of a preset category from the multiple cells; by extracting features from the target data of the target cells, the target gene program of the target cells and the feature genes in the target gene program are obtained, and thus a feature gene set that can be used to indicate the gene features of the target cells is obtained; by determining the enrichment degree of the genes in the reference gene set in the feature gene set, the molecular subtype characteristics of the target gene program are determined, and thus the type of the target cells corresponding to the target gene program is determined. The present application can infer the molecular subtype characteristics of the gene programs of disease cells with high heterogeneity by performing data analysis on the spatial transcriptome data, while reducing the difficulty of performing data analysis on the spatial transcriptome data and improving the accuracy and reliability of classifying and annotating disease cells.

[0040] Figure 1 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 10 can be a mobile phone, a tablet computer, a smart wearable device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, a netbook, an energy storage device, a power distribution device, a vehicle-mounted device, a self-mobile device, and other electronic devices. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device.

[0041] As Figure 1 shown, the electronic device 10 may include a communication module 101, a memory 102, a processor 103, an input / output (I / O) interface 104, and a bus 105. The processor 103 is coupled to the communication module 101, the memory 102, and the I / O interface 104 respectively through the bus 105.

[0042] The communication module 101 may include a wired communication module and / or a wireless communication module. The wired communication module may provide one or more of the solutions for wired communication such as a universal serial bus (USB), a controller area network bus (CAN). The wireless communication module may provide one or more of the solutions for wireless communication such as wireless fidelity (Wi-Fi), Bluetooth (BT), a mobile communication network, frequency modulation (FM), near field communication (NFC), and infrared technology (IR).

[0043] The memory 102 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM). The random access memory can be directly read and written by the processor 103, and can be used to store the operating system or executable programs of other running programs (such as machine instructions), and can also be used to store user and application data, etc. The random access memory can include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc.

[0044] The non-volatile memory can also store executable programs and store user and application data, etc., and can be pre-loaded into the random access memory for direct reading and writing by the processor 103. The non-volatile memory can include disk storage devices and flash memory.

[0045] The memory 102 is used to store one or more computer programs. The one or more computer programs are configured to be executed by the processor 103. The one or more computer programs include a plurality of instructions, and when the plurality of instructions are executed by the processor 103, a cell feature typing method executable on the electronic device 10 can be implemented.

[0046] In other embodiments, the electronic device 10 further includes an external memory interface for connecting to an external memory to implement the expansion of the storage capacity of the electronic device 10.

[0047] The processor 103 may include one or more processing units. For example, the processor 103 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0048] The processor 103 provides computing and control capabilities. For example, the processor 103 is used to execute the computer program stored in the memory 102 to implement the above-mentioned cell feature typing method.

[0049] The I / O interface 104 is used to provide a channel for user input or output. For example, the I / O interface 104 can be used to connect various input and output devices, such as a mouse, a keyboard, a touch device, a display screen, etc., so that the user can enter information or visualize the information.

[0050] The bus 105 is at least used to provide a communication channel for mutual communication between the communication module 101, the memory 102, the processor 103, and the I / O interface 104 in the electronic device 10.

[0051] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 10. In other embodiments of the present application, the electronic device 10 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.

[0052] Figure 2 is a flowchart of a cell feature typing method provided by an embodiment of the present application. The cell feature typing method is applied to an electronic device, such as Figure 1 the electronic device 10 in, and specifically includes the following steps. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0053] S21, obtain cell data of multiple cells.

[0054] In one embodiment, the multiple cells may be cells belonging to the same biological tissue sample or organ sample (hereinafter simply referred to as a sample), or may be cells belonging to multiple samples. For example, the multiple cells may be multiple cells in multiple samples of glioblastoma, where glioblastoma is a malignant tumor of neural tissue.

[0055] The cell data includes the spatial transcriptome data of the multiple cells, and the spatial transcriptome data includes, but is not limited to, the gene expression matrix of each cell among the multiple cells and the spatial position coordinates of each cell among the multiple cells.

[0056] In one embodiment, spatial transcriptome sequencing technology can be used to obtain the cell data of the multiple cells. For example, the spatial transcriptome sequencing technology may be Stereo-seq (Spatial Transcriptomics through Ribonucleic Acid sequencing) technology. Specifically, the Stereo-seq technology can obtain the cell data of the multiple cells through the following process: (1) Perform image segmentation on the single-stranded Deoxyribonucleic Acid (ssDNA) image of a slice of a biological tissue sample containing multiple cells, determine the region where each cell is located in the ssDNA image, and use the coordinates of the center point position of each cell's region in the slice coordinate system as the spatial position coordinates of the corresponding cell; (2) Within the region where each cell is located, determine the nuclear range where the cell nucleus is located in each cell by identifying DNA chromosomes or performing nuclear labeling; (3) Register the nuclear range of each cell to the DNB (Digital Nucleic Acid Barcoding) expression profile to obtain the mRNA (messenger Ribonucleic Acid) transcript information related to each cell; (4) According to the registration information of the nuclear range of each cell on the DNB expression profile, segment the DNB expression profile to obtain the gene expression information of each cell, and obtain the gene expression matrix of each cell according to the gene expression information, where each element of the gene expression matrix can represent the gene expression level of a gene.

[0057] In one embodiment, the method further includes: preprocessing the cell data, and the preprocessing includes one or more of the following processing methods: data denoising, normalization processing, and data correction processing.

[0058] In one embodiment, the data denoising includes: performing quality control and filtering on the cell data, reducing the noise in the cell data and improving the signal-to-noise ratio of the cell data by operations such as removing low-quality readings in the cell data, controlling the gene expression levels of all cells within a certain range, and removing missing values in the gene expression data. Specifically, the methods used for the data denoising include, but are not limited to, one or a combination of the following methods: mean filtering, median filtering, Gaussian filtering, wavelet transform.

[0059] In one embodiment, the normalization processing includes: using data processing tools such as scanpy to normalize the cell data of each cell. Specifically, the gene expression matrix can be loaded using scanpy, and functions or methods such as scanpy.pp.normalize_per_cell() can be used to normalize the gene expression matrix, so that the scales of different gene expression matrices of different cells are consistent and comparable. In other embodiments, a logarithmic transformation method can also be used to process the gene expression matrix, thereby compressing the value range of each element in each gene expression matrix and reducing the skewness between the gene expression data.

[0060] In one embodiment, since the spatial transcriptome data may be obtained from multiple chips within a relatively long experimental period, and the users who perform the acquisition of the spatial transcriptome data for each chip may be different, and the reagent batches used by the users may also be different, data deviation may occur in the cell data due to the above objective factors. To remove this deviation, data correction processing can be performed on the cell data. For example, batch effect correction can be performed by adding data of different batches to correct possible different deviations in the sequencing process of different technical platforms, so as to perform more accurate comparison and analysis between the data.

[0061] In other embodiments, the preprocessing can also include dimensionality reduction of the cell data. For example, principal component analysis can be used to perform dimensionality reduction on the normalized cell data. Specifically, the following process can be included: performing feature extraction on the normalized gene expression matrix, performing principal component analysis on the feature set obtained by the feature extraction, obtaining multi-dimensional (for example, 50 dimensions to 200 dimensions) principal component information of the entire gene expression matrix, and retaining the principal component information of the gene expression profile, thereby realizing dimensionality reduction processing of the normalized gene expression matrix. By performing dimensionality reduction on the normalized cell data, the cell data can be mapped into a lower-dimensional space to reduce data redundancy and improve the calculation efficiency of subsequent steps; in addition, using principal component analysis to perform dimensionality reduction on the normalized gene expression matrix can also better extract the most representative gene features in the gene expression matrix to achieve feature enhancement.

[0062] In other embodiments, preprocessing the cell data further includes: performing the normalization process and the PCA dimensionality reduction process on the spatial position coordinates of each cell. In addition, when the multiple cells do not belong to the same section, preprocessing the cell data may further include: registering the three-dimensional spatial position coordinates of each cell so that the spatial position coordinates of the multiple cells are in the same coordinate system.

[0063] S22. Perform clustering analysis on the multiple cells according to the cell data, determine multiple target cells from the multiple cells, and select target data corresponding to the multiple target cells from the cell data.

[0064] In one embodiment, according to the similarity between the gene expression matrices of different cells, a clustering analysis method can be used to perform a preliminary grouping of the cells. Specifically, clustering analysis is an unsupervised learning algorithm that can be used to divide samples of unknown categories, divide them into several clusters according to certain rules, cluster samples with higher similarity in the same cluster, and divide dissimilar samples into different clusters, thereby revealing the inherent properties of the samples and the connection rules between them. The methods used in clustering analysis include but are not limited to the K-means clustering algorithm, the hierarchical clustering algorithm, etc.

[0065] In one embodiment, the gene expression of disease cells is different from that of normal cells. By observing the gene expression patterns and spatial distributions of each group of cells in the multiple cells grouped preliminarily, disease cells (such as tumor cells) can be preliminarily identified from the multiple cells.

[0066] In one embodiment, the expression levels of genes in different types of disease cells are different. By analyzing the gene expression levels of specific genes in disease cells, cells of the required target category can be determined from the disease cells as target cells. For example, by determining the cells in which the gene expression levels of genes such as EGFR, NFIB, PTPRZ1, and SOX9 in tumor cells are higher than a preset threshold, glioma cells in malignant tumor cells can be determined. The method for determining multiple target cells can also refer to the embodiments as Figure 3 shown.

[0067] In one embodiment, the target data includes the target spatial transcriptome data of each target cell. For example, the target data includes the count matrix of each target cell. The count matrix refers to the discrete count data of gene expression in different samples. This matrix shows the read counts of each gene in different samples and is used to describe the level of gene expression.

[0068] S23. Extract features from the target data to obtain the target gene program in the target data, determine the characteristic genes in the target gene program, and construct a characteristic gene set of the target gene program using the characteristic genes.

[0069] In one embodiment, feature extraction can be performed on the target data based on the consensus non - negative matrix factorization algorithm, and each gene program in each target cell can be determined according to the results of feature extraction. Subsequently, by screening the gene programs, clustering and grouping the screened gene programs, and obtaining a target gene program (Meta Program, MP) composed of multiple gene programs (GeneProgram, GP) according to the grouping results. Then, according to the occurrence frequency of genes in the target gene program, the characteristic genes in the target gene program and the characteristic gene set composed of the characteristic genes are determined.

[0070] Among them, GP describes a mode of gene regulation and expression in a specific cell. Through matrix factorization, a "consensus" file about the relative strength or usage of each GP in each cell and a "score frequency" file about the weight information of each gene in each GP can be obtained. These files provide important information for studying cell function and gene regulation. MP is a further combination and division of gene expression patterns of GP. They are obtained through multiple - sample clustering analysis and have similar characteristic genes, which are statistically obtained according to the frequency of each GP and the NMF score.

[0071] In one embodiment, the method for determining the target gene program can refer to the embodiment as Figure 4 shown. The method for determining the characteristic genes in the target gene program can refer to the embodiment as Figure 5 shown.

[0072] S24. Determine the enrichment degree of each reference gene in the preset reference gene set in the characteristic gene set, and determine the molecular subtype characteristics of the target gene program corresponding to the characteristic gene set according to the enrichment degree.

[0073] In one embodiment, the reference gene set can be obtained from an open - source database. For example, the HALLMARK gene set, c2.KEGG gene set, c5.GO gene set, and C2.Reactome gene set can be downloaded from MSigDB (Molecular Signatures Database) as 4 reference gene sets.

[0074] In one embodiment, the genes in the characteristic gene set can be compared with the genes in the reference gene set, and the enrichment degree of the genes in each reference gene set pathway in the characteristic gene set can be calculated.

[0075] For example, based on the enrichment analysis of the HALLMARK gene set, the enrichment results of gene sets related to specific biological processes and pathways in cancer samples can be obtained. Based on the enrichment analysis of the GO gene set, the significant enrichment results in three aspects of the GO genes (molecular function, biological process, and cellular component) can be obtained. Based on the enrichment analysis of the Reactome gene set pathways, the enrichment results of signals, metabolic molecules, and their organization into biological pathways and processes in the Reactome database can be obtained.

[0076] It is also possible to obtain the significant enrichment results of pathways in the KEGG gene metabolic pathways based on the enrichment analysis of the KEGG gene set pathways. For example Figure 6 As shown, it is an example diagram of the KEGG pathway enrichment bubble chart provided by the embodiment of the present application for MP. Among them, the abscissa represents the number of matches of the top 50 characteristic genes in each MP in the number of KEGG pathway genes; the ordinate represents the significantly enriched KEGG pathways, and the adjusted p-value is less than 0.05. The p-value is used to measure the difference between the genes in the KEGG pathway and the random expectation. An adjusted p-value less than 0.05 indicates that the pathway shows significant enrichment in the gene set, that is, compared with the random expectation, the frequency of genes in this pathway in the gene set is higher than expected; Figure 6 The size of the bubbles in it represents the proportion of genes. Specifically, it refers to the proportion of the number of matches of the top 50 characteristic genes in each expression program (MP) in the KEGG pathway to the number of KEGG pathway genes; Figure 6 The depth of the color in it represents the size of the adjusted p-value.

[0077] In one embodiment, the enrichment results can be visualized and summarized. For the enrichment analysis results of single cells, multiple chart types are used, including bar charts, heatmaps, Venn diagrams, scatter plots, and other methods for visualization and summary, and then the molecular subtype characteristics of each MP are defined according to the characteristics of malignant tumor cells.

[0078] Among them, since the characteristic genes of MP appear with different frequencies in different GPs, the number of characteristic genes of each MP is inconsistent. Therefore, errors will be caused due to the difference in the number of genes in the input gene set when showing the number of gene intersections. Therefore, when visualizing, the number of genes cannot be directly shown, but the enrichment ratio is shown.

[0079] For example Figure 7 As shown, it is an example diagram of the bar chart of the MP enrichment pathway summary provided by the embodiment of the present application. Among them, the abscissa represents the total number of enrichments of the characteristics of multiple GPs contained in each of the 7 MPs in a certain pathway, and the ordinate represents the top five pathways with the largest total number of genes.

[0080] For example Figure 8As shown in the figure, it is an example diagram of the visualization heatmap of the GP expression correlation included in the MP provided by the embodiment of the present application. Among them, the intersection area of the horizontal and vertical coordinates represents multiple GPs included in the MP. Pearson correlation coefficient clustering is performed according to the program.Zscore of the GPs, and the GPs with high correlation will gather together to form an MP. Specifically, the depth of the color represents the size of the Pearson correlation coefficient, and different colors represent different GPs.

[0081] For example Figure 9 As shown in the figure, it is an example diagram of the Venn diagram of the characteristic genes of the MP provided by the embodiment of the present application. The Venn diagram can intuitively show the gene intersections and unique genes among the characteristic gene sets of 7 MPs. Among them, the bar chart on the left represents the number of genes in the characteristic gene set of each MP, which is 50 in this embodiment; the bar chart above shows the number of genes in the intersection of the characteristic gene sets of each MP. Each gene intersection is formed by the combination of two or more characteristic gene sets, and the height of the bar chart represents the number of characteristic gene sets in the gene intersection; the dot matrix (matrix diagram) below represents the intersection situation among the characteristic gene sets of different MPs. Each dot represents a gene intersection, each row of the matrix represents a characteristic gene set, and each column of the matrix represents a gene intersection.

[0082] For example Figure 10 As shown in the figure, it is an example diagram of the scatter plot of the characteristic genes included in the pathway-enriched genes provided by the embodiment of the present application. Among them, the abscissa represents 50 characteristic genes of each of the 7 MPs, and the ordinate represents the characteristic gene set of the enriched pathway; the size of the dot is used to indicate the number of genes matched by the 50 characteristic genes of each MP in the enriched pathway; each MP is uniformly represented by red.

[0083] In one embodiment, as Figure 3 As shown in the figure, it is a flowchart of a method for determining multiple target cells provided by an embodiment of the present application, specifically including the following steps:

[0084] S31, according to the first similarity between the gene expression matrices, cluster and group the multiple cells to obtain multiple groups of cells.

[0085] In one embodiment, the first similarity can be determined by calculating the distance between the gene expression matrices, where the distance is inversely proportional to the first similarity, that is, the greater the distance between the gene expression matrices, the smaller the first similarity between the gene expression matrices. Among them, the distance includes but is not limited to Euclidean distance, Manhattan distance, Chebyshev distance, Minkowski distance, etc.

[0086] In one embodiment, taking the example of clustering and grouping multiple cells using K-means clustering, the specific process includes: (1) randomly selecting K initial centroids from all gene expression matrices, where K represents a preset positive integer and can be selected according to actual needs. For example, the value of K can be proportional to the total number of the multiple cells; (2) for each gene expression matrix outside the centroids, according to the distance between each gene expression matrix and each centroid, classify it into the group where the centroid with the smallest distance is located; (3) after all gene expression matrices are grouped, recalculate the centroids of each group according to the grouping situation; (4) repeat the above processes (2) and (3) until the centroids no longer change or the maximum number of iterations is reached to obtain the grouping result of the gene expression matrices.

[0087] In other embodiments, in addition to the method of clustering and grouping multiple cells according to the first similarity between gene expression matrices used in the above embodiments, the similarity between cells can also be determined by combining the spatial distance between cells. Specifically, corresponding weights can be set for the similarity of gene expression and the similarity of spatial position respectively, and the weighted sum of the similarity of gene expression and the similarity of spatial position is used as the first similarity.

[0088] S32, determine the cells of the first category as the first cells from the multiple groups of cells according to the second similarity between each group of cells and the cells of the preset first category.

[0089] In one embodiment, the cells of the first category are disease cells of a preset category, which can be selected according to actual needs. For example, the cells of the first category can be malignant tumor cells. Since the gene expression matrix of the cells of the first category is known and the characteristics of the gene expression of the cells of the first category are known, therefore, the cells of the first category can be determined as the first cells from the multiple groups of cells by determining the second similarity between each group of cells and the cells of the first category. Among them, the method for determining the second similarity is similar to the method for determining the first similarity and will not be described repeatedly.

[0090] In one embodiment, cells with a second similarity greater than a preset similarity threshold can be determined as the first cells from the multiple groups of cells. Among them, since each group of cells can contain multiple cells, the average value of the second similarities of all cells in each group of cells can be used as the value of the second similarity of the group of cells.

[0091] S33, determine the gene expression level of the first gene in the first cells, and use the first cells corresponding to the gene expression levels greater than the preset threshold as the target cells.

[0092] In one embodiment, the same disease may include multiple different subtypes. For example, according to the origin and development process of malignant tumors, they can be divided into different types and subtypes of malignant tumors, including: malignant epithelial tumors - malignant tumors originating from epithelial tissues such as lung cancer, breast cancer, gastric cancer, liver cancer, colon cancer, cervical cancer, ovarian cancer, thyroid cancer, etc.; malignant connective tissue tumors - malignant tumors originating from mesenchymal tissues, connective tissues, bones and other tissues, such as osteosarcoma, soft tissue sarcoma, etc.; malignant lymphoid tissue tumors - malignant tumors originating from lymphoid tissues, such as Hodgkin lymphoma, non-Hodgkin lymphoma, etc.; malignant vascular tissue tumors - malignant tumors originating from tissues such as vascular endothelium and vascular smooth muscle cells, such as hepatic hemangioma, vascularized nasopharyngeal carcinoma, etc.; malignant nerve tissue tumors - malignant tumors originating from nerve tissues, such as glioma, medulloblastoma, etc. Correspondingly, the same disease cells may include multiple different subtypes of cells.

[0093] In one embodiment, the first gene includes a preset second category of genes. Since the expression levels of genes in different types of disease cells are different, by analyzing the gene expression levels of the second category of genes in disease cells, cells containing the first gene with a gene expression level greater than a preset threshold can be determined as target cells. The preset threshold can be determined according to prior data, and the present application does not limit this. For example, the first gene can be genes such as EGFR, NFIB, PTPRZ1, SOX9, etc., and the target cells determined according to the first gene are glioblastoma cells.

[0094] As Figure 4 shown, it is a flowchart of a method for determining a target gene program provided by an embodiment of the present application, specifically including the following steps:

[0095] S41, perform a first matrix decomposition operation on the target spatial transcriptome data based on a preset algorithm, determine the functional subtype to which each target cell belongs, and determine the number of the functional subtypes corresponding to multiple target cells.

[0096] In one embodiment, the performing a first matrix decomposition operation on the target spatial transcriptome data based on a preset algorithm to determine the functional subtype to which each target cell belongs includes: obtaining a gene pattern matrix of each target spatial transcriptome data based on the first matrix decomposition operation, and determining the functional subtype to which each target cell belongs according to the gene pattern matrix, where the gene pattern matrix is used to indicate the expression pattern of the genes of the target cell.

[0097] In one embodiment, the preset algorithm includes a constrained Non - negative Matrix Factorization (cNMF) algorithm. The cNMF algorithm is an unsupervised learning method that can extract meaningful information from complex transcriptome data, including cell types, cell subtypes, and the relative positions of cells in tissues. The principle of the cNMF algorithm includes: through matrix factorization, the original gene expression data is decomposed into two matrices, namely the gene pattern matrix and the cell pattern matrix, so as to obtain information on cell types and spatial distributions.

[0098] Specifically, the cNMF algorithm decomposes the original gene expression data matrix into two non - negative matrices, namely the gene pattern matrix and the cell pattern matrix, by minimizing the data reconstruction error and adding constraint terms. Among them, the gene pattern matrix represents the expression pattern of genes, and the cell pattern matrix represents the spatial distribution pattern of cells. By adjusting the weight of the constraint terms, the smoothness and sparsity of the cell pattern matrix can be controlled, so as to obtain more accurate information on cell types and spatial distributions. This is very helpful for understanding the heterogeneity of malignant tumors, the functional differences of cell subtypes, and the interactions between cells. In addition, the cNMF algorithm can also provide the gene expression patterns of cell types, which helps to further study the functions and regulatory networks of different cell types.

[0099] In one embodiment, when determining the number of functional subtypes corresponding to multiple target cells, the initial number of functional subtypes of the target cells can be obtained by receiving the preliminary annotation results of the user on the functional subtypes of the target cells; the initial number can also be randomly initialized according to prior knowledge, for example, the initial number is set to any integer in [2, 10]; after determining the initial number, the cNMF algorithm can be used to perform multiple matrix factorization operations on the target spatial transcriptome data according to the initial number. Among them, after each matrix factorization operation, the stability and accuracy of the matrix obtained by the matrix factorization are determined, and the initial number is iteratively updated according to the stability and accuracy, and the updated number is used for the next matrix factorization operation until the number with the stability and accuracy meeting the preset requirements is obtained as the number of functional subtypes corresponding to multiple target cells.

[0100] In one embodiment, after identifying different cell types through the cNMF algorithm, other biological information can be combined for the interpretation and verification of cell type results. In the spatial transcriptome analysis of malignant tumor cells, known cell markers, identified cell types, and the results of other analysis methods can be used to verify the identification results of the cNMF algorithm, and the algorithm can be tested using a validation dataset. Compared with other methods, the cNMF algorithm has better performance in the classification analysis of malignant tumor cells, can explain the cell characteristics of each malignant tumor module, and helps to further study the cell heterogeneity of malignant tumors and the corresponding biological significance.

[0101] S42. Based on the number of the functional subtypes, perform a second matrix decomposition operation on the target spatial transcriptome data using the preset algorithm to obtain a first score of each gene program of each target cell and a first weight of each second gene in each gene program.

[0102] In one embodiment, by determining the first score of each gene program and the first weight of each second gene in each gene program, it is convenient to analyze the action intensity of each gene program in the cell group of the target category where the target cell is located, thereby facilitating the filtering of gene programs. The detailed process of step S42 can refer to the embodiment as Figure 11 shown.

[0103] S43. Determine a second score of each gene program of each target cell based on the first score, and based on the second score, use the gene programs that meet a preset first requirement among multiple gene programs as the first gene programs.

[0104] In one embodiment, the gene programs (GP) obtained after matrix decomposition of the target cells need to be filtered because not every GP will be used in the subsequent target gene programs (Meta Program, MP). Among them, the gene programs retained after filtering are used as the first gene programs. The detailed process of step S43 can refer to the embodiment as Figure 12 shown.

[0105] S44. Determine a second weight of each first gene program based on the first weight, and based on the second weight, use the first gene programs that meet a preset second requirement as the second gene programs.

[0106] In one embodiment, taking the first weight of the first gene program as the second weight, the first gene programs corresponding to the preset number (e.g., 50) of the second weights sorted in descending order can be used as the second gene programs meeting the second requirement. The second gene programs are, so to speak, a part of the gene programs with the highest usage rate among all cells of the target category selected by further filtering and screening the first programs.

[0107] S45. Cluster and group the second gene programs to obtain multiple groups of second gene programs, aggregate each group of second gene programs into a target gene program, and determine the corresponding relationship between the target gene program and the second gene program.

[0108] In one embodiment, the methods that can be used for clustering and grouping the second gene programs include but are not limited to the Wilks clustering method. Specifically, each second gene program includes multiple fifth genes. It is possible to determine the intersection and union between two sets of fifth genes included in every two fifth gene programs, use the ratio of the number of fifth genes in the intersection to the number of fifth genes in the union as the Jaccard coefficient between the two second gene programs, and perform Wilks clustering on the second gene programs according to the Jaccard coefficient.

[0109] In one embodiment, according to the grouping result obtained by Wilks clustering, each group of second gene programs is regarded as a target gene program. Multiple target gene programs may be obtained, so the inclusion relationship between each target gene program and the second gene program can be recorded.

[0110] Reference Figure 5 shown in Figure 5 is a flowchart of a method for determining characteristic genes in a target gene program provided by an embodiment of the present application. The method for determining characteristic genes in a target gene program includes:

[0111] S51. Determine the first occurrence frequency of each third gene in each target gene program, and determine the fourth gene from the third genes based on the first occurrence frequency.

[0112] In one embodiment, the first quantity of third genes corresponding to the first occurrence frequencies sorted in the top preset first quantity (e.g., 50) in descending order of the first occurrence frequencies can be used as the fourth genes.

[0113] In one embodiment, the fourth gene represents a gene with a relatively high occurrence frequency in each target gene program. The fourth genes in each target gene program can be retained to obtain a set of all fourth genes of all target gene programs. For example, if there are m target gene programs and each target gene program includes n fourth genes with relatively high occurrence frequencies, then the set of all fourth genes of all target gene programs includes a total of m × n genes. Among them, the same fourth gene may be included in different target gene programs.

[0114] S52. Determine the second occurrence frequency of each fourth gene in all target gene programs, and determine the characteristic genes in the target gene program according to the second occurrence frequency.

[0115] In one embodiment, according to the order from large to small, the second quantity of fourth genes corresponding to the second occurrence frequencies ranked among the top preset second quantity in the second occurrence frequencies can be used as the characteristic genes.

[0116] In one embodiment, the characteristic genes represent genes with relatively high occurrence frequencies in all target gene programs and are the genes that can best reflect the characteristics of the target gene programs. Therefore, using the characteristic genes to analyze the molecular subtype characteristics of the target gene programs can improve the accuracy of classification.

[0117] In one embodiment, as shown in Figure 11 the detailed process of step S42 includes:

[0118] S61. Based on the second matrix decomposition operation, obtain the consensus matrix and gene matrix of the gene expression program of each target cell.

[0119] In one embodiment, based on the number of functional subtypes, the cNMF algorithm can be used to perform matrix decomposition operation on the target spatial transcriptome data to obtain the "consensus" file of the consensus matrix of cell × expression program of each target cell and the "score frequency" file of the gene matrix of expression program × gene.

[0120] Among them, the "consensus" file includes the first score of each gene program (Gene Program, GP) of each target cell, and the first score is used to indicate the relative strength or usage, utilization rate of each GP in each target cell. Each gene program includes multiple second genes; the "score frequency" file includes the first weight of each second gene in each GP, and the first weight is used to indicate the weight of each gene in the corresponding GP. Among them, the gene program is also called the expression program.

[0121] S62. Determine the first score of the gene program according to the consensus matrix, and the first score is used to indicate the relative strength degree of the gene program.

[0122] S63. Take the genes in the gene program as the second genes, and determine the weight of each second gene in the corresponding gene program as the first weight according to the gene matrix.

[0123] Reference Figure 12 As shown, the refinement process of step S43 includes:

[0124] S71. Determine the third weight of each target cell in the target category.

[0125] In one embodiment, since the target cells of the target category are a group of cells obtained by gene screening after clustering analysis of all cells, the third weight of each target cell in the target category can be determined according to the distance of gene expression between each target cell and the cell corresponding to the centroid in this group of cells. Specifically, the closer the distance between the target cell and the cell corresponding to the centroid in this group of cells, the greater the corresponding third weight.

[0126] S72. Determine the second score of each gene program in all cells of the target category according to the product of the first score and the third weight.

[0127] In one embodiment, the first score represents the usage rate of each gene program in each target cell, and the product of the first score and the third weight can be used to indicate the usage rate of each gene program in all cells of the target category.

[0128] S73. Determine the average value of the second scores of all gene programs, and take the gene programs with the second score greater than or equal to the average value as the first gene programs.

[0129] In one embodiment, gene programs with the second score greater than or equal to the average value can be regarded as gene programs with a relatively high usage rate in all cells of the target category. Taking the gene programs with the second score greater than or equal to the average value as the first gene programs can filter out gene programs with a relatively small usage rate. The first requirement includes that the second score is greater than or equal to the average value.

[0130] The cell feature typing method provided by the embodiments of the present application obtains the spatial transcriptome data of multiple cells, performs clustering analysis on the multiple cells according to the spatial transcriptome data, and thus preliminarily determines target cells of a preset category from the multiple cells; by extracting features from the target data of the target cells, the target gene program of the target cells and the characteristic genes in the target gene program are obtained, and thus a characteristic gene set that can be used to indicate the gene characteristics of the target cells is obtained; by determining the enrichment degree of the genes in the reference gene set in the characteristic gene set, the molecular subtype characteristics of the target gene program are determined, and thus the type of the target cells corresponding to the target gene program is determined. The present application can infer the molecular subtype characteristics of the gene programs of disease cells with high heterogeneity by performing data analysis on the spatial transcriptome data, while reducing the difficulty of performing data analysis on the spatial transcriptome data and improving the accuracy and reliability of classifying and annotating disease cells.

[0131] In one embodiment, for example Figure 13 As shown, it is a flowchart of the cell feature typing method provided by another embodiment of the present application. The method provided by the embodiments of the present application first performs spatial transcriptome sequencing and analysis on the collected glioma samples through the Stereo-seq technology; then performs preprocessing such as basic data quality control filtering and standardization on each sample of the output glioma data; then performs dimensionality reduction clustering and preliminary annotation on it to determine the cells with malignant tumor characteristics in each sample and extract the counts matrix of the malignant tumor cells; performs cNMF analysis on the counts matrix of the malignant tumor cells, calculates and confirms the N value of each sample through multiple iterations, where the N value represents the number of types of malignant tumor cells contained in each sample, and obtains the most representative genes that can represent each GP; calculates the Jaccard coefficients between the top 50 most representative genes of each GP, clusters the tumor GPs through the correlation of the Jaccard coefficients, and divides the tumor GPs into multiple MPs; finally, the final MP characteristic gene set is confirmed by the frequency method, and the functional definition of the MP characteristic gene set is performed through enrichment analysis.

[0132] In one embodiment, for example Figure 14 As shown, it is a flowchart of the method for determining the characteristic gene set provided by another embodiment of the present application. The method provided by the embodiments of the present application extracts the characteristics of the glioma MP based on the cNMF algorithm. The specific process includes: first confirming the number N value of tumor subtypes in the sample; then using the cNMF algorithm to perform matrix decomposition on the cell data of the tumor cells to obtain the GPs of the tumor cells; then filtering the GPs by calculating the usage ratio of the GPs in the tumor cells; then calculating the Jaccard coefficients between the filtered GPs; clustering the tumor GPs through the correlation of the Jaccard coefficients, and dividing the tumor GPs into multiple MPs; finally, the MP characteristic gene set is confirmed by the frequency method.

[0133] In other embodiments, other unsupervised feature extraction methods similar to the cNMF algorithm can also be used to replace the cNMF algorithm. For example, Independent Component Analysis (ICA), Non-negative Matrix Factorization-Cube Root (NMF-CubeRoot), Gaussian Mixture Model (GMM), Self-Organizing Map (SOM), etc.; these methods can all be used for unsupervised learning tasks and can decompose and cluster malignant tumor cells of different samples into GPs. They may have different advantages and applicability in different situations, so algorithms for other cell type decomposition tasks can be selected according to needs to replace the cNMF algorithm.

[0134] In other embodiments, the method provided in the above embodiments can also be performed on other non-glioma malignant tumor samples, such as spatial transcriptome tumor data with high tumor heterogeneity. The tumor feature recognition process of this process can also be considered and replaced with an algorithm for other cell type decomposition tasks.

[0135] Figure 15 It is a structural diagram of a cell feature typing device provided by an embodiment of the present application.

[0136] In some embodiments, the cell feature typing device 120 may include a plurality of functional modules composed of computer program segments. The computer programs of each program segment in the cell feature typing device 120 can be stored in the memory of the electronic device and executed by at least one processor to perform the function of cell feature typing (see details in Figure 2 description).

[0137] In this embodiment, the cell feature typing device 120 can be divided into a plurality of functional modules according to the functions it performs. The functional modules may include: a data acquisition module 1201, a clustering analysis module 1202, a feature extraction module 1203, and a feature classification module 1204. The module referred to in this application means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, for the implementation manners of the functions of each module in the cell feature typing device 120, reference can be made to the definition of the cell feature typing method above, and the description will not be repeated here.

[0138] The data acquisition module 1201 is used to acquire cell data of a plurality of cells, and the cell data includes spatial transcriptome data of the plurality of cells;

[0139] The clustering analysis module 1202 is configured to perform clustering analysis on the multiple cells according to the cell data, determine multiple target cells from the multiple cells, and select target data corresponding to the multiple target cells from the cell data;

[0140] The feature extraction module 1203 is configured to extract features from the target data, obtain a target gene program in the target data, determine feature genes in the target gene program, and construct a feature gene set of the target gene program by using the feature genes;

[0141] The feature classification module 1204 is configured to determine the enrichment degree of each reference gene in a preset reference gene set in the feature gene set, and determine the molecular subtype characteristics of the target gene program corresponding to the feature gene set according to the enrichment degree.

[0142] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program includes program instructions, and the method implemented when the program instructions are executed may refer to the methods in the above various embodiments of the present application.

[0143] Among them, the computer-readable storage medium may be an internal memory of the electronic device described in the above embodiment, such as a hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device.

[0144] In some embodiments, the computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the electronic device.

[0145] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0146] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0147] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical or other forms.

[0148] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A method for cell characteristic typing, characterized in that, The method includes: Obtaining cell data of multiple cells, where the cell data includes spatial transcriptome data of the multiple cells; Performing clustering analysis on the multiple cells according to the cell data, determining multiple target cells from the multiple cells, and selecting target data corresponding to the multiple target cells from the cell data; Performing feature extraction on the target data to obtain a target gene program in the target data, determining characteristic genes in the target gene program, and constructing a characteristic gene set of the target gene program by using the characteristic genes; Determining the enrichment degree of each reference gene in a preset reference gene set in the characteristic gene set, and determining the molecular subtype characteristics of the target gene program corresponding to the characteristic gene set according to the enrichment degree.

2. The cell characteristic typing method according to claim 1, wherein Before performing clustering analysis on the multiple cells according to the cell data, the method further includes: Performing preprocessing on the cell data, where the preprocessing includes one or more of the following processing methods: data denoising, normalization processing, and data correction processing.

3. The cell characteristic typing method according to claim 1, wherein The spatial transcriptome data includes a gene expression matrix of each cell in the multiple cells; The performing clustering analysis on the multiple cells according to the cell data includes: Performing clustering grouping on the multiple cells according to a first similarity between the gene expression matrices to obtain multiple groups of cells; Determining the cells of the first category as first cells from the multiple groups of cells according to a second similarity between each group of cells and the cells of a preset first category; Determining the gene expression level of a first gene in the first cells, and taking the first cells corresponding to the gene expression levels greater than a preset threshold as the target cells, where the first gene includes genes of a preset second category.

4. The cell characteristic typing method according to claim 1, wherein The target data includes target spatial transcriptome data of each target cell; The performing feature extraction on the target data to obtain a target gene program in the target data includes: Performing a first matrix decomposition operation on the target spatial transcriptome data based on a preset algorithm to determine the functional subtype to which each target cell belongs, and determining the number of the functional subtypes corresponding to the multiple target cells, where the preset algorithm includes a consensus-based non-negative matrix decomposition algorithm; Based on the number of the functional subtypes, using the preset algorithm to perform a second matrix decomposition operation on the target spatial transcriptome data to obtain a first score of each gene program of each target cell and a first weight of each second gene in each gene program, where each gene program includes multiple second genes; Determining a second score of each gene program of each target cell based on the first score, and taking the gene programs that meet a preset first requirement among the multiple gene programs as the first gene program according to the second score; Determining a second weight of each first gene program according to the first weight, and taking the first gene programs that meet a preset second requirement as the second gene program according to the second weight; Cluster and group the second gene programs to obtain multiple groups of second gene programs, aggregate each group of second gene programs into a target gene program, and determine the correspondence between the target gene program and the second gene program.

5. The cell characteristic typing method according to claim 4, wherein The first matrix decomposition operation is performed on the target spatial transcriptome data based on a preset algorithm to determine the functional subtype to which each target cell belongs, including: Based on the first matrix decomposition operation, a gene pattern matrix of each target spatial transcriptome data is obtained, and the functional subtype to which each target cell belongs is determined according to the gene pattern matrix, where the gene pattern matrix is used to indicate the expression pattern of the genes of the target cell.

6. The cell characteristic typing method according to claim 4, wherein Based on the number of the functional subtypes, the second matrix decomposition operation is performed on the target spatial transcriptome data using the preset algorithm to obtain the first score of each gene program of each target cell and the first weight of each second gene in each gene program, including: Based on the second matrix decomposition operation, a consensus matrix and a gene matrix of the gene expression program of each target cell are obtained; The first score of the gene program is determined according to the consensus matrix, and the first score is used to indicate the relative strength degree of the gene program; The genes in the gene program are used as the second genes, and the weight of each second gene in the corresponding gene program is determined according to the gene matrix as the first weight.

7. The cell characteristic typing method according to claim 4, wherein The second score of each gene program of each target cell is determined based on the first score, and according to the second score, the gene programs that meet the preset first requirement in the gene program are used as the first gene programs, including: Determine the third weight of each target cell in the target category, where the target category represents the category to which the target cell belongs; According to the product of the first score and the third weight, determine the second score of each gene program in all cells of the target category; Determine the average value of the second scores of all gene programs, and use the gene programs with the second score greater than or equal to the average value as the first gene programs.

8. The cell characteristic typing method according to claim 1, wherein Determining the characteristic genes in the target gene program includes: Determine the first occurrence frequency of each third gene in each target gene program, and determine the fourth gene from the third genes based on the first occurrence frequency, including: according to the order from large to small, use the first number of third genes corresponding to the first number of the first occurrence frequencies sorted in the front in the first occurrence frequencies as the fourth gene; Determine the second occurrence frequency of each fourth gene in all target gene programs, and determine the characteristic genes in the target gene program according to the second occurrence frequency, including: according to the order from large to small, use the second number of fourth genes corresponding to the second number of the second occurrence frequencies sorted in the front in the second occurrence frequencies as the characteristic genes.

9. A cell characteristic typing device, characterized in that, The device includes: A data acquisition module for acquiring cell data of multiple cells, where the cell data includes the spatial transcriptome data of the multiple cells; A clustering analysis module, configured to perform clustering analysis on the multiple cells according to the cell data, determine multiple target cells from the multiple cells, and select target data corresponding to the multiple target cells from the cell data; A feature extraction module, configured to extract features from the target data, obtain a target gene program in the target data, determine feature genes in the target gene program, and construct a feature gene set of the target gene program by using the feature genes; A feature classification module, configured to determine the enrichment degree of each reference gene in a preset reference gene set in the feature gene set, and determine the molecular subtype characteristics of the target gene program corresponding to the feature gene set according to the enrichment degree.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory. When the processor executes a computer program stored in the memory, the cell feature typing method according to any one of claims 1 to 8 is implemented.