Subcellular level space transcriptome function analysis method and system
By developing a multi-module automation process, the shortcomings of existing tools in processing high-resolution subcellular-level spatial transcriptome data are solved, and more efficient and accurate analysis is achieved, suitable for multi-sample and big data scenarios.
Patent Information
- Application Number
- CN202510275928.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
Existing spatial transcriptome data analysis tools lack the ability to process high-resolution subcellular data, and lack coherence among tools, affecting analysis efficiency and accuracy.
Develop a multi-module automation process, including automated processing of sample information, optimizing spatial transcriptome expression matrix data, performing gene ID conversion and cell unit type selection, statistical and visualizing spatial data characteristics, performing standardization and clustering analysis, and implementing subcellular-level spatial transcriptome functional analysis.
It significantly improves the analysis efficiency, accuracy and biological interpretability of spatial transcriptomics and single-cell data, and is suitable for multi-sample, high-resolution and big data analysis scenarios.
Smart Images

Figure CN120220805A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular, to a method and system for subcellular-level spatial transcriptome functional analysis. Background Art
[0002] With the rapid development of biotechnology, spatial transcriptomics, as an emerging field that studies the gene expression patterns of cells in spatial positions and their relationships with biological functions, is gradually becoming a hot topic in life science research. Spatial transcriptome data not only contains the gene expression information of cells but also incorporates spatial position information, providing a unique perspective for in-depth understanding of cell-cell interactions, tissue structure, and disease mechanisms.
[0003] The existing spatial transcriptome data analysis processes and technical tools generally have the following main problems: Existing tools generally lack dedicated modules for processing high-resolution spatial transcriptome data, which limits the ability to analyze subcellular-level data and the understanding of dynamic changes in cell states. Different tools have limited support for data from specific technical platforms. For example, Squidpy lacks support for BGI Stereo data, which restricts its application scope. Some tools are slow in processing large-scale spatial transcriptome data and lack sufficient consideration of spatial positions, affecting the analysis efficiency. There is a lack of coherence in upstream and downstream analysis between tools, especially during the data processing connection process, which affects the fluency of analysis and the accuracy of results. Existing tools have limitations in data output and statistics. The lack of example statements for data table text output affects the further analysis and application of data. At the same time, the lack of statistical and calculation methods for cell proportions limits the understanding of cell composition and function. Some tools are not meticulous enough in the selection of bin cell sizes, with too large a precision span, affecting the fineness of analysis. Existing tools lack an automated bin selection function, etc., which increases the operation difficulty for users and limits the application of automated analysis. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and system for subcellular-level spatial transcriptome functional analysis, which significantly improves the analysis efficiency, accuracy, and biological interpretability of spatial transcriptomics and single-cell data through a multi-module automated process, and is applicable to multi-sample, high-resolution, and big data analysis scenarios.
[0005] To solve the above technical problem, the technical solution of the present invention is as follows:
[0006] In a first aspect, a method for subcellular-level spatial transcriptome functional analysis, the method includes:
[0007] Automatically process the core information of the sample and optimize the spatial transcriptome expression matrix data to achieve the reception and processing of the expression matrix data;
[0008] Add functions such as timed polling to automate the conversion from spatial gene expression matrix to standard spatial transcriptome analysis. According to the expression matrix data, perform gene ID conversion, cell unit type and cell unit selection to obtain the selected data;
[0009] According to the converted and selected data, calculate the spatial data features and visualize the data features;
[0010] Standardize the data features and automatically select a clustering algorithm to cluster the dimensionality-reduced data to obtain the clustered data;
[0011] According to the clustered data, implement subcellular-level spatial transcriptome functional analysis, including automated cell type annotation, gene function enrichment analysis, and cell proportion analysis.
[0012] Furthermore, automatically process the core information of the samples and optimize the spatial transcriptome expression matrix data to implement the reception and processing of expression matrix data, including:
[0013] According to the sample results, automatically fill in the core information of the samples, including species, tissue type, kit version, and sequencing type, to implement sample management and the reception and processing of expression matrix data;
[0014] By automatically processing the core information of the samples, receive the spatial transcriptome expression matrix data and add the statistical analysis results of the SAW expression matrix to obtain the multi-sample statistical results;
[0015] According to the multi-sample statistical results, provide 15 subcellular statistical gradients from bin20, 50, 200 extended to bin1 to bin200, and automatically generate the corresponding cell diameter size ranges.
[0016] Furthermore, according to the expression matrix data, perform gene ID conversion, cell unit type and cell unit selection to obtain the selected data, including:
[0017] Add functions such as timed polling to automate the conversion from spatial gene expression matrix to standard spatial transcriptome analysis, receive the spatial transcriptome expression matrix data, and automatically convert ensembl to gene symbol to implement gene ID conversion;
[0018] Obtain the predefined analysis requirements, and automatically select cell bin or square bin based on the analysis requirements, and dynamically adjust the bin size to obtain the converted and selected data.
[0019] Furthermore, according to the converted and selected data, calculate the spatial data features and visualize the data features, including:
[0020] According to the data after conversion and selection, calculate the proportion of spatial cells and the expression of spatially differential genes, automatically select highly variable genes and characteristic genes, and obtain the spatial data characteristics;
[0021] Based on the spatial data characteristics, generate UMAP and spatial distribution maps to visually display the characteristics of spatial transcriptomics and achieve data feature visualization.
[0022] Furthermore, perform automated standardization on the data features, automatically select a clustering algorithm for clustering after dimensionality reduction to obtain clustering data, including:
[0023] Perform automated standardization on the data features to obtain standardized data, and then perform dimensionality reduction;
[0024] According to the data after dimensionality reduction, automatically select neighbors and spatial_neighbors to construct a neighborhood graph;
[0025] According to the constructed neighborhood graph, automatically select a clustering algorithm to cluster the data to obtain clustering data.
[0026] Furthermore, according to the clustered data, perform subcellular-level spatial transcriptome functional analysis. The analysis includes automated cell type annotation, gene function enrichment analysis, and cell proportion analysis, including:
[0027] According to the clustered data, perform cell spatial GO / KEGG / Reactom enrichment analysis on the automated cell type annotation to achieve subcellular-level spatial transcriptome functional analysis. The basic analysis includes automated cell type annotation, gene function enrichment analysis, and cell proportion analysis.
[0028] In the second aspect, a system for subcellular-level spatial transcriptome functional analysis includes:
[0029] A subcellular and expression matrix optimization module for automatically processing the core information of samples and optimizing the statistical analysis of spatial transcriptome expression matrix data to achieve the reception and processing of expression matrix data;
[0030] Add automated implementation of spatial gene expression matrix to spatial transcriptome standard analysis such as timed polling. A gene dictionary and cell unit selection module for performing gene ID conversion, cell unit type and cell unit size selection based on the expression matrix data to obtain the selected data;
[0031] A spatial data feature visualization module for calculating spatial data features based on the converted and selected data and visualizing the data features;
[0032] A dimensionality reduction and clustering automatic selection module, which is used to automatically standardize data features, and automatically select a clustering algorithm to cluster the data after dimensionality reduction to obtain clustered data;
[0033] A cell annotation module, which is used to perform subcellular-level spatial transcriptome functional analysis based on the clustered data. The analysis includes automatic cell type annotation, differential expression analysis between cell types, functional enrichment analysis of differential genes, and cell proportion analysis.
[0034] In a third aspect, a computing device includes:
[0035] One or more processors;
[0036] A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method.
[0037] In a fourth aspect, a computer-readable storage medium stores a program, which when executed by a processor implements the method.
[0038] The above solution of the present invention has at least the following beneficial effects:
[0039] By developing a special algorithm to process high-resolution subcellular-level spatial transcriptome expression matrix data, supporting 15 subcellular statistical gradients from bin1 to bin200, and automatically generating corresponding size ranges in combination with the cell diameter module, the present invention realizes refined analysis of subcellular-level data and significantly improves the accuracy of data analysis. The gene dictionary module automatically converts ensembl gene identifiers into common gene names, and automatically selects the cell unit type and cell unit size according to the actual situation of the data, improving the readability and automation of biological data and facilitating researchers to understand and interpret analysis results. Integrating a variety of standardization, dimensionality reduction, and clustering methods, and realizing the intelligence of dimensionality reduction and clustering through automatic selection, avoiding deviations caused by improper algorithm selection, and significantly improving the analysis efficiency. By adding a gene dictionary module and a cell unit selection module, the present invention optimizes the data processing flow, enhances the compatibility and connection with tools such as Saw, and improves the coherence of analysis.
[0040] The present invention provides richer data output options, including spatial cell proportion statistics, differential gene expression statistics, gene spatial visualization, and text table output, etc., facilitating further analysis and application of data. The present invention integrates functions such as automatic dimensionality reduction, clustering, and cell annotation, reducing the operation complexity of users and improving the automation degree of analysis. Description of the Drawings
[0041] Figure 1It is a schematic flowchart of a method for subcellular-level spatial transcriptome functional analysis provided by an embodiment of the present invention.
[0042] Figure 2 It is a schematic diagram of the principle of the subcellular space bin statistics and expression matrix statistics module of a system for subcellular-level spatial transcriptome functional analysis provided by an embodiment of the present invention.
[0043] Figure 3 It is a schematic diagram of the principle of the gene dictionary module and cell unit selection module of a system for subcellular-level spatial transcriptome functional analysis provided by an embodiment of the present invention.
[0044] Figure 4 It is a schematic diagram of the principle of the spatial cell proportion and gene spatial visualization module of a system for subcellular-level spatial transcriptome functional analysis provided by an embodiment of the present invention.
[0045] Figure 5 It is a schematic diagram of the principle of the dimensionality reduction clustering automatic selection module of a system for subcellular-level spatial transcriptome functional analysis provided by an embodiment of the present invention.
[0046] Figure 6 It is a schematic diagram of the principle of a system for subcellular-level spatial transcriptome cell annotation and functional analysis provided by an embodiment of the present invention. Detailed implementation manners
[0047] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0048] As Figure 1 shown, an embodiment of the present invention provides a method for subcellular-level spatial transcriptome functional analysis, and the method includes the following steps:
[0049] Step 11, automatically process the core information of the sample and optimize the spatial transcriptome expression matrix data to implement the reception and processing of the expression matrix data;
[0050] Step 12, add a timed poll, etc. to implement the automation of the spatial gene expression matrix to the standard analysis of the spatial transcriptome; according to the expression matrix data, perform gene ID conversion and selection of the cell unit type and cell unit size to obtain the selected data;
[0051] Step 13, according to the converted and selected data, statistically analyze the spatial data features and visualize the data features;
[0052] Step 14: Standardize the data features and automatically select a clustering algorithm to cluster the dimension-reduced data to obtain clustered data;
[0053] Step 15: Based on the clustered data, perform subcellular-level spatial transcriptome functional analysis, including automated cell type annotation, gene function enrichment analysis, and cell proportion analysis.
[0054] In the embodiment of the present invention, by receiving spatial transcriptome expression matrix data, performing gene dictionary conversion and unit selection, converting ensembl gene identifiers into common gene symbols, the readability of the data is improved. At the same time, cell bin or square bin is automatically selected according to the analysis requirements, and the bin size is dynamically adjusted to optimize the accuracy of data analysis. By statistically analyzing and visualizing spatial data features, such as generating spatial cell proportion maps, differential gene expression maps, etc., the characteristics of spatial transcriptomics are intuitively displayed. Through automated dimensionality reduction and clustering analysis, the processing process of high-dimensional data is simplified, and a suitable dimensionality reduction method and clustering algorithm are automatically selected according to the data characteristics, improving the efficiency and accuracy of the analysis, and helping to identify different cell types or spatial regions. Subcellular-level spatial transcriptome functional analysis is realized through analyses such as automated cell type annotation, differential expression analysis between cell types, and differential enrichment analysis.
[0055] In a preferred embodiment of the present invention, the above step 11 may include:
[0056] Step 111: Automatically fill in the core information of the sample according to the sample results, including species, tissue type, kit version, and sequencing type, to realize sample management and the reception and processing of expression matrix data.
[0057] Step 112: Receive spatial transcriptome expression matrix data by automatically processing the core information of the sample, and add the statistical analysis results of the SAW expression matrix to obtain multi-sample statistical results;
[0058] Step 113: According to the multi-sample statistical results, provide 15 subcellular statistical gradients from bin20, 50, 200 to bin1 to bin200, and automatically generate the corresponding cell diameter size ranges;
[0059] In the embodiments of the present invention, the core information of the samples is automatically filled, including species, tissue type, kit version, sequencing type, etc. This process realizes the standardized and automated management of sample information, reduces manual input errors and omissions, and improves the efficiency and accuracy of data processing. By receiving spatial transcriptome expression matrix data and adding the statistical analysis results of the SAW expression matrix, the integration of multi-sample data is achieved, which not only enriches the content of the data set but also provides cross-sample statistical analysis information. According to the multi-sample statistical results, 15 subcellular statistical gradients from bin20, 50, 200 to bin1 to bin200 are provided. The gradients represent different spatial resolutions, enabling researchers to analyze subcellular-level data more precisely. At the same time, the corresponding cell diameter size range is automatically generated, further enhancing the readability and visualization effect of the data.
[0060] In the embodiments of the present invention, the specific steps include:
[0061] Step 111: Automatically process the core information of the samples, including organisms, tissue types, kit versions, sequencing types, etc. These information are usually stored as additional columns or metadata in the expression matrix data and are automatically filled into the corresponding data structures according to the extracted sample information. By automatically filling the sample information, sample management functions are realized, including operations such as storage, query, update, and deletion of sample information. After the above steps, the reception and processing of the spatial transcriptome expression matrix data are completed, including data integration, setting of subcellular statistical gradients, generation of cell diameter size ranges, and automatic filling of sample information.
[0062] Step 112: Receive spatial transcriptome expression matrix data from different sources. The data usually exists in matrix form, where rows represent genes and columns represent the expression levels at spatial positions (such as cells or spatial regions). Read the statistical analysis results of the spatial transcriptome expression matrix generated by the SAW tool, including statistical summaries of gene expression levels and spatial distribution characteristics. According to the statistical analysis results of SAW and the received expression matrix data, directly add the statistical analysis results as additional columns to the expression matrix for integration. By integrating the original expression matrix data and the statistical analysis results of SAW, multi-sample statistical results are generated.
[0063] Step 113, the multi-sample statistical results are improved from the Tissuesquare binstatistics of bins 20, 50, and 200 to 15 subcellular statistical gradients of bins 1, 3, 5, 10, 12, 20, 24, 30, 40, 50, 60, 80, 100, 150, and 200. Corresponding cell size dimensions of 0.7, 2, 3.5, 7, 8, 14, 16, 24, 28, 35, 42, 56, 70, 105, and 140 microns are added according to the subcellular statistical gradients.
[0064] In a preferred embodiment of the present invention, the above step 12 may include:
[0065] Step 121, add functions such as timed polling to automate the conversion from the spatial gene expression matrix to the standard analysis of spatial transcriptomics. Receive the spatial transcriptome expression matrix data, automatically convert ensembl to gene symbol, and achieve gene ID conversion;
[0066] Step 122, obtain predefined analysis requirements, and automatically select cell bin or square bin based on the analysis requirements, dynamically adjust the bin size to obtain the data after conversion and selection.
[0067] In the embodiments of the present invention, by automatically converting ensembl to gene symbol, the gene names in the spatial transcriptome expression matrix data become more intuitive and easier to understand, improving the readability of the data. By obtaining predefined analysis requirements and automatically selecting cell bin or square bin as the analysis unit based on these requirements, while dynamically adjusting the cell bin size to optimize the accuracy of data analysis, this automated unit selection process reduces the manual operations of users and improves the efficiency and accuracy of data analysis. By selecting the appropriate cell analysis unit type and the bin size of the cell unit, researchers can more precisely analyze the spatial transcriptome data and reveal the interactions between cells and the spatial patterns of gene expression. By integrating the gene dictionary conversion and cell unit selection functions, the method of the present invention provides a more flexible and adaptable data analysis framework.
[0068] In the embodiments of the present invention, the specific steps include:
[0069] Step 121: Add functions such as timed polling to automate the conversion from the spatial gene expression matrix to the standard analysis of spatial transcriptomics. Receive the spatial transcriptomics expression matrix data, where rows represent genes and columns represent the expression levels at spatial positions (such as cells or spatial regions). Gene identifiers are usually stored in the ensembl format in the expression matrix. Load a pre-constructed or obtained ensembl gene dictionary. The gene dictionary is a mapping table that maps ensembl gene identifiers to common gene names, such as gene symbol. Traverse each gene identifier in the expression matrix, use the ensembl gene dictionary to convert it to the corresponding gene symbol. During the conversion process, record the conversion results for each gene and replace the original genes in the expression matrix.
[0070] Step 122: Obtain the analysis requirements from user input, configuration files, or predefined analysis workflows. The requirements include the analysis objectives, such as cell type identification and gene expression pattern analysis; the precision requirements of the analysis, such as subcellular level and tissue level; and the output format of the analysis. Parse the obtained analysis requirements to determine the required analysis unit type (cell bin or square bin) and the desired data analysis precision. According to the parsed analysis requirements, automatically select cell bin or square bin as the analysis unit. Dynamically adjust the size of the bin according to the selected analysis unit type and the desired data analysis precision. For example, if the analysis requirement is to finely analyze subcellular-level data, select a smaller bin size; if more attention is paid to a larger spatial distribution, select a larger bin size. Apply the selected analysis unit and bin size to the expression matrix data. For each spatial position (such as cells or spatial regions), classify it into the corresponding analysis unit according to the bin size it belongs to, and output the data after conversion and selection, where the gene names have been converted to gene symbols, and the appropriate analysis unit and bin size have been selected according to the analysis requirements.
[0071] In a preferred embodiment of the present invention, the above step 13 may include:
[0072] Step 131: Add functions such as timed polling to automate the conversion from the spatial gene expression matrix to the standard analysis of spatial transcriptomics. According to the data after conversion and selection, calculate the spatial cell proportion and spatial differential gene expression, automatically select highly variable genes and characteristic genes to obtain spatial data features;
[0073] Step 132: Generate UMAP and spatial distribution maps based on the spatial data features to visually display the characteristics of spatial transcriptomics and achieve data feature visualization.
[0074] In the embodiments of the present invention, by statistically analyzing the proportion of spatial cells and the expression of spatially differential genes, highly variable genes and characteristic genes are automatically selected, comprehensively extracting the characteristics of spatial transcriptome data. The characteristics not only reflect the spatial distribution of cells but also reveal the spatial patterns of gene expression and their relationships with cell types, developmental stages, or disease states. By generating UMAP and spatial distribution maps, the spatial data characteristics are visually displayed in a graphical manner. The visualization results not only help researchers quickly capture key information in the data but also enhance their understanding and interpretation abilities of the characteristics of spatial transcriptomics. Through intuitive visualization, researchers can more easily discover patterns and trends in the data. By automatically extracting spatial data characteristics and generating visualization results, the method of the present invention significantly improves the efficiency and accuracy of data analysis. Compared with manually extracting features and creating visualization charts, the automated method reduces human errors and omissions while accelerating the analysis speed. The extraction and visualization of spatial data characteristics are important steps in data-driven biological research. By deeply analyzing these characteristics and visualization results, researchers can discover new biological phenomena, reveal the interaction mechanisms between cells, and explore the occurrence and development processes of diseases.
[0075] In the embodiments of the present invention, the specific steps include:
[0076] Step 131: Add functions such as timed polling to automate the conversion from the spatial gene expression matrix to the standard analysis of spatial transcriptomics, enabling the automatically execution of the spatial basic analysis of the stereopy framework for the expression matrix generated by SAW. Read the spatial transcriptome expression matrix data after gene dictionary conversion and cell unit selection. The data contains gene names converted to gene symbols, selected cell analysis units (cell bin or square bin), and corresponding expression level information. According to the selected analysis unit, count the number of cells or the proportion of expression levels in each unit to obtain the spatial cell proportion information. Use statistical methods to calculate the gene expression differences between different spatial positions or analysis units. By comparing the expression levels at different positions or units, identify genes with significant differences, i.e., spatially differential genes. Automatically select highly variable genes according to certain criteria, such as the variability of expression levels and their relevance to biological processes. Integrate information such as spatial cell proportion, spatially differential gene expression, highly variable genes, and characteristic genes to form spatial data characteristics.
[0077] Step 132, select UMAP as the visualization method. UMAP is a technique for dimensionality reduction and visualization. Apply the gene expression data in the spatial data features to the UMAP algorithm for dimensionality reduction processing. Generate a spatial distribution map based on the dimensionality-reduced data and the selected analysis unit. The spatial distribution map graphically shows the cell distribution and gene expression at different spatial positions or analysis units. Add annotations and labels to the spatial distribution map. The annotations and labels can include information such as cell type, gene name, and expression level, and output the generated spatial distribution map.
[0078] In a preferred embodiment of the present invention, the above step 14 may include:
[0079] Step 141, perform automated standardization on the data features to obtain standardized data, and then perform dimensionality reduction;
[0080] Step 142, according to the dimensionality-reduced data, automatically select neighbors and spatial_neighbors to construct a neighborhood graph;
[0081] Step 143, according to the constructed neighborhood graph, automatically select a clustering algorithm to cluster the data to obtain clustered data.
[0082] In the embodiments of the present invention, through various standardization techniques, such as normalize_total, log1p, scale, sctransform, quantile, and gaussian_smooth, etc., the data dimension is effectively reduced, while the main structure and features of the original data are retained, reducing the complexity and computational amount of data processing. By automatically selecting neighbors and spatial_neighbors, a neighborhood graph reflecting the similarity and spatial relationship between data points is constructed, which not only simplifies the construction process of the neighborhood graph, but also improves the construction efficiency and accuracy. According to the constructed neighborhood graph, a suitable clustering algorithm (such as leiden, louvain, and phenograph, etc.) is automatically selected to perform intelligent clustering on the data, reducing manual intervention by users and improving the accuracy and efficiency of clustering. By integrating multiple dimensionality reduction methods, neighborhood graph construction methods, and clustering algorithms, a more flexible and adaptable data analysis framework is provided. Data dimensionality reduction, neighborhood graph construction, and clustering analysis are important steps in data-driven biological research. By deeply analyzing the dimensionality-reduced data, neighborhood graph, and clustering results, new biological phenomena can be discovered, the interaction mechanism between cells can be revealed, and the occurrence and development process of diseases can be explored.
[0083] In the embodiments of the present invention, the specific steps include:
[0084] Step 141, read the spatial transcriptome data features after preprocessing (such as gene dictionary conversion, cell unit selection, spatial data feature extraction, etc.). These features include gene expression levels and spatial location information. According to the user-defined analysis requirements or built-in strategies, select a suitable normalization method, including normalize_total (total expression normalization), log1p (logarithmic transformation of expression levels and adding 1 to avoid 0 values), scale (normalization processing to make the data have zero mean and unit variance), SCTransform (a normalization method specifically for single-cell data), quantile (quantile normalization), and gaussian_smooth (an entropy-based adaptive normalization method). Different normalization methods will perform different transformations on the data to remove noise, reduce dimensions, and retain key information. For example, the normalize_total method will divide the expression level of each cell by the total expression level of that cell, making the sum of the expression levels of each cell equal to 1; the log1p method will perform a logarithmic transformation on the expression levels to compress the range of high expression levels and enhance the visibility of low expression levels. Subsequently, perform dimensionality reduction on the data to obtain the dimensionality-reduced data. The data after dimensionality reduction will have a lower dimension while retaining the main structural information of the original data.
[0085] Step 142, read the dimensionality-reduced data, which is now in a low-dimensional space for easy calculation and analysis. According to the user-defined analysis requirements or built-in strategies, select a suitable neighborhood graph construction method, including neighbors (neighborhood graph construction based on Euclidean distance) and spatial_neighbors (neighborhood graph construction considering spatial location). Automatically select the parameters required for constructing the neighborhood graph, such as neighborhood size and distance metric method; the selection of these parameters is based on the characteristics of the data, the requirements of the analysis, or the built-in default settings. According to the selected method and parameters, construct a neighborhood graph. A neighborhood graph is a graph structure where nodes represent data points (such as cells or spatial locations) and edges represent the adjacency relationships between nodes. During the construction process, the distance between each node and its neighbors will be calculated, and the graph structure will be constructed based on the distance relationship.
[0086] Step 143: Read the constructed neighborhood graph. According to the user's predefined analysis requirements or built-in strategies, select a suitable clustering algorithm, such as leiden, louvain, and phenograph. Automatically execute the selected clustering algorithm to cluster the data points in the neighborhood graph. During the clustering process, the algorithm will divide the data points into different clusters (i.e., cell types or spatial regions) according to the structure of the neighborhood graph and the attributes of the nodes. After clustering analysis, clustering data is obtained, including the cluster label to which each data point belongs, the number and size of the clusters, so as to obtain the clustering results, including clustering labels, the number and size of the clusters, and evaluation metrics for clustering quality.
[0087] In a preferred embodiment of the present invention, the above step 15 may include:
[0088] Based on the clustering data, perform automatic annotation of cell types, and then perform differential expression analysis between cell types and GO / KEGG / Reactom enrichment analysis of differential genes to achieve subcellular-level spatial transcriptome function analysis.
[0089] In the embodiments of the present invention, by automatically annotating humans and mice in the standard analysis results according to the clustering data and embedding gene dictionary transformation, rapid and accurate annotation of cell types is achieved, which not only reduces the time and labor costs of manual annotation, but also improves the accuracy and consistency of annotation. Through automatic annotation of cell types, differential expression analysis between cell types, and GO / KEGG / Reactom enrichment analysis of differential genes in different spatial regions, it helps researchers better understand the mechanism of action and biological significance of cells in organisms. By integrating functions such as automatic annotation of cell types and enrichment analysis, the intelligent level of data analysis is improved. This intelligent data analysis method not only improves the efficiency and accuracy of analysis, but also promotes data-driven biological discoveries.
[0090] In the embodiments of the present invention, the specific steps include:
[0091] Read the data after clustering, which contains the cluster label to which each data point (such as a cell or a spatial location) belongs. According to the predefined strategy or built-in knowledge base, formulate an automatic annotation strategy for cell types. For human and mouse data, use the automatic annotation strategy to annotate the cell types of each cluster in the clustering results. During the annotation process, compare the expression pattern of the cluster with the known cell type expression patterns, and select the most matching cell type as the annotation result. During the annotation process, use the gene dictionary to convert the gene identifiers in the annotation result into common gene names to obtain the automatic annotation result, including information such as the cell type annotation of each cluster and the confidence level of the annotation;
[0092] For each cluster (i.e., each annotated cell type), count the number of cells belonging to that cluster and obtain the proportion of each cell type in the total number of cells. Using databases such as GO, KEGG, and Reactome, perform functional annotation on the automatically annotated cell types, provide a refined cell annotation interface, support users to perform refined annotation of cell types through external databases, upload their own data or query external databases through the interface to obtain more detailed cell type information and marker genes, and integrate the automatic annotation results, functional annotation results, and cell proportion information.
[0093] Such as Figure 1-6 , in an embodiment of the system for subcellular-level spatial transcriptome functional analysis of the present invention, the implementation principle includes:
[0094] S1. System subcellular and expression matrix optimization module, such as Figure 2 shown:
[0095] 1. Sample information automation module: SAW automatically adds organism, tissue, kit-version, and sequencing-type modules, and automatically executes spatial gene expression analysis and spatial basic analysis functions.
[0096] 2. Multi-sample statistics module: Add the statistical results of the SAW expression matrix to facilitate comparison of multi-sample and big data analysis.
[0097] 3. Subcellular high-resolution processing module: Improve the tissue square bin statistics from bin20, 50, 200 to 15 subcellular statistical gradients of more refined bin1, 3, 5, 10, 12, 20, 24, 30, 40, 50, 60, 80, 100, 150, 200.
[0098] 4. Cell diameter module: Add corresponding cell size dimensions of 0.7, 2, 3.5, 7, 8, 14, 16, 24, 28, 35, 42, 56, 70, 105, 140 microns according to the subcellular statistical gradient for easy understanding and visualization.
[0099] S2. System gene dictionary and cell unit selection module, such as Figure 3 shown:
[0100] 1. Gene dictionary module: Add functions such as timed polling to automate the conversion from spatial gene expression matrix to standard analysis of spatial transcriptome, automatically convert Ensembl identifiers to common gene names of Gene system Symbol, and enhance the biological readability of downstream analysis.
[0101] 2. Cell Unit Selection Module: The cell unit automatically selects the cell bin mode or the square bin mode according to the chip type; if it is the square bin mode, it automatically selects the unit size of the cell bin.
[0102] S3. System Spatial Data Feature Visualization Module, such as Figure 4 shown:
[0103] 1. Spatial Cell Proportion Statistics Module: Visualize the proportion of spatial cluster cells and annotated cells.
[0104] 2. Differential Gene Expression Statistics Module: Statistic and visualize spatial differential genes.
[0105] 3. Gene Spatial Visualization Module: Automatically select the top 10 highly variable genes, feature gene umap, and spatial visualization.
[0106] S4. System Standardized Clustering Automatic Selection Module, such as Figure 5 shown:
[0107] 1. Standardization Module: Automatically select normalize_total, log1p, scale, SCTransform, quantile, EAGS.
[0108] 2. Neighborhood Graph Module: Automatically select neighbors, spatial_neighbors.
[0109] 3. Clustering Module: Automatically select Leiden, Louvain, and Phenograph.
[0110] S5. System Cell Annotation, such as Figure 6 shown:
[0111] 1. Cell Type Automatic Annotation Module: Automatic annotation of human and mouse for standard analysis results, embedded gene dictionary conversion module.
[0112] 2. Number and Proportion of Cell Types: Number of cells of cell types, average of median of genes.
[0113] 3. Differential Expression Analysis: Differentially expressed genes of cell types.
[0114] 4. Enrichment Analysis Module: GO / KEGG / Reactom enrichment analysis of differentially expressed genes.
[0115] The object of the present invention is:
[0116] S1. System High-Resolution Spatial Transcriptome Processing Module: By expanding the statistical results of multiple samples and refining the subcellular statistical gradient, this module provides higher-resolution analysis capabilities and supports an automated sample information filling module. Combined with the cell diameter module, it facilitates intuitive understanding and visualization by researchers.
[0117] S2. System Analysis Coherence: Through process control, the coherence from the expression matrix to the spatial basic analysis is achieved. The automatic conversion of gene names is realized through the gene dictionary module, enhancing the readability of the data. Meanwhile, the cell unit selection module automatically selects appropriate analysis units (such as square bin or cell bin) according to different requirements. The present invention can optimize the data processing flow, add the gene dictionary module and the cell unit selection module, enhance the compatibility and connection with tools such as Saw, and improve the coherence of the analysis.
[0118] S3. System Data Output and Statistical Issues: For spatial characteristics, by adding visualization functions of spatial cluster and annotated cell proportion, as well as spatial statistics functions of differential gene expression, the spatial distribution characteristics of the data are intuitively displayed. The present invention can provide richer data output options, including spatial cell proportion statistics module, differential gene expression statistics module, gene spatial visualization module, and text table output, facilitating further analysis and application of the data.
[0119] S4. System Automation and Selection Function: By using standardized methods (such as SCTransform, EAGS) and automated neighborhood graph and clustering algorithm selection, the intelligent level of dimensionality reduction and clustering is improved, avoiding biases caused by improper algorithm selection. The present invention can provide the functions of an automated dimensionality reduction module, neighborhood graph module, and clustering module, reducing the operation complexity of users and improving the automation degree of the analysis.
[0120] S5. System Cell Annotation Module: Automatic cell type annotation, functional annotation (such as GO / KEGG / Reactom), gene function annotation. The present invention can integrate an automatic cell type annotation module, enrichment analysis module, and automated data deep mining module, filling the gaps in Scanpy, Squidpy, and Seruat, and solving the deficiencies of Stereopy software in this regard. Using machine learning algorithms and rich reference databases to improve the accuracy of cell annotation.
[0121] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for subcellular spatial transcriptome functional analysis, characterized in that: The method comprises: Automatically process the core information of samples and optimize the spatial transcriptome expression matrix data to realize the reception and processing of expression matrix data; Add timed polling to automate the analysis of spatial gene expression matrix to spatial transcriptome standard. According to the expression matrix data, perform gene ID conversion, cell unit type and cell unit selection to obtain the selected data. According to the transformed and selected data, the spatial data features are counted and the data features are visualized; Standardize the data features and automatically select a clustering algorithm to cluster the reduced-dimensional data to obtain clustered data; Based on the clustered data, subcellular spatial transcriptome functional analysis is achieved, including automated annotation of cell types, gene function enrichment analysis, and cell ratio analysis.
2. The method for subcellular spatial transcriptome functional analysis according to claim 1, characterized in that: Automatically process the core information of samples and optimize spatial transcriptome expression matrix data to realize the reception and processing of expression matrix data, including: According to the sample results, the core information of the sample is automatically filled in, including species, tissue type, kit version and sequencing type, to achieve sample management and expression matrix data reception and processing; By automatically processing the core information of the sample, receiving the spatial transcriptome expression matrix data, and adding the statistical analysis results of the SAW expression matrix, we can obtain multi-sample statistical results; According to the statistical results of multiple samples, 15 subcellular statistical gradients extending from bin20, 50, 200 to bin1 to bin200 are provided, and the corresponding cell diameter size range is automatically generated.
3. The method for subcellular spatial transcriptome functional analysis according to claim 2, characterized in that: According to the expression matrix data, gene ID conversion, cell unit type and cell unit selection are performed to obtain the selected data, including: Add timed polling to automate the analysis of spatial gene expression matrix to spatial transcriptome standard, receive spatial transcriptome expression matrix data, automatically convert ensembl to gene symbol, and realize gene ID conversion; Get predefined analysis requirements, automatically select cell bins or square bins based on the analysis requirements, and dynamically adjust the bin size to obtain transformed and selected data.
4. The method for subcellular spatial transcriptome functional analysis according to claim 3, characterized in that: According to the transformed and selected data, the spatial data features are counted and visualized, including: According to the converted and selected data, the spatial cell proportion and spatial differential gene expression are counted, and highly variable genes and characteristic genes are automatically selected to obtain spatial data features; According to the characteristics of spatial data, UMAP and spatial distribution maps are generated to intuitively display the spatial transcriptomics characteristics to achieve data feature visualization.
5. The method for subcellular spatial transcriptome functional analysis according to claim 4, characterized in that: Automatically standardize the data features, and automatically select a clustering algorithm for clustering after dimensionality reduction to obtain clustered data, including: Automatically standardize the data features to obtain standardized data, and then perform dimensionality reduction; According to the reduced-dimensional data, neighbors and spatial_neighbors are automatically selected to construct a neighborhood graph; According to the constructed neighborhood graph, a clustering algorithm is automatically selected to cluster the data to obtain clustered data.
6. The method for subcellular spatial transcriptome functional analysis according to claim 5, characterized in that: Based on the clustered data, subcellular spatial transcriptome functional analysis is performed, including automatic annotation of cell types, gene function enrichment analysis, and cell ratio analysis, including: According to the clustered data, the cell type automatic annotation is subjected to cellular space GO / KEGG / Reactom enrichment analysis to achieve subcellular spatial transcriptome functional analysis. The basic analysis includes cell type automatic annotation, gene function enrichment analysis and cell ratio analysis.
7. A system for subcellular spatial transcriptome functional analysis, the system implementing the method according to any one of claims 1 to 6, characterized in that: include: Subcellular and expression matrix optimization module, which is used to automatically process the core information of samples and optimize the spatial transcriptome expression matrix data to realize the reception and processing of expression matrix data; Gene dictionary and cell unit selection module, used to convert gene ID, cell unit type and cell unit selection according to expression matrix data to obtain selected data; Spatial data feature visualization module, used to count spatial data features and visualize data features based on the transformed and selected data; The dimensionality reduction clustering automatic selection module is used to automatically standardize data features and automatically select a clustering algorithm to cluster the reduced-dimensional data to obtain clustered data after dimensionality reduction. The cell annotation module is used to perform subcellular spatial transcriptome functional analysis based on clustered data. The analysis includes automated annotation of cell types, gene function enrichment analysis, and cell ratio analysis.
8. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as claimed in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Cited By
Ultrahigh-resolution space transcriptome subcellular expression feature extraction method and application thereof
CN121997037A