A method and system for identifying spatial domains of spatial transcriptomics based on multi-scale neighborhoods

By constructing a multi-scale neighborhood map and adaptive weighting method, multi-scale highly variable genes are identified, spatial embedding features and gene embedding features are extracted, and feature capture in the existing technology is solved, and accurate and efficient recognition of spatial domains is achieved.

CN119905144BActive Publication Date: 2025-08-01SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510386513.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-08-01
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

When identifying the spatial domain, existing spatial transcriptomics methods have problems such as poor spatial continuity, high algorithm complexity and time complexity caused by insufficient feature capture, and cannot accurately map the true structure of the organization.

Method used

Based on spatial transcriptome data of multi-scale neighborhoods, multi-scale highly variable genes are identified by constructing neighborhood maps of different scales, mask maps are generated, and multi-scale spatial embedding features and gene embedding features are extracted, adaptive weighting and iterative clustering are performed, and spatial domain identification is combined with spatial location information.

Benefits of technology

It realizes accurate and efficient identification of the spatial domain, accurately maps the true structure of the organization, improves the accuracy and efficiency of spatial domain identification, and solves the problem of insufficient feature capture in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119905144B_ABST
    Figure CN119905144B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of spatial transcriptome sequencing data processing, and provides a method and system for identifying spatial domains of spatial transcriptomes based on multi-scale neighborhoods. The technical solution is to construct graphs of different-scale neighborhoods based on spatial transcriptome data, and calculate highly variable genes corresponding to the graphs of different-scale neighborhoods; generate a mask graph based on the multi-scale highly variable genes and their corresponding spatial position information, extract multi-scale spatial embedding features based on the mask graph, and adaptively weight the multi-scale spatial embedding features to obtain spatial embedding features; extract multi-scale gene embedding features based on the multi-scale highly variable genes, and adaptively weight the multi-scale gene embedding features to obtain gene embedding features; splice the spatial embedding features and the gene embedding features to obtain embedding features, and perform iterative clustering on the embedding features to obtain the spatial domain identification result. It solves the problem of poor spatial continuity caused by insufficient feature capture, and realizes the accurate identification of spatial domains.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of spatial transcriptome sequencing data processing, and particularly relates to a method and system for identifying spatial domains of spatial transcriptomes based on multi-scale neighborhoods. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] With the development of biotechnology, spatial transcriptomics technology can provide a comprehensive gene expression profile while retaining spatial location information. This technology provides a new perspective for understanding tissue structure, tissue function, and disease mechanisms. Spatial domain identification is one of the most important research contents in spatial transcriptomics. However, due to the characteristics of high discreteness, sparsity, and multimodality of the data, accurately identifying spatial domains from these sequencing data is a challenging task.

[0004] Traditional spatial domain identification methods include non-spatial clustering methods (such as Louvain), which only use gene expression as input, ignoring spatial location information and histological pathological image information, resulting in the lack of spatial continuity in the clustering results and being unable to accurately map the true structure of the tissue. Although some methods have improved the accuracy of spatial domain identification to a certain extent by using spatial location information or further integrating histological pathological images, there are still problems such as poor spatial continuity caused by insufficient feature capture, high algorithmic spatial complexity, and high time complexity. Summary of the Invention

[0005] In order to solve at least one of the technical problems in the above background art, the present invention provides a method and system for identifying spatial domains of spatial transcriptomes based on multi-scale neighborhoods, which identify multi-scale highly variable genes based on spatial transcriptome data, extract spatial embedding features and gene embedding features based on the multi-scale highly variable genes, and make full use of spatial features to improve the accuracy and efficiency of spatial domain identification.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The first aspect of the present invention provides a method for identifying spatial domains of spatial transcriptomes based on multi-scale neighborhoods, including the following steps:

[0008] Obtain spatial transcriptome data;

[0009] Based on the spatial transcriptome data, construct graphs of different-scale neighborhoods and calculate the highly variable genes corresponding to the graphs of different-scale neighborhoods;

[0010] Generate a mask map based on multi-scale highly variable genes and their corresponding spatial position information, extract multi-scale spatial embedding features from the mask map, and adaptively weight the multi-scale spatial embedding features to obtain spatial embedding features;

[0011] Extract multi-scale gene embedding features based on multi-scale highly variable genes, and adaptively weight the multi-scale gene embedding features to obtain gene embedding features;

[0012] Concatenate the spatial embedding features and the gene embedding features to obtain embedding features, and perform iterative clustering on the embedding features to obtain the spatial domain recognition result.

[0013] Further, after obtaining the spatial transcriptome data, preprocess the spatial transcriptome data, including: performing quality control on the gene expression profile data, screening out the sites and genes that meet the conditions, and then performing normalization, logarithmic transformation, and standardization to obtain the preprocessed gene expression profile data.

[0014] Further, the construction process of the graphs of different scales of neighborhoods includes:

[0015] Based on the spatial position information in the spatial transcriptome data, calculate the Euclidean distance between any two sites for subsequent nearest neighbor judgment;

[0016] Let the simple graph where is a set of When and only when is 's nearest neighbor or is 's nearest neighbor, and there is an edge between is composed of the set of Take take and and to obtain graphs of three different scales of neighborhoods and The graphs and The size relationship of the resolutions is: , m is the number of genes, is the site.

[0017] Further, calculate the highly variable genes corresponding to the graphs of different scales of neighborhoods, including:

[0018] Constructing adjacency matrices based on graphs of neighborhoods at different scales and diagonal matrices ;

[0019] Based on the adjacency matrix and diagonal matrices , calculate the Fourier coefficients and Fourier modes ;

[0020] Computational genetics , Defined as: , ,in, is the Laplace matrix No. eigenvalues, is the first A quantity, is the first in the unnormalized initial spectrum domain A quantity, is the gene expression level, n is the total number of sites, m is the number of genes;

[0021] Determine whether each gene is a highly variable gene: When a gene meets two conditions, it is considered a highly variable gene. First, its Greater than all genes The inflection point of the distribution, second, its preceding Low frequency of than before High frequency of There is statistical difference.

[0022] Furthermore, based on the adjacency matrix and diagonal matrices , calculate the Fourier coefficients and Fourier modes When combined with the adjacency matrix and diagonal matrices , calculate the Laplacian matrix , n is the number of sites, using spectral decomposition , , get the eigenvalue , eigenvector , corresponding to the Fourier coefficients and Fourier modes .

[0023] Further, the mask graph generated based on the multi-scale highly variable genes and their corresponding spatial position information can be replaced with a pruning graph, and the spatial embedding features corresponding to the scale are obtained based on the pruning graph and the graph variational autoencoder.

[0024] Further, the spatial embedding features are represented as:

[0025] ,

[0026] wherein, represents the spatial embedding features, represents the spatial embedding features of the first scale, represents the importance of the spatial embedding features of the first scale for spatial domain recognition, represents the spatial embedding features of the second scale, represents the importance of the spatial embedding features of the second scale for spatial domain recognition, represents the spatial embedding features of the third scale, represents the importance of the spatial embedding features of the third scale for spatial domain recognition;

[0027] The gene embedding features are represented as:

[0028] ,

[0029] wherein, represents the gene embedding features, represents the gene embedding features of the first scale, represents the importance of the gene embedding features of the first scale for spatial domain recognition, represents the gene embedding features of the second scale, represents the importance of the gene embedding features of the second scale for spatial domain recognition, represents the gene embedding features of the third scale, represents the importance of the gene embedding features of the third scale for spatial domain recognition.

[0030] The second aspect of the present invention provides a spatial domain recognition system for spatial transcriptomics based on multi-scale neighborhoods, including:

[0031] A data acquisition module for acquiring spatial transcriptomics data;

[0032] A multi-scale highly variable gene recognition module for constructing graphs of different scale neighborhoods based on the spatial transcriptomics data and calculating the highly variable genes corresponding to the graphs of different scale neighborhoods;

[0033] A spatial embedding feature extraction module, which is used to generate a mask map based on multi-scale highly variable genes and their corresponding spatial position information, extract multi-scale spatial embedding features based on the mask map, and adaptively weight the multi-scale spatial embedding features to obtain spatial embedding features;

[0034] A gene embedding feature extraction module, which is used to extract multi-scale gene embedding features based on multi-scale highly variable genes, and adaptively weight the multi-scale gene embedding features to obtain gene embedding features;

[0035] A spatial domain recognition module, which is used to splice the spatial embedding features and the gene embedding features to obtain embedding features, and perform iterative clustering on the embedding features to obtain a spatial domain recognition result.

[0036] Furthermore, the system further includes a data preprocessing module, which is used to preprocess the spatial transcriptome data after obtaining the spatial transcriptome data, including: performing quality control on the gene expression profile data, screening out the sites and genes that meet the conditions, and then performing normalization, logarithmic transformation, and standardization to obtain the preprocessed gene expression profile data.

[0037] Furthermore, in the spatial embedding feature extraction module, the mask map generated based on the multi-scale highly variable genes and their corresponding spatial position information can be replaced with a pruning map, and the corresponding scale of spatial embedding features is obtained based on the pruning map and the graph variational autoencoder.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] The present invention extracts spatial embedding features and gene embedding features based on the identified multi-scale highly variable genes, considers the spatial continuity, accurately maps the true structure of the tissue, solves the poor spatial continuity caused by insufficient feature capture in the prior art, and realizes the accurate and efficient recognition of the spatial domain.

[0040] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0041] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0042] Figure 1 It is a flowchart of a method for identifying spatial domains of spatial transcriptomes based on multi-scale neighborhoods provided by an embodiment of the present invention;

[0043] Figure 2It is a block diagram of a method for identifying spatial domains of spatial transcriptomics based on multi-scale neighborhoods provided by an embodiment of the present invention;

[0044] Figure 3 It is a comparison diagram of effects with a pathological image as the negative film provided by an embodiment of the present invention. Among them, (a) is the recognition effect diagram of manual annotation, (b) is the recognition effect diagram of the stLearn method, (c) is the recognition effect diagram of the Seurat method, (d) is the recognition effect diagram of the SpaceFlow method, (e) is the recognition effect diagram of the stAA method, and (f) is the recognition effect diagram of the method of the present invention. Detailed implementation manners

[0045] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0046] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0047] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0048] Regarding the existing methods mentioned in the background technology, by using spatial position information or further integrating histological pathological images, the accuracy of spatial domain recognition has been improved to a certain extent, but there are still problems such as poor spatial continuity caused by insufficient feature capture, high algorithmic spatial complexity and time complexity. The present invention performs graph Fourier transform of multi-scale neighborhoods based on spatial transcriptome data and identifies multi-scale highly variable genes; generates a mask graph based on the multi-scale highly variable genes and their corresponding spatial position information, extracts multi-scale spatial embedding features based on the mask graph, and adaptively weights the multi-scale spatial embedding features to obtain spatial embedding features; based on the multi-scale highly variable genes, extracts multi-scale gene embedding features, and adaptively weights the multi-scale gene embedding features to obtain gene embedding features; splices the spatial embedding features and the gene embedding features to obtain embedding features, and performs iterative clustering on the embedding features to obtain the spatial domain recognition result. It solves the problem of poor spatial continuity caused by insufficient feature capture and realizes the accurate and efficient recognition of spatial domains (regions that are spatially consistent with respect to gene expression profiles and histological pathological images).

[0049] Embodiment 1

[0050] As Figure 1 and Figure 2 shown, this embodiment provides a method for identifying the spatial domain of spatial transcriptomics based on multi-scale neighborhoods, including the following steps:

[0051] Step 1: Obtain spatial transcriptome data and perform preprocessing;

[0052] The spatial transcriptome data includes gene expression profile data , spatial position information and pathological images. For the gene expression profile data, quality control is performed to screen out sites and genes that meet the conditions. Secondly, normalization, logarithmic transformation, and standardization are performed to obtain the preprocessed gene expression data , where n is the number of sites, m |is the number of genes;

[0053] Step 2: Convert the preprocessed gene expression data , spatial position information from graph signals into frequency domain information and extract multi-scale highly variable genes;

[0054] Specifically, it includes the following steps:

[0055] Step 201: Based on the spatial position information , calculate the Euclidean distance between any two for subsequent nearest neighbor judgment;

[0056] Step 202: Construct graphs of different scale neighborhoods , ;

[0057] In this embodiment, let the simple graph , where is a set of . When and only when is 's nearest neighbor or is |'s nearest neighbor ( idea), and there is an edge between them. The set composed of constitutes . Take and take , , , we get the graphs of neighborhoods of three different scales 、 and .

[0058] Among them, Figure 、 and The relationship between resolutions is: ;

[0059] Step 203: Calculate the graphs of neighborhoods at different scales 、 and The corresponding highly variable genes;

[0060] The specific steps include:

[0061] Step 2031: Construct an adjacency matrix based on graphs of neighborhoods of different scales and diagonal matrices ;

[0062] In this embodiment, , diagonal matrix ,in for degree;

[0063] Step 2032: Based on the adjacency matrix and diagonal matrices , calculate the Fourier coefficients , Fourier mode ;

[0064] In this embodiment, combined with the adjacency matrix and diagonal matrices , calculate the Laplacian matrix , using spectral decomposition , , get the eigenvalue , eigenvector , corresponding to 、 ;

[0065] Step 2033, calculate the gene ;

[0066] The gene expression Converting to the spectral domain, we get , Defined as ,in ,in, is the Laplace matrix No. eigenvalues, is the first component is the th component in the initial spectral domain that has not been standardized.

[0067] Step 2034: Determine whether each gene is a highly variable gene: When a gene meets two conditions, it is regarded as a highly variable gene, including: (1) its is greater than the inflection point of the distribution of all genes , and (2) the low-frequency of its first is statistically different from the high-frequency of its first ; finally, the corresponding highly variable genes of , and are obtained;

[0068] Step 3: Generate a mask map based on the multi-scale highly variable genes and their corresponding spatial position information, encode and decode the mask map to capture the multi-scale spatial embedding features, and adaptively weight the multi-scale spatial embedding features to obtain the spatial embedding features;

[0069] In this embodiment, graphs are respectively constructed based on the low-resolution highly variable genes, medium-resolution highly variable genes, high-resolution highly variable genes and their corresponding spatial position information and are respectively input into 3 graph variational autoencoders to obtain the low-scale spatial embedding features , medium-scale spatial embedding features and high-scale spatial embedding features .

[0070] Then, use the attention mechanism to adaptively learn the importance of the low-scale spatial embedding features , medium-scale spatial embedding features and high-scale spatial embedding features for spatial domain recognition;

[0071] The importance of the low-scale spatial embedding features for spatial domain recognition is , where is the activation function, is the weight matrix, is the bias vector; similarly, calculate the importance of the medium-scale spatial embedding features and the importance of the high-scale spatial embedding features for spatial domain recognition. Finally, the weighted spatial embedding feature is 。

[0072] In this embodiment, in the adaptive graph variational autoencoder module, a pruning module can be selectively adopted;

[0073] The low-resolution highly variable genes, medium-resolution highly variable genes, and high-resolution highly variable genes are respectively dimensionality-reduced to obtain 、 、 ;

[0074] For 、 、 perform pre-clustering, and pruning is performed according to the pre-clustering results.

[0075] Regarding , if in the pre-clustering results, if one 's neighbors belong to different clusters, then prune this to obtain a pruned graph ;

[0076] Use the pruned graph instead of the ordinary graph directly constructed .

[0077] Step 4: Encode and decode the multi-scale highly variable genes, capture the multi-scale gene embedding features, and adaptively weight the multi-scale gene embedding features to obtain gene embedding features;

[0078] In this embodiment, the low-resolution highly variable genes, medium-resolution highly variable genes, and high-resolution highly variable genes are respectively input into the corresponding masked autoencoders to obtain low-scale gene embedding features 、medium-scale gene embedding features and high-scale gene embedding features .

[0079] Then, use the attention mechanism to adaptively learn 、 and 's importance for spatial domain recognition. The importance of the low-scale spatial embedding feature for spatial domain recognition is , and similarly calculate 、 , and the spatial embedding feature is 。

[0080] Based on the identified multi-scale highly variable genes, this invention extracts spatial embedding features and gene embedding features, considers the continuity of space, accurately maps the true structure of tissues, solves the poor spatial continuity caused by insufficient feature capture in the prior art, and realizes accurate and efficient identification in the spatial domain.

[0081] Step 5: Concatenate the spatial embedding features and gene embedding features to obtain embedding features, and perform iterative clustering on the embedding features to obtain clustering results;

[0082] In this embodiment, and are concatenated to obtain embedding features . Perform iterative clustering on to obtain clustering results .

[0083] As Figure 3 shown is the effect comparison diagram of this invention and 4 existing methods. Among them, (a) is the recognition effect diagram of manual annotation, (b) is the recognition effect diagram of the stLearn method, (c) is the recognition effect diagram of the Seurat method, (d) is the recognition effect diagram of the SpaceFlow method, (e) is the recognition effect diagram of the stAA method, and (f) is the recognition effect diagram of the method of this invention. The experimental results show that the method of this invention performs excellently in spatial domain recognition, can effectively improve the accuracy of spatial domain recognition, has stronger discrimination ability and data adaptability than other methods, and has higher precision and robustness in dealing with the area boundaries that are difficult to clearly identify by other methods.

[0084] Embodiment 2

[0085] This embodiment provides a spatial domain recognition system for spatial transcriptomics based on multi-scale neighborhoods, including:

[0086] A data acquisition module for acquiring spatial transcriptomic data;

[0087] A multi-scale highly variable gene recognition module for constructing graphs of different-scale neighborhoods based on the spatial transcriptomic data and calculating the highly variable genes corresponding to the graphs of different-scale neighborhoods;

[0088] A spatial embedding feature extraction module for generating a mask graph based on the multi-scale highly variable genes and their corresponding spatial position information, extracting multi-scale spatial embedding features based on the mask graph, and adaptively weighting the multi-scale spatial embedding features to obtain spatial embedding features;

[0089] A gene embedding feature extraction module for extracting multi-scale gene embedding features based on the multi-scale highly variable genes, and adaptively weighting the multi-scale gene embedding features to obtain gene embedding features;

[0090] A spatial domain recognition module, which is used to splice the spatial embedding feature and the gene embedding feature to obtain an embedding feature, and perform iterative clustering on the embedding feature to obtain a spatial domain recognition result.

[0091] Furthermore, the system further includes a data preprocessing module, which is used to preprocess the spatial transcriptome data after obtaining the spatial transcriptome data, including: performing quality control on the gene expression profile data, screening out sites and genes that meet the conditions, and then performing normalization, logarithmic transformation, and standardization to obtain the preprocessed gene expression profile data.

[0092] In the spatial embedding feature extraction module, the mask graph generated based on the multi-scale highly variable genes and their corresponding spatial position information can be replaced with a pruning graph, and the spatial embedding feature corresponding to the scale is obtained based on the pruning graph and the graph variational autoencoder.

[0093] Since the embodiments in the system part correspond to the embodiments in the method part, please refer to the description of the embodiments in the method part for the embodiments in the system part, which will not be elaborated here. And it has the same beneficial effects as the above-mentioned method for spatial domain recognition of spatial transcriptome based on multi-scale neighborhoods. The spatial embedding feature and the gene embedding feature are extracted based on the identified multi-scale highly variable genes, considering the continuity of space, accurately mapping the true structure of the tissue, solving the poor spatial continuity caused by insufficient feature capture in the prior art, and realizing accurate and efficient recognition of the spatial domain.

[0094] Embodiment III

[0095] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the above-mentioned method for spatial domain recognition of spatial transcriptome based on multi-scale neighborhoods.

[0096] Embodiment IV

[0097] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the above-mentioned method for spatial domain recognition of spatial transcriptome based on multi-scale neighborhoods.

[0098] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for identifying spatial domains of spatial transcriptomics based on multi-scale neighborhoods, characterized in that, Including the following steps: Obtain spatial transcriptome data; Based on the spatial transcriptome data, construct graphs of different-scale neighborhoods, and calculate the highly variable genes corresponding to the graphs of different-scale neighborhoods; Generate a mask graph based on the multi-scale highly variable genes and their corresponding spatial position information, extract multi-scale spatial embedding features based on the mask graph, and adaptively weight the multi-scale spatial embedding features to obtain spatial embedding features; Extract multi-scale gene embedding features based on the multi-scale highly variable genes, and adaptively weight the multi-scale gene embedding features to obtain gene embedding features; wherein, the gene embedding features are expressed as: , Among them, represents the gene embedding feature, represents the first-scale gene embedding feature, represents the first-scale gene embedding feature importance for spatial domain recognition, represents the second-scale gene embedding feature, represents the second-scale gene embedding feature importance for spatial domain recognition, represents the third-scale gene embedding feature, represents the third-scale gene embedding feature importance for spatial domain recognition; Concatenate the spatial embedding features and the gene embedding features to obtain embedding features, and perform iterative clustering on the embedding features to obtain the spatial domain recognition result.

2. The spatial domain recognition method of spatial transcriptome based on multi-scale neighborhood according to claim 1, wherein, After obtaining the spatial transcriptome data, preprocess the spatial transcriptome data, including: performing quality control on the gene expression profile data, screening out the sites and genes that meet the conditions, and then performing normalization, logarithmic transformation, and standardization to obtain the preprocessed gene expression profile data.

3. The spatial domain recognition method for spatial transcriptomics based on multi-scale neighborhood according to claim 1, characterized in that, The construction process of the graphs of different-scale neighborhoods includes: Based on the spatial location information in the spatial transcriptome data, calculate the Euclidean distance between any two loci for subsequent nearest neighbor judgment; Let the simple graph , where is a set of . An edge exists betweenand if and only if is the nearest neighbor of or is the nearest neighbor of . The graph is composed of the set of . Taking as , , , three graphs and with different scales of neighborhoods are obtained. The size relationship of the resolutions of the graphs , is: , m where is the number of genes and is the locus.

4. The spatial domain recognition method for spatial transcriptomics based on multi-scale neighborhoods according to claim 1, wherein, Calculating the highly variable genes corresponding to the graphs of different-scale neighborhoods includes: Construct the adjacency matrix of the graph based on neighborhoods of different scales and the diagonal matrix ; Based on the adjacency matrix and the diagonal matrix , the Fourier coefficients and the Fourier modes ; Calculating genes , is defined as: , where, is the Laplacian matrix of the th eigenvalue, is the th component in the normalized spectral domain, is the th component in the unnormalized initial spectral domain, is the gene expression level, n is the total number of loci, m is the number of genes; Determine whether each gene is a highly variable gene: When a gene meets two conditions, it is regarded as a highly variable gene. First, its is greater than the inflection point of the distribution of all genes Second, the low-frequency before it is statistically different from the high-frequency before it. There is a statistical difference.

5. The spatial domain recognition method of spatial transcriptome based on multi-scale neighborhood according to claim 4, wherein Based on the adjacency matrix and the diagonal matrix , when calculating the Fourier coefficients and the Fourier modes , combine the adjacency matrix and the diagonal matrix to calculate the Laplacian matrix . n is the number of sites. Using spectral decomposition . , the eigenvalues and eigenvectors are obtained, corresponding to the Fourier coefficients and the Fourier modes respectively.

6. The spatial domain recognition method of spatial transcriptome based on multi-scale neighborhood according to claim 1, characterized in that The mask graph generated based on the multi-scale highly variable genes and their corresponding spatial position information can be replaced by a pruning graph, and the corresponding-scale spatial embedding features are obtained based on the pruning graph and the graph variational autoencoder.

7. A method for identifying the spatial domain of spatial transcriptomics based on multi-scale neighborhoods according to claim 1, characterized in that, The spatial embedding features are expressed as: , Among them, represents the spatial embedding feature, represents the first-scale spatial embedding feature, represents the first-scale spatial embedding feature importance for spatial domain recognition, represents the second-scale spatial embedding feature, represents the second-scale spatial embedding feature importance for spatial domain recognition, represents the third-scale spatial embedding feature, represents the third-scale spatial embedding feature importance for spatial domain recognition.

8. A spatial domain recognition system for spatial transcriptomics based on multi-scale neighborhoods, characterized in that, Including: A data acquisition module for obtaining spatial transcriptome data; A multi-scale highly variable gene recognition module for constructing graphs of different-scale neighborhoods based on the spatial transcriptome data and calculating the highly variable genes corresponding to the graphs of different-scale neighborhoods; A spatial embedding feature extraction module for generating a mask graph based on the multi-scale highly variable genes and their corresponding spatial position information, extracting multi-scale spatial embedding features based on the mask graph, and adaptively weighting the multi-scale spatial embedding features to obtain spatial embedding features; A gene embedding feature extraction module for extracting multi-scale gene embedding features based on the multi-scale highly variable genes, and adaptively weighting the multi-scale gene embedding features to obtain gene embedding features; wherein, the gene embedding features are expressed as: , Among them, represents the gene embedding feature, represents the first-scale gene embedding feature, represents the first-scale gene embedding feature importance for spatial domain recognition, represents the second-scale gene embedding feature, represents the second-scale gene embedding feature importance for spatial domain recognition, represents the third-scale gene embedding feature, represents the third-scale gene embedding feature importance for spatial domain recognition; A spatial domain recognition module for concatenating the spatial embedding features and the gene embedding features to obtain embedding features, and performing iterative clustering on the embedding features to obtain the spatial domain recognition result.

9. A spatial transcriptome spatial domain recognition system based on a multi-scale neighborhood as claimed in claim 8, wherein The system further includes a data preprocessing module for preprocessing the spatial transcriptome data after obtaining the spatial transcriptome data, including: performing quality control on the gene expression profile data, screening out the sites and genes that meet the conditions, and then performing normalization, logarithmic transformation, and standardization to obtain the preprocessed gene expression profile data.

10. A spatial domain recognition system for spatial transcriptomics based on multi-scale neighborhoods as claimed in claim 8, wherein In the spatial embedding feature extraction module, the mask graph generated based on the multi-scale highly variable genes and their corresponding spatial position information can be replaced by a pruning graph, and the corresponding-scale spatial embedding features are obtained based on the pruning graph and the graph variational autoencoder.

Citation Information

Patent Citations

  • Space transcriptome data processing method and system based on hypergraph

    CN117457081A

  • BERT-based cell type deconvolution method and system

    CN119724372A