Spatial domain identification method and system based on spatial multi-omics data
Through the anchor concept decomposition and dual-graph regularization method, the spatial multiomics data is integrated, which solves the problems of insufficient utilization of space information and noise interference in the existing technology, and achieves more accurate cellular heterogeneity analysis, which is suitable for tumor microenvironment and cellular heterogeneity research.
Patent Information
- Application Number
- CN202510256192.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-08
AI Technical Summary
When integrating spatial multiomics data in the prior art, spatial information is insufficient, susceptible to noise interference and limited generalization ability, making it difficult to accurately characterize the nonlinear relationship between cells and adapt to diversified application scenarios.
By acquiring multi-omics data and performing data preprocessing, a consistent low-dimensional representation is performed using the anchor concept decomposition framework, a dual-graph regularization objective function is constructed by combining spatial position maps and element-weighted ensemble maps, an alternating optimization algorithm is used to solve the final low-dimensional feature representation, and a k-means algorithm is used for cluster recognition.
It significantly improves the local spatial dependence and robustness of the model, can more accurately analyze the cellular heterogeneity and functional regulation mechanisms in the tissue microenvironment, reduces redundant information interference, and is suitable for tumor microenvironment analysis and cell heterogeneity research.
Smart Images

Figure CN120277435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and more specifically, to a method and system for spatial domain recognition based on spatial omics data. Background Art
[0002] In recent years, breakthroughs in single-cell sequencing technology have made it possible to simultaneously analyze multi-omics data such as mRNA, chromatin accessibility, and DNA methylation at single-cell resolution, providing a more comprehensive molecular map for analyzing cell heterogeneity. However, such technologies lack the spatial location information of cells in tissues, while spatial multi-omics technologies (such as DBiT-seq, CUT&Tag-RNA-seq, etc.) provide an important means for studying the cell microenvironment and tissue structure by integrating multi-modal data such as transcriptomics, proteomics, and metabolomics and retaining spatial information.
[0003] Currently, multi-omics data integration methods are mainly divided into two categories: those based on traditional statistical models (such as MOFA+ and Seurat v5) and those based on deep learning (such as TotalVI and MultiVI). However, their designs are mainly for single-cell data and do not fully combine spatial location information. Although some spatial multi-omics integration methods (such as SpatialGlue and PRAGA) have emerged recently, the existing technologies still have the following key problems:
[0004] Insufficient utilization of spatial information: Most methods focus on the fusion of multi-omics expression profiles and do not effectively explore the guiding role of spatial location in cell-cell interactions and tissue structure, resulting in limited model performance.
[0005] Redundancy and noise interference: Existing methods are vulnerable to noise during cross-omics feature fusion and ignore the complex manifold structure of spatial data, making it difficult to accurately characterize the non-linear relationships between cells.
[0006] Limited generalization ability: Ensemble learning strategies usually do not consider the differential effects of different cell subpopulations on the clustering results, resulting in insufficient algorithm robustness and difficulty in adapting to diverse application scenarios.
[0007] Therefore, there is an urgent need to develop a computational method that can efficiently integrate spatial multi-omics data, reduce noise interference, and improve the generalization ability of the model to more accurately analyze cell heterogeneity and functional regulation mechanisms in the tissue microenvironment. Summary of the Invention
[0008] The present invention aims to overcome the defects of the prior art that cannot efficiently integrate spatial multi-omics data, reduce noise interference, and improve the generalization ability of the model, and provides a method and system for spatial domain recognition based on spatial multi-omics data.
[0009] To solve the above technical problems, the technical solutions of the present invention are as follows:
[0010] The present invention provides a spatial domain recognition method based on spatial omics data, including:
[0011] Obtain omics data, where the omics data includes expression profile information and spatial position coordinates of different omics, and perform data preprocessing on the omics data;
[0012] Extract a consistent low-dimensional representation of the omics data to obtain an association matrix of multiple base clustering results;
[0013] Calculate the distance between each cell based on the spatial position coordinates, and construct a spatial position map considering the higher-order information between cells;
[0014] Integrate the association matrices of multiple base clustering results to obtain an element-weighted integrated map;
[0015] Construct a bi-graph regularization objective function based on the spatial position map and the element-weighted integrated map, and solve the objective function to obtain the final low-dimensional feature representation;
[0016] Perform clustering recognition based on the final low-dimensional feature representation to obtain the spatial domain recognition result.
[0017] Preferably, the data preprocessing includes:
[0018] Use the SCANPY package to remove cells outside the tissue section;
[0019] Retain the top preset number of genes in the RNA-seq data in the expression profile information of different omics according to gene expression differences;
[0020] Perform latent semantic indexing dimensionality reduction on the ATAC-seq data in the expression profile information of different omics;
[0021] Perform normalization and scaling processing on the ADT data in the expression profile information of different omics;
[0022] Fill in the sparse omics data through spatial similarity and gene expression similarity to fill in the missing expression values of the omics data.
[0023] Preferably, use the anchor concept decomposition framework to extract a consistent low-dimensional representation of the omics data. The formula of the anchor concept decomposition framework is:
[0024]
[0025] where α v is the omics weight of adaptive learning; X (v) is the feature matrix of different omics; A (v) ∈R k×m is the anchor base of a specific omics; is the transpose of matrix A (v) ; W (v) ∈ R m×d is the association matrix related to the concept - related anchor bases; P ∈ R d×n is the low - dimensional representation matrix of cells; I m is the identity matrix; ||·|| F is the Frobenius norm.
[0026] Preferably, the association matrices of multiple base clustering results are integrated using an element - wise weighted integration strategy, and the element - wise weighted integration strategy includes:
[0027] Construct a base clustering association matrix, and its formula is:
[0028]
[0029] where E ij d is the association matrix; cls d (x i ) is the category to which the i - th cell belongs in the d - th clustering;
[0030] Determine the cell - level weights through the nearest - neighbor set to generate a consensus association matrix.
[0031] Preferably, the formula of the consensus association matrix is:
[0032]
[0033] where D is the total number of selected base clustering methods; E d is the d - th base clustering association matrix; d is the number of base clustering methods; w d is the high - order weight matrix; is the transpose of matrix w d ; ⊙ represents the dot - product of matrices.
[0034] Preferably, construct a bi - graph regularization objective function based on the spatial position graph and the element - weighted integration graph, including:
[0035] Obtain the initial objective function based on the association matrices of multiple base clustering results and the high - order network containing multi - modal high - order information of spatial omics data;
[0036] In the initial objective function, obtain the bi - graph regularization function by considering the spatial position graph and the element - weighted integration graph.
[0037] Preferably, the bi - graph regularization objective function is:
[0038]
[0039] where, λ is the first balance parameter; β is the second balance parameter; L E is the Laplacian matrix of the consensus correlation matrix; L M is the Laplacian matrix of the high-order adjacency matrix; α v is the weight of omics υ; X (v) is the feature matrix of different omics; A (v) is the anchor base of omics v; is the matrix A (v) transpose; W (v) is the correlation matrix between omics v and the concept-related anchor base; P is the common cell low-dimensional representation matrix between different omics; P T is the transpose of the matrix P; I m is the identity matrix; ||·|| F is the Frobenius norm.
[0040] Preferably, the bi-graph regularization objective function is solved by an alternating optimization algorithm to obtain the final low-dimensional feature representation, including:
[0041] Decompose the bi-graph regularization objective function into three sub-problems, and solve the three sub-problems separately;
[0042] Fix the preset variables of the first sub-problem as constants for solution to obtain the update rule of the first sub-problem; keep the preset variables of the second sub-problem unchanged for derivation to obtain the update rule of the second sub-problem;
[0043] Use derivative zeroing to solve the third sub-problem to obtain the update rule of the third sub-problem;
[0044] Fuse the update rules of the first sub-problem, the second sub-problem, and the third sub-problem to obtain the final low-dimensional feature representation.
[0045] Preferably, the k-means algorithm is used to cluster and identify the final low-dimensional feature representation to obtain the spatial domain identification result.
[0046] The present invention also provides a spatial domain identification system based on spatial multi-omics data for implementing the above method, including:
[0047] A data acquisition module, where the multi-omics data includes the expression profile information and spatial position coordinates of different omics, and preprocesses the multi-omics data;
[0048] A feature extraction module, which extracts the consistent low-dimensional representation of the multi-omics data to obtain the correlation matrix of multiple base clustering results;
[0049] A position map construction module, which calculates the distance between each cell based on the spatial position coordinates and constructs a spatial position map considering the high-order information between cells;
[0050] An integrated graph generation module that integrates the association matrices of multiple base clustering results to obtain an element-weighted integrated graph;
[0051] A target function solving module that constructs a bi-graph regularization target function based on the spatial position graph and the element-weighted integrated graph, and solves the target function to obtain the final low-dimensional feature representation;
[0052] A spatial domain recognition module that performs clustering recognition based on the final low-dimensional feature representation to obtain a spatial domain recognition result.
[0053] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0054] The present invention proposes a spatial domain recognition method and system based on spatial omics data. By using the anchor concept decomposition technology, the omics data is projected into a shared low-dimensional space, breaking through the limitation of traditional methods that only rely on expression profile information. Combining graph regularization constraints, it effectively integrates the high-order relationships and spatial position information between cells, significantly improving the learning ability of local spatial dependence. A per-element weighting strategy is proposed to fully exploit the complementarity of multiple base clustering results, avoiding the error interference caused by traditional integration methods ignoring the performance differences of base clustering, so as to show stable and excellent performance in various application scenarios. Based on the bi-graph regularization framework, while retaining the geometric structure of the original data manifold, it reduces the interference of redundant information, realizes the efficient integration of spatial omics data, and provides a computationally lightweight and highly scalable solution for the analysis of complex biological systems. Compared with the prior art, the present invention has made significant breakthroughs in algorithm accuracy and robustness, and is particularly suitable for scenarios such as tumor microenvironment analysis and cell heterogeneity research, providing a more reliable analysis tool for the fields of precision medicine and bioinformatics. Description of the Drawings
[0055] Figure 1 It is a flowchart of the spatial domain recognition method based on spatial omics data described in Example 1;
[0056] Figure 2 It is an analysis diagram of parameter sensitivity described in Example 2;
[0057] Figure 3 It is a comparison schematic diagram of the spatial domain recognition effects of different methods on the test data set described in Example 2;
[0058] Figure 4 It is a schematic diagram of the index comparison of the test data set on different methods described in Example 2;
[0059] Figure 5 It is a structural schematic diagram of the spatial domain recognition system based on spatial omics data described in Example 3. Detailed implementation manners
[0060] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent;
[0061] To better illustrate this embodiment, some components in the accompanying drawings are omitted, enlarged or reduced, and do not represent the size of the actual product;
[0062] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0063] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0064] Embodiment 1
[0065] This embodiment provides a spatial domain recognition method based on spatial omics data, as Figure 1 shown, including:
[0066] Obtain omics data, where the omics data includes expression profile information and spatial position coordinates of different omics, and perform data preprocessing on the omics data;
[0067] The data preprocessing includes:
[0068] Use the SCANPY package to remove cells outside the tissue section;
[0069] Retain the top preset number of genes in the RNA-seq data in the expression profile information of different omics according to gene expression differences;
[0070] Perform latent semantic indexing dimensionality reduction on the ATAC-seq data in the expression profile information of different omics;
[0071] Perform normalization and scaling processing on the ADT data in the expression profile information of different omics;
[0072] Fill in the sparse omics data through spatial similarity and gene expression similarity to fill in the missing expression values of the omics data;
[0073] Use the SCANPY package to preprocess the data. For all datasets, remove cells outside the tissue section. For RNA-seq data, filter and logarithmically transform the original expression data according to the library size, and only retain the top 3000 genes for downstream analysis according to gene expression differences. Use latent semantic indexing (LSI) to perform dimensionality reduction on ATAC-seq data. ADT data is preprocessed by normalization and scaling. The expression matrices of different omics are represented as The spatial position coordinates are denoted as For sparse omics, the expression profiles in the omics data are imputed by considering the spatial similarity of the closest cells and the expression similarity of the closest genes.
[0074] Extract the consistent low-dimensional representation of multi-omics data to obtain the association matrix of multiple base clustering results;
[0075] Use the anchor concept decomposition framework to extract the consistent low-dimensional representation of multi-omics data. The formula of the anchor concept decomposition framework is:
[0076]
[0077] where, α v can adaptively learn the corresponding weights from multi-omics data; A (v) ∈R k×m is the anchor base of a specific omics. An orthogonality constraint is imposed on it to make the anchor bases of different omics independent of each other and more discriminative; W (v) ∈R m×d represents the association matrix of the anchor bases related to the concept and is the connector between multi-omics data, anchor bases, and low-dimensional representation. It can collect interaction information and correctly indicate semantic clues; P ∈ R d×n is the low-dimensional representation matrix of cells, recording the projection of cell expression data into the concepts defined by W (v) ; is the transpose of matrix A (v) ; I m is the identity matrix; ||·|| F is the Frobenius norm.
[0078] Calculate the distance between each cell based on the spatial position coordinates, and construct a spatial position graph considering the high-order information between cells;
[0079] The Laplacian kernel can capture the non-linear relationships in the data and is not easily affected by outliers. We use the Laplacian kernel to calculate the similarity matrix K of cells in the spatial position Y, and use the p-nearest neighbor graph to make similar cells close enough in the cell space to obtain the cell space network G. The multi-hop and other indirect relationships between cells more accurately characterize the topological structure of the network. Since the cell network G is too sparse, we use the pointwise mutual information matrix as the high-order adjacency matrix of the cell space network G (denoted by M). And, the element m ij in M is defined as:
[0080]
[0081] where, G ij is the element in the i-th row and j-th column of network G; l is the total number of cells; d lis the degree value of the l-th cell in network G; k is the number of non-negative samples (default is 1). M, as the high-order topological structure of network G, enhances the representation ability of the similarity relationship between cells.
[0082] Integrate the association matrices of multiple base clustering results to obtain an element-weighted integrated graph;
[0083] Integrate the association matrices of multiple base clustering results using an element-wise weighted integration strategy, and the element-wise weighted integration strategy includes:
[0084] Construct a base clustering association matrix, and its formula is:
[0085]
[0086] where cls d (x i ) is the category to which the i-th cell belongs in the d-th clustering; cls d (x j ) is the category to which the j-th cell belongs in the d-th clustering;
[0087] Denote the initial base clustering association matrix as E0, and the formula of the base clustering association matrix is:
[0088]
[0089] where D is the total number of selected base clustering methods; E d is the d-th base clustering association matrix.
[0090] Determine the cell-level weights through the nearest neighbor set to generate a consensus association matrix.
[0091] The formula of the consensus association matrix is:
[0092]
[0093] where D is the total number of selected base clustering methods; d is the d-th base clustering method, and w d is the high-order weight matrix.
[0094] To utilize the different cell clustering results obtained by existing models, we introduce an element-wise ensemble learning strategy. Denote the d base clustering results as B = {π 1 , …, π d}, and cls d (x i ) represents the category to which the i-th cell belongs in the d-th clustering. Construct the association matrix E d of the d-th base clustering as:
[0095]
[0096] Given the significant diversity among cells, we found that different basic clustering methods perform unevenly when dealing with different types of cells. Some methods may perform excellently on certain types of cells but poorly on others. To make more full use of the advantages of each base clustering, we decided to weight the base clusters element by element at the cell level. For the set matrix E0, we constructed the nearest neighbor set τ of cell i, with the formula:
[0097]
[0098] where i1, …, i τ is the column number corresponding to the largest τ element in the i-th row of E0;
[0099] The high-order weight matrix w is calculated through the Jaccard similarity. Specifically, the high-order weight matrix of cell i is: d
[0100]
[0101] where τ i is the nearest neighbor set of cell i; cls d (x i ) is the category to which the i-th cell belongs in the d-th cluster; ∩ is the intersection of two sets; |·| represents the number of elements in a set. The larger w id is, the better the clustering effect of the d-th base clustering method on the i-th cell.
[0102] The consistent consensus association matrix is:
[0103]
[0104] Construct a bi-graph regularization objective function based on the spatial position graph and the element-weighted integration graph, and solve the objective function to obtain the final low-dimensional feature representation;
[0105] Construct a bi-graph regularization objective function based on the spatial position graph and the element-weighted integration graph. The construction of the bi-graph regularization function includes:
[0106] Obtain the initial objective function based on the association matrix of multiple base clustering results and the high-order network containing multi-modal high-order information of spatial multi-omics data;
[0107] In the initial objective function, incorporate the spatial position graph and the element-weighted integration graph to obtain the bi-graph regularization function.
[0108] The bi-graph regularization objective function is:
[0109]
[0110] Among them, λ is the first balance parameter; β is the second balance parameter; L E is the Laplacian matrix of the consensus association matrix; L M is the Laplacian matrix of the high-order adjacency matrix; α v is the weight of omics υ; X (v) is the feature matrix of different omics; A (v) is the anchor base of omics υ; W (v) is the association matrix between omics v and the concept-related anchor base;
[0111] P is the common low-dimensional cell representation matrix between different omics; P T is the transpose of matrix P; I m is the identity matrix; ||·|| F is the Frobenius norm.
[0112] The bi-graph regularization objective function is solved by an alternating optimization algorithm to obtain the final low-dimensional feature representation, including:
[0113] Updating the association matrix through singular value decomposition;
[0114] Updating the anchor base matrix through the gradient descent method;
[0115] Determining the final low-dimensional feature representation by solving a linear equation system.
[0116] Based on the final low-dimensional feature representation, clustering recognition is performed to obtain the spatial domain recognition result.
[0117] The k-means algorithm is used to perform clustering recognition on the final low-dimensional feature representation to obtain the spatial domain recognition result.
[0118] To achieve the performance of optimal results in multiple scenarios, we use the consensus association matrix E containing the clustering results of multiple method bases and the high-order network M containing the multi-modal high-order information of spatial multi-omics data to guide the learning of the cell low-dimensional representation matrix P, and add a graph regularization term to the original objective function to obtain the final objective function:
[0119]
[0120] Among them, λ is the first balance parameter; β is the second balance parameter; L E is the Laplacian matrix of the consensus association matrix; L M is the Laplacian matrix of the high-order adjacency matrix; α v is the weight of omics v; A (v) is the anchor base of omics v; W (v) is the association matrix between omics v and the concept-related anchor base; P is the common low-dimensional cell representation matrix between different omics; PT is the transpose of matrix P; I m is the identity matrix; ||·|| F is the Frobenius norm. When solving this problem, the alternating optimization algorithm can be used. We decompose the entire objective function into several sub-problems and optimize one variable alone while keeping other variables unchanged.
[0121] For the A (v) sub-problem, we fix all other variables as constants and can obtain by solving:
[0122]
[0123] where B (v) = X (v) P T W (v)T = X (v) [W (v) P] T , X (v) is the feature matrix of different omics; W (v) is the association matrix between omics v and concept-related anchor bases; P is the common low-dimensional cell representation matrix between different omics; P T is the transpose of matrix P; Further, the singular value decomposition (SVD) method is used to obtain
[0124] For the W (v) sub-problem, the update rule of W (v) can be derived by keeping other variables unchanged:
[0125]
[0126] where X (v) is the feature matrix of different omics; A (v) is the anchor base of omics v; W (v) is the association matrix between omics v and concept-related anchor bases; P is the common low-dimensional cell representation matrix between different omics; P T is the transpose of matrix P; (PPT) -1 is the inverse matrix of matrix PP T .
[0127] For the P sub-problem, we set the derivative of the formula of the objective function with respect to P equal to zero and solve to obtain the update rule of P:
[0128] P = sylvester(X, Y, Z)
[0129] Here, sylvester(·) represents solving the Sylvester equation; X, Y, Z are obtained by solving through the following formulas respectively:
[0130]
[0131] Y = λ[βL E +(1 - β)L M
[0132]
[0133] where λ is the first balance parameter; β is the second balance parameter; L E is the Laplacian matrix of the consensus correlation matrix; L M is the Laplacian matrix of the high - order adjacency matrix; α v is the weight of omics v; X (v) is the feature matrix of different omics; A (v) is the anchor base of omics v; W (v) is the correlation matrix between omics v and the concept - related anchor base; P is the common low - dimensional cell representation matrix between different omics; P T is the transpose of matrix P.
[0134] By performing k - means clustering on the finally learned low - dimensional representation P to identify specific cell types, the functional associations and heterogeneities in different regions within the tissue can be better understood.
[0135] Example 2
[0136] This example provides a spatial domain identification method based on spatial multi - omics data, including:
[0137] 1) Data collection: This invention is tested on a simulated three - omics dataset. The simulated three - omics data contains the expression profile information of three omics, namely RNA, ATAC, and ADT and spatial location information Perform data pre - processing on the three - omics data and obtain the base clustering results of SpatialGlue and scMIC on the three - omics dataset.
[0138] 2) Model construction: Construct the model according to the steps described above, and input the processed data and the base clustering results into the network model.
[0139] 3) Parameter update: We use an alternating iterative optimization strategy to solve the above - mentioned objective function. Specifically, decompose the entire objective function into several sub - problems, and optimize one variable alone while fixing other variables. As Figure 2 shown, the update order is A (υ) , P, W (υ) , α υ , the model achieved the best performance when λ = 1 and β = 0.9. In addition, the invention has the ability to adaptively adjust parameters for different datasets.
[0140] 4) Test results: As Figure 3 shown, compared with the baseline methods (SpatialGlue, scMIC), this embodiment shows superior performance in matching the ground truth. The single omics expression profile information (omics 1, omics 2, and omics 3) shows incomplete or noisy patterns. Specifically, spatial cluster 1 is determined by omics 1 and omics 3, spatial cluster 2 is determined by omics 1, omics 2, and omics 3, while spatial cluster 3 is uniquely determined by omics 2, and spatial cluster 4 is uniquely determined by omics 3. The augmented omics is obtained from the sparse RNA modality, which improves the initial pattern but is still lower than other methods in terms of accuracy and resolution. The present invention can accurately identify the spatial domain with minimal noise. The poor scMIC results may be related to the failure to fully utilize spatial information, which emphasizes the important role of spatial information in spatial multi-omics analysis. As Figure 4 shown, the present invention selects accuracy (ACC), normalized mutual information (NMI), purity, F1-score, and adjusted Rand index (ARI) to evaluate the spatial domain recognition performance. The present invention achieves the highest NMI and ARI scores, indicating that the present invention is superior to SpatialGlue, scMIC, and other omics, and has excellent clustering and spatial integration capabilities.
[0141] The specific optimization algorithm steps are shown in Table 1:
[0142] Table 1 Algorithm optimization process
[0143]
[0144]
[0145] Example 3
[0146] This embodiment also provides a spatial domain recognition system based on spatial multi-omics data for implementing the method described in Example 1 or Example 2. As Figure 5 shown, it includes:
[0147] A data acquisition module, where the multi-omics data includes the expression profile information and spatial position coordinates of different omics, and preprocesses the multi-omics data;
[0148] A feature extraction module that extracts a low-dimensional representation of consistency for the multi-omics data to obtain an association matrix of multiple base clustering results;
[0149] A position map construction module that calculates the distance between each cell based on the spatial position coordinates and constructs a spatial position map considering the high-order information between cells;
[0150] An integrated graph generation module integrates the correlation matrices of multiple base clustering results to obtain an element-weighted integrated graph;
[0151] An objective function solving module constructs a bi-graph regularization objective function based on the spatial position graph and the element-weighted integrated graph, and solves the objective function to obtain the final low-dimensional feature representation;
[0152] A spatial domain recognition module performs clustering recognition based on the final low-dimensional feature representation to obtain a spatial domain recognition result.
[0153] The same or similar reference numerals correspond to the same or similar components;
[0154] The terms describing the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0155] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for spatial domain recognition based on spatial omics data, characterized in that, Including: Obtain multi-omics data, where the multi-omics data includes expression profile information and spatial position coordinates of different omics, and perform data preprocessing on the multi-omics data; Extract a consistent low-dimensional representation of the multi-omics data to obtain an association matrix of multiple base clustering results; Calculate the distance between each cell based on the spatial position coordinates, and construct a spatial position map considering the high-order information between cells; Integrate the association matrices of multiple base clustering results to obtain an element-weighted integrated map; Construct a bi-graph regularization objective function based on the spatial position map and the element-weighted integrated map, and solve the objective function to obtain the final low-dimensional feature representation; Perform clustering recognition based on the final low-dimensional feature representation to obtain the spatial domain recognition result.
2. The spatial domain recognition method based on spatial omics data according to claim 1, wherein, The data preprocessing includes: Use the SCANPY package to remove cells outside the tissue section; Retain the top pre-set number of genes in the RNA-seq data of the expression profile information of different omics according to gene expression differences; Perform latent semantic indexing dimensionality reduction on the ATAC-seq data in the expression profile information of different omics; Perform normalization and scaling processing on the ADT data in the expression profile information of different omics; Fill in the sparse omics data through spatial similarity and gene expression similarity to fill in the missing expression values of the omics data.
3. The spatial domain recognition method based on spatial omics data according to claim 1, wherein, Adopt an anchor concept decomposition framework to extract a consistent low-dimensional representation of the multi-omics data. The formula of the anchor concept decomposition framework is: s.t.A (υ)T A (υ) = I m Among them, α c is the omics weight for adaptive learning; X (v) ∈R n×k is the feature matrix of different omics; A (v) ∈R k×m is the anchor base of a specific omics; is the transpose of matrix A (υ) ; W (v) ∈R m×d is the association matrix of anchor bases related to concepts; P ∈ R d×n is the low-dimensional representation matrix of cells; I m is the identity matrix; ||·|| F is the Frobenius norm.
4. The spatial domain recognition method based on spatial omics data according to claim 1, wherein Use an element-wise weighted integration strategy to integrate the association matrices of multiple base clustering results. The element-wise weighted integration strategy includes: Construct a base clustering association matrix, and its formula is: where, cls d (x i ) is the class to which the i-th cell belongs in the d-th cluster; cls d (x j ) is the class to which the j-th cell belongs in the d-th cluster; Denote the initial base clustering association matrix as E0, and the formula of the base clustering association matrix is: Among them, D is the total number of selected base clustering methods; E d is the d-th base clustering association matrix; Determine the cell-level weights through the nearest neighbor set τ to generate a consensus association matrix.
5. The spatial domain recognition method based on spatial omics data according to claim 4, wherein, The formula of the consensus association matrix is: where d is the number of base clustering methods; w d is the high-order weight matrix; is the transpose of matrix w d ; ⊙ represents the dot product of matrices.
6. The spatial domain recognition method based on spatial omics data according to claim 1, wherein Construct a bi-graph regularization objective function based on the spatial position map and the element-weighted integrated map. The construction of the bi-graph regularization function includes: Obtain the initial objective function based on the association matrix of multiple base clustering results and the high-order network containing the multi-modal high-order information of spatial multi-omics data; Obtain the bi-graph regularization function by taking the spatial position map and the element-weighted integrated map in the initial objective function.
7. The spatial domain recognition method based on spatial omics data according to claim 6, characterized in that, The bi-graph regularization objective function is: s.t.A (υ)T A (υ) = I m where, λ is the first balance parameter; β is the second balance parameter; L E is the Laplacian matrix of the consensus correlation matrix; L M is the Laplacian matrix of the high-order adjacency matrix; α v is the weight of omics v; X (v) is the feature matrix of different omics; A (v) is the anchor base of omics v; is the matrix A (υ) transpose; W (v) is the correlation matrix between omics v and the concept-related anchor base; P is the common cell low-dimensional representation matrix between different omics; P T is the transpose of the matrix P; I m is the identity matrix; ||·|| F is the Frobenius norm.
8. The spatial domain recognition method based on spatial omics data according to claim 1, characterized in that The bi-graph regularization objective function is solved through an alternating optimization algorithm to obtain the final low-dimensional feature representation, including: Decompose the bi-graph regularization objective function into three sub-problems, and solve the three sub-problems separately; Fix the preset variables of the first sub-problem as constants and solve to obtain the update rule of the first sub-problem; Keep the preset variables of the second sub-problem unchanged and derive to obtain the update rule of the second sub-problem; Solve the third sub-problem by setting the derivative to zero to obtain the update rule of the third sub-problem; Fuse based on the update rules of the first sub-problem, the second sub-problem, and the third sub-problem to obtain the final low-dimensional feature representation.
9. The spatial domain recognition method based on spatial omics data according to claim 1, characterized in that Adopt the k-means algorithm to perform clustering recognition on the final low-dimensional feature representation to obtain the spatial domain recognition result.
10. A spatial domain recognition system based on spatial omics data, for implementing the method according to any one of claims 1-9, characterized in that, Including: A data acquisition module, where the multi-omics data includes expression profile information and spatial position coordinates of different omics, and perform data preprocessing on the multi-omics data; Feature extraction module, which performs consistent low-dimensional representation extraction on multi-omics data to obtain the association matrix of multiple base clustering results; Position map construction module, which calculates the distance between each cell based on spatial position coordinates and constructs a spatial position map considering the high-order information between cells; Integrated graph generation module, which integrates the association matrices of multiple base clustering results to obtain an element-weighted integrated graph; Objective function solving module, which constructs a bi-graph regularization objective function based on the spatial position map and the element-weighted integrated graph, and solves the objective function to obtain the final low-dimensional feature representation; Spatial domain identification module, which performs clustering identification based on the final low-dimensional feature representation to obtain the spatial domain identification result.
Citation Information
Cited By
Spatial domain identification method and device based on multi-modal topology consistency
CN120932746A