Transcriptomics spatial domain identification method

Through a multi-task graph comparative learning method, combined with graph neural networks and zero-inflated negative binomial distribution models, the problems of poor data denoising and recognition in spatial transcriptomics are solved, high-precision spatial domain recognition and data completion are achieved, and the accuracy of transcriptomics analysis is improved.

CN120600121APending Publication Date: 2025-09-05NORTHEAST FORESTRY UNIV
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510673347.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing spatial transcriptomics methods have problems with low data denoising accuracy and poor recognition effect when identifying spatial domains. This is mainly due to the Dropout problem, which causes the information of gene expression differences to be hidden, affecting the application of transcriptomics in complex cellular microenvironment research.

Method used

A multi-task graph contrastive learning method is adopted, combined with graph neural network and contrastive representation learning mechanism, features are extracted through graph convolutional neural network, and the zero-inflated negative binomial distribution model is used to fit the reconstruction matrix. Combined with neighborhood graph augmentation and contrastive loss function, the model is optimized to improve recognition accuracy.

Benefits of technology

It effectively reduces the noise of gene expression data, improves the accuracy of spatial domain recognition, and achieves data completion in the case of missing data, thereby improving the accuracy and stability of transcriptomics analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600121A_ABST
    Figure CN120600121A_ABST
Patent Text Reader

Abstract

The invention discloses a transcriptomics spatial domain identification method, and belongs to the technical field of transcriptomics. The objective of the invention is to solve the problems of low data noise reduction precision and poor recognition effect of an existing spatial transcriptional spatial domain recognition method. The method comprises the following steps: firstly, obtaining an undirected neighborhood graph according to a gene expression matrix, obtaining embedded representation of the gene expression matrix by utilizing an encoder, obtaining a corresponding reconstruction matrix by utilizing a decoder, and further determining reconstruction loss; meanwhile, a ZINB model is used for fitting a reconstruction matrix, and a ZINB loss function is obtained; then, an augmented graph is constructed based on the undirected neighborhood graph, respective embedded matrixes are obtained through an encoder, the comparison loss of the undirected neighborhood graph and the comparison loss of the augmented graph are obtained through a comparison representation learning mechanism, and then the neighbor comparison loss is obtained; total target loss is obtained based on all losses, a joint optimization strategy is adopted for training, and after training of the whole model is completed, dimensionality reduction and spatial domain recognition are carried out on a generated reconstruction matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of transcriptomics, and particularly relates to a transcriptomics spatial domain identification method. Background Art

[0002] In multicellular organisms, different cell types play unique functional roles in specific tissues, and changes in their life cycles and functions are profoundly influenced by the microenvironment surrounding the cells. The spatial location information of cells and their spatial neighborhood structure in tissues are crucial for understanding the biological functions and disease mechanisms of cells. Although single-cell RNA sequencing technology (scRNA-seq) can obtain gene expression signals at cellular resolution, which has greatly promoted the development of epigenetics, this type of technology lacks the acquisition and retention of cell spatial location information during the sequencing process, making it impossible for downstream analysis to effectively measure the spatial topological relationship within the cell, thereby limiting its application potential in the field of complex cellular microenvironment research. The emergence of spatial transcriptomics technology (ST) combines gene expression data with the spatial location information of cells, making it possible to accurately understand the histological function of cells in the spatial dimension.

[0003] In spatial transcriptome research, accurately identifying spatial domains where gene expression is consistent with histological location is one of the core tasks of spatial transcriptome analysis. However, spatial transcriptome data suffers from high sparsity in the expression matrix, i.e., a large number of zero values ​​are present in the matrix. Numerous studies have demonstrated that the dropout problem is the primary cause of expression matrix sparsity. The dropout problem refers to the fact that at certain sites, although certain genes are actually expressed, their expression is not detected due to technical or biological reasons, resulting in zero values ​​in the data. The dropout problem hides a large amount of gene expression differences between cells, leading to the current research status of poor accuracy in spatial domain identification at the transcriptome level. Therefore, spatial transcriptomics urgently needs to design multi-task spatial transcriptome algorithms that combine data denoising with spatial domain identification. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems of low data noise reduction accuracy and poor recognition effect in existing spatial transcription spatial domain recognition methods.

[0005] A transcriptomics spatial domain identification method, comprising:

[0006] Step 1: Filter out the top M highly variable genes from the original gene expression data and obtain a feature matrix based on the M genes That is, the gene expression matrix, N represents the number of cells; according to the gene expression matrix, the undirected neighborhood graph G is obtained;

[0007] Step 2: The adjacency matrix A and the gene expression matrix X are taken as input and passed to the encoder. The encoder uses a single-layer graph convolutional neural network GCN to extract features and obtain the embedded representation Z. The obtained embedded representation Z is input into a GCN decoder with a symmetrical structure with the encoder GCN to obtain the gene expression matrix H, which is denoted as the reconstruction matrix H. The reconstruction loss l is determined based on the reconstruction matrix H and the gene expression matrix X. recon ;

[0008] Use the ZINB model to fit the reconstruction matrix H and obtain the ZINB loss function l zinb ;

[0009] Step 3: Name the original undirected neighborhood graph G as G1, and perform data augmentation on G1 by applying random perturbations to its node feature vectors to generate an augmented feature matrix While keeping the original topological structure unchanged, the augmented graph G2 is constructed; G1 and G2 are input into a shared parameter encoder respectively to obtain two intermediate embedding matrices Z 1 and Z 2 ;

[0010] For node v in G1 i Embed Get positive sample pairs. The sources of positive sample pairs include: (1) node v i The corresponding embedding in the augmented graph G2 (2) v in Figure G1 i The embedding of the neighbor nodes of in Indicates v i The set of neighbor nodes in G1; (3), graph G w Medium v i The embedding of the neighbor nodes of in Indicates v i The set of neighbor nodes in G2; the selection of node positive sample pairs in G2 is the same as that in G1;

[0011] For samples other than positive sample pairs as negative sample pairs, through each node v i Embedding z i Filter out incorrect negative sample pairs;

[0012] Based on G1, embedding between graphs G1 and G2 with anchor points The relevant contrastive loss function:

[0013]

[0014] in, The “1,1” in the superscript represents the treatment within the figure, and “1,2” represents the treatment between figures; Node v i The corresponding sample pair sets in G1 and G2, For node v i The number of sample pairs constructed; τ is the temperature parameter, θ(·) is the cosine similarity; intra-graphneighbor por represents the positive sample within the intra-graph neighborhood, intra-graph neg represents the negative sample within the intra-graph; inter-graphneighbor pos represents the positive sample within the inter-graph neighborhood, inter-graph neg represents the negative sample between graphs;

[0015] Based on G2, the contrast loss function is determined in the same way Then we get the neighbor comparison loss

[0016] Step 4: For the overall model from step 1 to step 3, a joint optimization strategy is used for training. The total target loss of the joint optimization strategy is L = α*l recon +β*l zinb +γ*l cos ; α, β, γ are weight balance coefficients;

[0017] After the overall model training is completed, the generated reconstruction matrix H is subjected to dimensionality reduction and spatial domain identification.

[0018] Furthermore, the process of obtaining the feature matrix X based on M genes includes:

[0019] The M gene expression matrices are normalized, logarithmically transformed, and scaled to ensure that they have unit variance and zero mean, thereby obtaining the feature matrix X.

[0020] Furthermore, the undirected neighborhood graph G = (V, E) is obtained according to the gene expression matrix, where V represents the set of N cells in the gene expression matrix {v1,…,v N}, Represents the set of edges between cells, and the corresponding adjacency matrix is ​​A∈{0,1} N×N The definition is as follows: If A ij =1, it means there is an edge (v i ,v j ); Based on the undirected neighborhood graph, determine the neighbor nodes for each cell.

[0021] Furthermore, based on the undirected neighborhood graph, in the process of determining neighbor nodes for each cell, the K-nearest neighbor method is used to connect each cell to its nearest k cells to form neighbor nodes.

[0022] Furthermore, the encoder uses a single-layer graph convolutional neural network GCN to extract features and obtain the embedded representation in, in It is the adjacency matrix A plus the self-loop term. The self-loop term is specifically a unit matrix with only the diagonal 1 and the rest are 0. D = ∑ j A ij is the degree matrix, W e and B e are the trainable weight matrix and bias vector in the encoder, respectively, and σ(·) is a nonlinear activation function.

[0023] Furthermore, the reconstruction loss l is determined based on the reconstruction matrix H and the gene expression matrix X. recon as follows:

[0024]

[0025] Among them, x i is the i-th sample of the gene expression matrix X, h i is the i-th sample in the reconstructed and denoised gene expression matrix H.

[0026] Furthermore, the ZINB model is used to fit the reconstruction matrix H and obtain the ZINB loss function l zinb The process includes:

[0027] The mean μ, the discreteness θ and the dropout rate π of the NB distribution corresponding to the ZINB model are calculated by three independent multi-layer perceptrons, and the ZINB model is used to fit the reconstructed matrix H; thus, the ZINB loss function l is obtained. zinb =∑-log(ZINB(X|π,μ,θ)).

[0028] Furthermore, the mean μ, the dispersion θ, and the dropout rate π of the NB distribution corresponding to the ZINB model are calculated through three independent multi-layer perceptrons. The process of fitting the reconstructed matrix H using the ZINB model includes:

[0029] Based on the reconstruction matrix H, the parameters required for the ZINB model are calculated through three independent multi-layer perceptrons:

[0030] μ=diag(S i )×exp(W μ H)

[0031] θ=exp(Wθ H)

[0032] π=σ(W π H)

[0033] Among them, W μ 、W θ and W π are the trainable weight matrices; H s is the reconstruction matrix of the input; diag represents the diagonal matrix; S i is the scaling factor;

[0034] The ZINB distribution ZINB(X|π,μ,θ) is determined based on the NB distribution NB(X|μ,θ).

[0035] Furthermore, for samples other than the positive sample pairs as negative sample pairs, through each node v i Embedding z i The process of filtering out false negative pairs includes:

[0036] Embedding feature Z based on gene expression 1 and Z 2 , using the k-means clustering algorithm, all cells are divided into Q categories and pseudo labels are generated y i Represents the category pseudo label; in the process of constructing negative sample pairs, all nodes with the same pseudo label are excluded.

[0037] Furthermore, after the overall model training is completed, the process of dimensionality reduction and spatial domain recognition of the generated reconstruction matrix H includes:

[0038] First, perform principal component analysis on the reconstructed matrix H and select the first P principal components to obtain the matrix after dimensionality reduction: Where N represents the number of cells and P represents the number of principal components;

[0039] Then for the matrix H after dimensionality reduction PCA , cluster analysis was performed using the Mclust algorithm, which determines the optimal distribution of data and its category division by maximizing the likelihood estimate.

[0040] Beneficial effects:

[0041] The transcriptomic spatial domain identification method of the present invention can achieve optimal distribution and category division on the basis of effectively improving the accuracy of data noise reduction; at the same time, the present invention can also achieve transcriptomic spatial domain identification with data completion effect when there is a certain amount of missing data. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1Flowchart for spatial domain identification in transcriptomics;

[0043] Figure 2 Logical block diagram for transcriptomics spatial domain identification;

[0044] Figure 3 Clustering effect diagram for slices of the human dorsolateral prefrontal cortex dataset;

[0045] Figure 4 Spatial domain visualization of the denoising effect on the human dorsolateral prefrontal cortex dataset. DETAILED DESCRIPTION

[0046] This paper proposes a spatial domain identification method for spatial transcriptomics data based on multi-task graph comparative learning. It takes the gene expression information and spatial position information of spatial transcriptomics data as input, and effectively reduces the noise of gene expression data based on graph neural network and comparative representation learning mechanism, and achieves high spatial domain identification accuracy. It has the ability to quickly and accurately process spatial transcriptomics data from different sequencing technology platforms.

[0047] Specific implementation method 1: Combination Figure 1 and Figure 2 To explain this embodiment,

[0048] This embodiment is a transcriptomics spatial domain identification method, comprising:

[0049] Step 1: Data preprocessing:

[0050] S1.1. Due to the high sparsity and high dimensionality of spatial transcriptome data, we first used to process the raw gene expression data.

[0051] Specifically, we use "Scanpy" or "Seurat_v3" to filter out the top M highly variable genes from the raw gene expression data. Then, the generated M gene expression matrices are normalized, logarithmically transformed, and scaled to ensure that they have unit variance and zero mean. Finally, the feature matrix input to the model is obtained. That is, the gene expression matrix, where N represents the number of cells and M represents the number of genes (feature number). During the experiment, M is set to 5000. The matrix consists of N M-dimensional eigenvectors, each of which corresponds to the gene expression information of a cell.

[0052] S1.2. In spatial domain recognition tasks, cells with similar gene expression are usually close to each other in space. Therefore, spatial information plays a crucial role in tasks such as cell type recognition and tissue substructure description. To better capture this spatial information, we model a set of gene expression matrices as an undirected neighborhood graph G = (V, E) with a pre-set number of neighbors, where V represents the set of N cells in the gene expression matrix {v1,…,v N}, Represents the set of edges between cells; the corresponding adjacency matrix A∈{0,1} B×B The definition is as follows: If A ij =1, it means there is an edge (v i ,v j ) and determine the edge connectivity by calculating a distance matrix, where the elements in the distance matrix are the Euclidean distances between any two cells. Then, using the K-nearest neighbor (KNN) method, each cell is connected to its k nearest cells to form neighbor nodes.

[0053] Step 2: Build a reconstruction model for distribution adaptive fitting:

[0054] S2.1. To efficiently reconstruct a denoised matrix that matches the real data, we design an autoencoder architecture based on a graph neural network (GNN):

[0055] Specifically, the adjacency matrix A and the gene expression matrix X are taken as input and passed into the encoder to obtain a low-dimensional embedding representation Z.

[0056] In the encoder, we use a single-layer graph convolutional neural network (GCN) to perform feature extraction and information aggregation, with the intermediate embedding dimension set to 256. Through the GCN's convolutional operation, the model can iteratively aggregate the gene expression features of each node and the spatial location information of its neighboring nodes. This approach not only enhances local spatial correlation but also extracts more representative low-dimensional embedding information, providing a solid foundation for subsequent reconstruction. Formally, the operation of the i-th layer of the graph convolutional neural network can be expressed as:

[0057]

[0058] in, in It is the adjacency matrix A plus the self-loop term. The self-loop term is specifically a unit matrix with only the diagonal 1 and the rest are 0. D = ∑ j A ij is the degree matrix, W e and Be are the trainable weight matrix and bias vector in the encoder, respectively. σ(·) is a nonlinear activation function (such as ReLU) used to introduce nonlinear features.

[0059] S2.2. After the encoder obtains the intermediate embedding representation Z, the embedded representation Z is input into a GCN decoder that is symmetrical with the encoder GCN structure to generate the denoised gene expression matrix H. The form is as follows:

[0060]

[0061] Among them, H represents the reconstructed denoised gene expression matrix, W d and B d are the trainable weight matrix and bias vector in the decoder, respectively.

[0062] In order to restore the denoised gene expression matrix H and the original gene expression matrix X to the greatest extent possible, we define the reconstruction loss as the root mean square error (RMSE) loss function, as follows:

[0063]

[0064] Among them, x i is the i-th sample of the gene expression matrix X, h i is the i-th sample in the reconstructed and denoised gene expression matrix H, and N is the total number of samples.

[0065] By minimizing this loss function, the reconstruction model (encoder + decoder) can reconstruct the denoised gene expression matrix H as accurately as possible and retain important expression information.

[0066] S2.3. Considering the high sparsity and overdispersion of biological data, we use the Zero-Inflated Negative Binomial (ZINB) model to fit the reconstructed matrix H output by the decoder. This ensures that the denoised matrix better conforms to the statistical characteristics of the real data. To achieve this goal, we use the ZINB model's distribution self-fitting loss function as the optimization objective.

[0067] Specifically, after the decoder outputs the matrix H, we use three independent multi-layer perceptrons (MLPs) to calculate the parameters required by the ZINB model. The input of these three MLPs is the denoised gene expression matrix H, and their outputs are:

[0068] The mean μ of the NB distribution: represents the expected value of gene expression;

[0069] The dispersion of the NB distribution, θ, controls the variance of gene expression;

[0070] Dropout rate π: describes the probability of the zero-inflated part.

[0071] The above parameters together form the zero-inflated negative binomial distribution.

[0072] The parameters are defined as follows:

[0073] μ=diag(S i )×exp(W μ H)

[0074] θ=exp(W θ H)

[0075] π=σ(W π H)

[0076] Among them, W μ 、W θ and W π are the trainable weight matrices; H is the input reconstruction matrix; diag represents the diagonal matrix; S i is the scaling factor.

[0077] The definitions of NB distribution and ZINB distribution are as follows:

[0078]

[0079] Among them, X is the observed data, Γ(X+θ) is the gamma function, X m is the observed value of X in m dimensions, δ(X m ) is the Dirac delta function;

[0080] Based on the above model, the final ZINB loss function is defined as follows:

[0081] l zinb =∑-log(ZINB(X|π,μ,θ))

[0082] By minimizing l zinb ,The zero-inflated negative binomial distribution model can effectively capture the sparsity and discreteness of gene expression data,,thereby generating a denoised gene expression matrix that is more consistent with the distribution of real biological data,,providing a higher quality data foundation for subsequent biological analysis.

[0083] Step 3: Learning contrastive representations of debiased images for embedding refinement:

[0084] S3.1. To improve the model's ability to represent node identification information, a novel contrastive representation learning framework is designed. The core idea is to maximize the consistency of node embeddings between two topological views, enabling the model to more effectively capture neighborhood structure information. The specific implementation process is as follows:

[0085] First, the original undirected neighborhood graph G is named G1, and data augmentation is performed on G1. The augmented feature matrix is ​​generated by applying random perturbations to its node feature vectors. At the same time, the original topological structure is kept unchanged, thereby constructing the augmented graph G2.

[0086] Next, g1 and G2 are input into a shared parameter encoder (encoder in S2.1) to obtain two intermediate embedding matrices Z 1 and Z 2 .

[0087] Subsequent contrastive learning aims to learn consistent embedding representations between the two different views. Specifically, in step S3.2, a neighborhood contrastive loss function is used to maximize the similarity of representations of identical nodes in the two views while minimizing the difference between representations of different nodes. This design helps the model pay more attention to node neighborhood information during learning, thereby improving its understanding of cell and gene expression characteristics.

[0088] S3.2, using neighbor contrast learning loss to ensure the embedding Z 1 and Z 2 To improve the consistency and discrimination, we first construct and compare positive and negative sample pairs.

[0089] Get positive samples, starting with node v in G1 i Embed For example, the sources of positive sample pairs are as follows:

[0090] (1) Node v i The corresponding embedding in the augmented graph G2

[0091] (2) v in Figure G1 i The embedding of the neighbor nodes of in Indicates v i The set of neighbor nodes in G1;

[0092] (3) In Figure G2, v i The embedding of the neighbor nodes of in Indicates v i The set of neighbor nodes in G2;

[0093] The selection of node positive sample pairs in G2 is the same as that in G1, so for v i The total number of positive sample pairs for anchor points is in Respectively represent v i The number of neighbor nodes in G1 and G2.

[0094] In order to solve the sampling bias problem in the selection of negative sample pairs in traditional contrastive learning (i.e., the problem caused by random sampling of negative samples without considering category information), a debiasing strategy is implemented. Specifically, each node v i Embedding z i Filter out false negative pairs:

[0095] First, based on the embedding feature Z of gene expression 1 and Z 2 , using the k-means clustering algorithm, all cells are divided into Q categories and pseudo labels are generated y i Represents the category pseudo label; in the process of constructing negative sample pairs, all nodes with the same pseudo label are excluded to reduce the interference of misclassified negative samples. By default, Q is 40.

[0096] Between graphs G1 and G2, embedded with anchors The relevant contrast loss function is defined as follows:

[0097]

[0098] in, The “1,1” in the superscript represents the treatment within the figure, and “1,2” represents the treatment between figures; Node v i The corresponding sample pair sets in G1 and G2, For node v i The number of sample pairs constituted by the network; τ is the temperature parameter, θ(·) is the cosine similarity; intra-graphneighbor por represents the positive sample within the intra-graph neighborhood, intra-graph neg represents the negative sample within the intra-graph; inter-graphneighbor pos represents the positive sample within the inter-graph neighborhood, and inter-graph neg represents the negative sample between graphs.

[0099] Further subdividing the positive and negative samples in the contrastive loss into positive samples within the intra-graph / inter-graph neighborhood (neighborpos) and negative samples outside the neighborhood can more accurately reflect the semantic and structural relationships between nodes in the graph structure. This fine-grained division significantly improves the quality of positive and negative samples, reduces the interference of pseudo-negative samples, and effectively enhances the model's ability to learn local structural consistency and cross-graph reinforcement stability. At the same time, it helps improve the discriminability and generalization of feature representations, providing a more robust representation learning foundation for downstream tasks such as clustering and classification in graph-structured data such as spatial transcriptomes.

[0100] The goal of minimizing the loss function is to enhance the similarity between positive samples while reducing the similarity between negative samples, thereby improving the representation and discrimination capabilities of the model.

[0101] Given the structural similarity between G1 and G2, the loss function and This ensures that the contrastive learning framework maintains symmetry when processing the two graph views.

[0102] The final neighbor contrastive loss is defined as follows:

[0103]

[0104] Step 4. Clustering via Joint Embedding Optimization:

[0105] S4.1. To improve the generalization performance of the overall model architecture, we adopted a joint optimization strategy during the overall model training process. This strategy comprehensively considers the reconstruction loss, the zero-inflated negative binomial distribution (ZINB) loss, and the contrastive learning loss to achieve the goal of multi-task balanced optimization. The overall objective loss function can be expressed as:

[0106] L=α*l recon +β*l zinb +γ*l cos

[0107] Among them, l recon represents the reconstruction loss, l zinb represents the ZINB reconstruction loss, l cos Denotes contrast and learning loss, and α, β, and γ are weight balancing coefficients. During the experiment, α, β, and γ were set to 20, 0.5, and 0.5, respectively. Adam was used as the optimizer algorithm during training, with a fixed-step learning rate decay strategy, and 600 training epochs.

[0108] S4.2. After model training is complete, we perform further dimensionality reduction and spatial domain recognition analysis on the generated reconstruction matrix H to achieve more accurate spatial domain recognition and reveal the underlying structure of cell types. The specific process is as follows:

[0109] Principal Component Analysis (PCA): First, we perform principal component analysis (PCA) on the reconstructed matrix H to reduce the dimensionality of the data and remove redundant information. Select the first P principal components and the matrix after dimensionality reduction is Where N represents the number of cells, P represents the number of principal components, and P is set to 20 during the experiment.

[0110] Mclust-based spatial domain identification: after dimension reduction, the matrix H PCA , we used the Mclust algorithm in the R package for cluster analysis, which determines the optimal distribution of data and its category division by maximizing the likelihood estimate.

[0111] Once the optimal distribution and classification are achieved, the transcriptomic spatial domain identification is achieved. It should be noted that the present invention not only achieves the process of transcriptomic spatial domain identification, but also achieves data noise reduction and data completion.

[0112] We selected one of the slices in the human dorsolateral prefrontal cortex dataset (DLPFC) as a case study and first performed a cluster analysis. The results are shown in the following figure. Figure 3 We then used the known specific marker genes in the DLPFC dataset to visualize the original gene expression data in the spatial domain, and the results are shown in Figure 4 Then, the gene expression data after noise reduction of the present invention is used for spatial domain visualization, and the results are shown in the first row. Figure 4 The second line shows the effect of data completion:

[0113] Using the human dorsolateral prefrontal cortex (DLPFC) dataset as a benchmark experimental dataset, we employed a masking strategy based on the Bernoulli distribution to randomly mask the feature vectors of each cell, thereby manually simulating noise. Specifically, we masked the input model feature matrix with masking rates of 0.1, 0.3, 0.5, 0.7, and 0.9. Subsequently, we calculated the root mean square error (MSE) between the denoised matrix output by the model and the unmasked feature matrix to measure the model's data completion ability under varying noise interference. A lower MSE indicates better model completion performance. The results are shown in Table 1.

[0114] Table 1 Comparison of data completion performance of the model with other methods under different mask probabilities

[0115]

[0116] Note: The indicator is MSE, the lower the better.

[0117] The above examples are merely illustrative of the calculation model and process of the present invention and are not intended to limit the embodiments of the present invention. Persons skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. This list of embodiments is not exhaustive; however, any obvious variations or modifications derived from the technical solution of the present invention remain within the scope of protection of the present invention.

Claims

1. A transcriptomics spatial domain identification method, characterized in that: include: Step 1: Filter out the top M highly variable genes from the original gene expression data and obtain a feature matrix based on the M genes That is, the gene expression matrix, N represents the number of cells; according to the gene expression matrix, the undirected neighborhood graph G is obtained; Step 2: The adjacency matrix A and the gene expression matrix X are taken as input and passed to the encoder. The encoder uses a single-layer graph convolutional neural network GCN to extract features and obtain the embedded representation Z. The obtained embedded representation Z is input into a GCN decoder with a symmetrical structure with the encoder GCN to obtain the gene expression matrix H, which is denoted as the reconstruction matrix H. The reconstruction loss l is determined based on the reconstruction matrix H and the gene expression matrix X. recon ; Use the ZINB model to fit the reconstruction matrix H and obtain the ZINB loss function l zinb ; Step 3: Name the original undirected neighborhood graph G as G1, and perform data augmentation on G1 by applying random perturbations to its node feature vectors to generate an augmented feature matrix While keeping the original topological structure unchanged, the augmented graph G2 is constructed; G1 and G2 are input into a shared parameter encoder respectively to obtain two intermediate embedding matrices Z 1 and Z 2 ; For node v in G1 i Embed Get positive sample pairs. The sources of positive sample pairs include: (1) node v i The corresponding embedding in the augmented graph G2 (2) v in Figure G1 i The embedding of the neighbor nodes of in Indicates v i The set of neighbor nodes in G1; (3) v in graph G2 i The embedding of the neighbor nodes of in Indicates v i The set of neighbor nodes in G2; the selection of node positive sample pairs in G2 is the same as that in G1; For samples other than positive sample pairs as negative sample pairs, through each node v i Embedding z i Filter out incorrect negative sample pairs; Based on G1, embedding between graphs G1 and G2 with anchor points The relevant contrastive loss function: in, The "1,1" in the superscript represents the treatment within the figure, and "1,2" represents the treatment between figures; Node v i The corresponding sample pair sets in G1 and G2, For node v i The number of sample pairs constructed; τ is the temperature parameter, θ(·) is the cosine similarity; intra-graphneighbor por represents the positive sample within the intra-graph neighborhood, intra-graph neg represents the negative sample within the intra-graph; inter-graphneighbor pos represents the positive sample within the inter-graph neighborhood, inter-graph neg represents the negative sample between graphs; Based on G2, the contrast loss function is determined in the same way Then we get the neighbor comparison loss Step 4: For the overall model from step 1 to step 3, a joint optimization strategy is used for training. The total target loss of the joint optimization strategy is L = α*l recon +β*l zinb +γ*l cos ; α, β, γ are weight balance coefficients; After the overall model training is completed, the generated reconstruction matrix H is subjected to dimensionality reduction and spatial domain identification.

2. A transcriptomics spatial domain identification method according to claim 1, characterized in that, The process of obtaining the feature matrix X based on M genes includes: The M gene expression matrices are normalized, logarithmically transformed, and scaled to ensure that they have unit variance and zero mean, thereby obtaining the feature matrix X.

3. A transcriptomics spatial domain identification method according to claim 1, characterized in that: The undirected neighborhood graph G = (V, E) obtained from the gene expression matrix, where V represents the set of N cells in the gene expression matrix {v1,…,v N }, Represents the set of edges between cells, and the corresponding adjacency matrix is ​​A∈{0,1} N×N The definition is as follows: If A ij =1, it means there is an edge (v i ,v j ); Based on the undirected neighborhood graph, determine the neighbor nodes for each cell.

4. A transcriptomics spatial domain identification method according to claim 3, characterized in that, Based on the undirected neighborhood graph, in the process of determining neighbor nodes for each cell, the K nearest neighbor method is used to connect each cell to its nearest k cells to form neighbor nodes.

5. The method for identifying transcriptomic spatial domains according to claim 1, wherein: The encoder uses a single-layer graph convolutional neural network GCN to extract features and obtain the embedded representation in, in It is the adjacency matrix A plus the self-loop term. The self-loop term is specifically a unit matrix with only the diagonal 1 and the rest are 0. D = ∑ j A ij is the degree matrix, W e and B e are the trainable weight matrix and bias vector in the encoder, respectively, and σ(·) is a nonlinear activation function.

6. A transcriptomics spatial domain identification method according to claim 1, characterized in that: Determine the reconstruction loss l based on the reconstruction matrix H and the gene expression matrix X recon as follows: Among them, x i is the i-th sample of the gene expression matrix X, h i is the i-th sample in the reconstructed and denoised gene expression matrix H.

7. The method for identifying transcriptomic spatial domains according to claim 1, wherein: Use the ZINB model to fit the reconstruction matrix H and obtain the ZINB loss function l zinb The process includes: The mean μ, the discreteness θ and the dropout rate π of the NB distribution corresponding to the ZINB model are calculated by three independent multi-layer perceptrons, and the ZINB model is used to fit the reconstructed matrix H; thus, the ZINB loss function l is obtained. zinb =∑-log(ZINB(X|π,μ,θ)).

8. A transcriptomics spatial domain identification method according to claim 7, characterized in that: The mean μ, the dispersion θ, and the dropout rate π of the NB distribution corresponding to the ZINB model are calculated through three independent multi-layer perceptrons. The process of fitting the reconstructed matrix H using the ZINB model includes: Based on the reconstruction matrix H, the parameters required for the ZINB model are calculated through three independent multi-layer perceptrons: μ=diag(S i )×exp(W μ H) θ=exp(W θ H) π=σ(W π H) Among them, W μ 、W θ and W π are the trainable weight matrices; H s is the reconstruction matrix of the input; diag represents the diagonal matrix; S i is the scaling factor; The ZINB distribution ZINB(X|π,μ,θ) is determined based on the NB distribution NB(X|μ,θ).

9. A transcriptomics spatial domain identification method according to any one of claims 1 to 8, characterized in that: For samples other than positive sample pairs as negative sample pairs, through each node v i Embedding z i The process of filtering out false negative pairs includes: Embedding feature Z based on gene expression 1 and Z 2 , using the k-means clustering algorithm, all cells are divided into Q categories and pseudo labels are generated y i Represents the category pseudo label; in the process of constructing negative sample pairs, all nodes with the same pseudo label are excluded.

10. A transcriptomics spatial domain identification method according to claim 9, characterized in that: After the overall model training is completed, the process of dimensionality reduction and spatial domain identification of the generated reconstruction matrix H includes: First, perform principal component analysis on the reconstructed matrix H and select the first P principal components to obtain the matrix after dimensionality reduction: Where N represents the number of cells and P represents the number of principal components; Then for the matrix H after dimensionality reduction PCA , cluster analysis was performed using the Mclust algorithm, which determines the optimal distribution of data and its category division by maximizing the likelihood estimate.

Citation Information

Cited By

  • Spatial transcriptome data spatial domain identification method based on multi-space self-supervised contrast learning

    CN120954500A

  • Multi-space self-supervised contrastive learning method for spatial transcriptome data spatial domain identification

    CN120954500B

  • Spatial domain identification method and device for biological space transcriptome slice

    CN121191599A

  • Spatial domain identification method and apparatus for biological spatial transcriptome slices

    CN121191599B

  • Spatial transcriptome data analysis method based on artificial intelligence

    CN121260260A