Niche recognition method and device, nonvolatile storage medium and electronic equipment

By combining graph neural networks with embedded neural networks and gating units to identify niches, the problem of incomplete feature information in spatial transcriptome analysis was solved, achieving high-precision niche identification of the target sample microenvironment and improving the comprehensiveness of tissue heterogeneity understanding.

CN122493954APending Publication Date: 2026-07-31SHENZHEN HUADA GENE INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN HUADA GENE INST
Filing Date
2025-01-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies in spatial transcriptome analysis fail to comprehensively analyze spatial location data and gene expression data, resulting in the inability to fully obtain niche and other characteristics in the sample data, thus limiting a comprehensive understanding of the tissue heterogeneity of the target sample.

Method used

By employing graph neural networks combined with embedded neural networks and gating units, an ecological niche identification map is constructed. This map comprehensively analyzes spatial location, gene expression, cell type, cell interaction, and pathway information. Graph attention networks are used to enhance the network's ability to process complex data, thereby improving the accuracy of ecological niche identification.

Benefits of technology

It achieves high-precision distribution identification of cell types, cell states and their molecular characteristics in the microenvironment of target samples, solves the problem of incomplete sample data feature information, and improves the comprehensiveness of tissue heterogeneity understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493954A_ABST
    Figure CN122493954A_ABST
Patent Text Reader

Abstract

This application discloses a niche identification method, apparatus, non-volatile storage medium, and electronic device. The method includes: acquiring spatial transcriptome data of a target sample, the spatial transcriptome data including multiple spatial points and the spatial location of each spatial point; extracting biological features from each spatial point in the spatial transcriptome data; constructing a target identification map structure based on the spatial location and biological features of each spatial point; and processing the target identification map structure using a target model to obtain a niche identification map corresponding to the target sample. The niche identification map indicates the distribution of cell types, cell states, and molecular characteristics in the microenvironment corresponding to the target sample. This application solves the technical problem of incomplete tissue heterogeneity and other characteristic information of the target sample obtained due to the inability to comprehensively analyze various feature data in the sample data in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more specifically, to a niche identification method, apparatus, non-volatile storage medium, and electronic device. Background Technology

[0002] In related technologies, spatial transcriptome analysis mainly relies on spatial location and gene expression data in the sample data. However, it fails to conduct a comprehensive analysis of spatial location and gene expression data, resulting in the inability to fully obtain niche and other characteristics in the sample data, thus limiting the comprehensive understanding of the tissue heterogeneity of the target sample.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a niche identification method, apparatus, non-volatile storage medium, and electronic device to at least solve the technical problem of incomplete characteristic information such as tissue heterogeneity of the target sample due to the inability to comprehensively analyze various characteristic data in sample data in related technologies.

[0005] According to one aspect of the embodiments of this application, a niche identification method is provided, comprising: acquiring spatial transcriptome data of a target sample, the spatial transcriptome data including multiple spatial points and the spatial location of each spatial point; extracting biological features of each spatial point in the spatial transcriptome data; constructing a map structure to be identified based on the spatial location and biological features of each spatial point; and processing the map structure to be identified through a target model to obtain a niche identification map corresponding to the target sample, wherein the niche identification map is used to indicate the distribution of cell types, cell states and their molecular features in the microenvironment corresponding to the target sample.

[0006] Optionally, the target model includes an embedded neural network model, a graph neural network model, and multiple gating units, wherein the embedded neural network model includes multiple embedded neural network layers, and the graph neural network model includes multiple graph neural network layers; each gating unit is connected to the output of an embedded neural network layer and the output of a graph neural network layer, respectively, and the output of the gating units other than the target gating unit is connected to the input of a graph neural network layer, and the target gating unit is the gating unit connected to the output of the last embedded neural network layer in the embedded neural network model and the output of the last graph neural network layer in the graph neural network model.

[0007] Optionally, processing the graph structure to be identified using a target model to obtain the niche partitioning map corresponding to the target sample includes: processing the node feature information in the graph structure to be identified using an embedded neural network model to obtain latent variables, wherein the latent variables include intermediate latent variables and target latent variables, the target latent variables include a low-dimensional feature representation of the feature information output by the embedded neural network model, and the intermediate latent variables are the intermediate processing results of the embedded neural network model when processing feature information; processing the graph structure to be identified using a graph neural network model, and updating the intermediate hidden state of the graph neural network model through gating units and intermediate latent variables during the processing of the graph structure to be identified to obtain the target hidden state output by the graph neural network model, wherein the target hidden state is a representation that integrates biological features and structural features of the graph structure to be identified; and determining the niche partitioning map based on the target hidden state.

[0008] Optionally, the embedded neural network model is constructed using an autoencoder or variational autoencoder to map feature information to a low-dimensional space to obtain the target latent variable.

[0009] Optionally, the embedded neural network model includes multiple embedded neural network layers, the graph neural network model includes multiple graph structure network layers, and the gating unit includes an update gate, a reset gate, and a combined gating layer. The combined gating layer includes multiple fully connected layers, used to send gating signals to the update gate and reset gate, and to concatenate the latent variables output by each embedded neural network layer and the hidden states output by each graph neural network layer to obtain the concatenation result. The update gate is used to determine the weights of the information provided by the hidden states retained in the concatenation result, as well as the weights of the information provided by the latent variables and feature vectors added, based on the gating signals. The reset gate is used to determine the contribution level of the hidden states during the concatenation process based on the gating signals.

[0010] Optionally, the loss of the target model includes embedding loss, clustering embedding loss, and clustering graph attention loss, wherein the clustering embedding loss is used to reflect the difference between the first predicted clustering distribution result of the embedded neural network model and the target clustering distribution result, and the clustering graph attention loss is used to reflect the difference between the second predicted clustering distribution result of the graph neural network model and the target clustering distribution result.

[0011] Optionally, the biological characteristics of each spatial point in the spatial transcriptome data include the cell type characteristics of each spatial point; extracting the biological characteristics of each spatial point in the spatial transcriptome data includes: collecting expression data of the target sample through a spatial transcriptomics platform, and determining gene expression characteristics based on the expression data, wherein the expression data is used to characterize the gene expression profile of the cell population at each spatial point in the tissue of the target sample; performing single-cell RNA sequencing on the target sample to obtain the reference cell type expression profile of the target sample, and determining the biological characteristics of each spatial point based on the gene expression profile of the cell population at each spatial point and the reference cell type expression profile.

[0012] Optionally, the biological features of each spatial point in the spatial transcriptome data include the biological pathway features of each spatial point; extracting the biological features of each spatial point in the spatial transcriptome data includes: for each spatial point in the spatial transcriptome data, calculating the average value of gene expression values ​​of all biological pathways corresponding to each spatial point, and using the average value corresponding to each spatial point as the biological pathway feature of each spatial point.

[0013] Optionally, the biological characteristics of each spatial point in the spatial transcriptome data include the intercellular interaction characteristics of each spatial point; extracting the biological characteristics of each spatial point in the spatial transcriptome data includes: determining the receptor expression value corresponding to each spatial point, and the ligand influence values ​​corresponding to each spatial point and a preset number of adjacent spatial points; determining the intercellular interaction characteristics of each spatial point based on the receptor expression value and the ligand influence values ​​corresponding to each of the multiple adjacent spatial points.

[0014] According to another aspect of the embodiments of this application, an ecological niche identification device is also provided, comprising: a first processing module for acquiring spatial transcriptome data of a target sample, the spatial transcriptome data including multiple spatial points and the spatial location of each spatial point; a second processing module for extracting biological features of each spatial point in the spatial transcriptome data; a third processing module for constructing a map structure to be identified based on the spatial location and biological features of each spatial point; and a fourth processing module for processing the map structure to be identified through a target model to obtain an ecological niche identification map corresponding to the target sample, wherein the ecological niche identification map is used to indicate the distribution of cell types, cell states and their molecular features in the microenvironment corresponding to the target sample.

[0015] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, wherein a program is stored in the non-volatile storage medium, wherein the program controls the device where the non-volatile storage medium is located to execute a niche identification method when it runs.

[0016] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes an economist identification method during runtime.

[0017] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that implements a niche identification method when executed by a processor.

[0018] In this embodiment, spatial transcriptome data of the target sample is acquired, including multiple spatial points and the spatial location of each point. Biological features of each spatial point are extracted from the spatial transcriptome data. A map structure to be identified is constructed based on the spatial location and biological features of each spatial point. The map structure is then processed using a target model to obtain a niche identification map corresponding to the target sample. This niche identification map indicates the distribution of cell types, cell states, and molecular characteristics in the microenvironment corresponding to the target sample. By constructing the map structure based on spatial location information and biological features from the spatial transcriptome data, the spatial and biological features of the sample are comprehensively analyzed, achieving the technical effect of obtaining a high-precision niche identification map. This solves the technical problem of incomplete tissue heterogeneity and other characteristic information of the target sample due to the inability to comprehensively analyze various feature data in the sample data in related technologies. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0020] Figure 1 This is a schematic diagram of the structure of a computer terminal (mobile terminal) according to an embodiment of this application;

[0021] Figure 2 This is a flowchart illustrating a niche identification method provided according to an embodiment of this application;

[0022] Figure 3 This is a schematic diagram of the architecture of a niche identification model provided according to an embodiment of this application;

[0023] Figure 4 This is a schematic diagram of a gate control unit according to an embodiment of this application;

[0024] Figure 5 This is a spatial transcriptome gene quantity heatmap provided according to an embodiment of this application;

[0025] Figure 6 This is a cell type annotation diagram provided according to an embodiment of this application;

[0026] Figure 7 This is a niche identification map provided according to an embodiment of this application;

[0027] Figure 8 This is a schematic diagram of the number of cells of different cell types in each niche according to the embodiments of this application;

[0028] Figure 9 This is a schematic diagram illustrating the spatial distribution of an ecological feature according to an embodiment of this application;

[0029] Figure 10 This is an ecological niche identification map based on embodiments of this application, which uses only gene expression as a feature for accuracy.

[0030] Figure 11 This is a niche identification map provided in the embodiments of this application without using a gating mechanism;

[0031] Figure 12 This is a schematic diagram of the structure of a niche identification device provided according to an embodiment of this application. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0034] The understanding of tumors within the medical field has evolved from a simple cluster of abnormally proliferating cells to a highly ordered "organ" entity. Modern research views tumors as complex cellular aggregates, comprising malignant cells, stromal cells, and immune cells. These cell populations not only exhibit significant heterogeneity but also collectively constitute the tumor microenvironment (TME). The TME is a dynamic region containing signaling networks that both inhibit and promote tumor growth, collectively regulating tumor progression. Although the specific composition of the TME varies significantly among different tumor types, research has revealed some universal characteristics, such as the presence of various immune cells and stromal cells, the peripheral vascular system, the extracellular matrix (ECM), and the role of various soluble factors.

[0035] To delve deeper into the unique properties of the tumor microenvironment (TME), research in this field has expanded beyond molecular characteristics and composition to include the spatial relationships and structural distribution of these factors. Spatial profiling of the TME primarily encompasses the following aspects: the spatial diversity and proportion of cells within tumor partitions; the distance between cells and their functionally related neighboring cells; and the spatial patterns of direct intercellular interactions and autocrine and paracrine signaling. While early studies have provided preliminary insights into the basic composition and spatial structure of the TME, modern advanced technologies have enabled precise, high-throughput, multi-dimensional, and high-resolution depictions of the TME. A deeper understanding of the TME has facilitated improvements in clinical prognosis and the optimization of immunotherapy strategies.

[0036] However, cancer-related cell-cell interactions can exist at different levels, including direct interactions between tumor cells and their surrounding non-tumor cells within the "microenvironment," as well as indirect interactions between tumor cells and more distant cells. The microenvironment is increasingly recognized by evidence as playing a crucial role in cancer phenotypes such as proliferation, invasion, metastasis, and drug resistance. Furthermore, the high heterogeneity of tumor cells themselves makes determining which specific subset of tumor cells directly interacts with surrounding non-tumor cells particularly challenging.

[0037] In deconstructing the tumor microenvironment (TME), various computational methods have been developed to analyze transcriptome sequencing data in tumor samples. These methods, by establishing gene expression profile models for specific cell types or applying deep learning techniques, reflect the complexity of the TME's cellular, molecular, and genetic heterogeneity, significantly advancing our understanding of the TME's cellular composition. Tumor pathology may depend on the spatial organization of the tumor stroma, just as the function of healthy tissue depends on the spatial organization of cells. While spatial technologies provide abundant cellular and neighborhood information, analyzing this information presents new challenges, particularly in identifying biologically significant microenvironments from rich spatial data and describing the correlation between microenvironments and disease.

[0038] To address the aforementioned issues, various spatial clustering and feature extraction algorithms have been developed in related technologies. Specifically, in spatial transcriptomics analysis, clustering analysis is often used to describe tissue morphology and structure. Conventional clustering analysis, such as the clustering workflow proposed for scRNA-seq analysis (e.g., Seurat V3 and SCANPY), is commonly used in spatial transcriptomics analysis. Traditional Leiden or Louvain clustering is typically applied to feature sets generated through dimensionality reduction using methods such as Principal Component Analysis (PCA), t-distributed stochastic neighbor embedding (t-SNE), and uniform manifold approximation and projection (UMAP) of gene expression data. However, these methods only use gene expression information and ignore the strength of the ST dataset.

[0039] To combine spatial distance or tissue morphology information with gene expression profiles, several spatial clustering and feature extraction algorithms have been developed, including BayesSpace, MUSE, BANKSY, SpaGCN, and SEDR. Among these, BayesSpace is a fully Bayesian model with a Markov random field that leverages neighborhood structure in spatial transcriptomics data to improve resolution to the sub-site level. BayesSpace divides observed points into multiple equally sized sub-sites, infers gene expression at each sub-site, and maintains the total expression level of the original points unchanged. By integrating spatial information and gene expression data, BayesSpace improves the resolution and accuracy of spatial transcriptomics research.

[0040] MUSE identifies high-resolution latent subpopulations in data by constructing a multi-view autoencoder that combines multimodal information. It employs a deep autoencoder and attention mechanisms to align the self-representations of different views. This approach enhances the model's nonlinear fitting capabilities while maintaining consistency and complementarity in multi-view learning. It can discover organizational subpopulations that might be missed by any single modality in the data, compensate for modality-specific noise issues, and enhance the overall analysis of multimodal datasets.

[0041] BANKSY unifies the tasks of cell type clustering and tissue region segmentation. The core of the algorithm lies in constructing a spatial representation of the transcriptome of a cell and its surrounding microenvironment, representing the cell state and the microenvironment, respectively. BANKSY combines the weighted average of cell expression with the expression of neighboring cells within its microenvironment as an enhancing feature for spatial clustering, thereby identifying niche-dependent cell subtypes and regions with consistent gene expression.

[0042] SpaGCN aggregates gene expression information for each sample point (spot) using graph convolution operations, while also considering information from neighboring sample points. This enables the algorithm to identify spatial domains with consistent expression and histological characteristics. SpaGCN also performs domain-guided differential expression (DE) analysis to detect genes with enriched expression patterns in specific spatial domains.

[0043] SEDR constructs low-dimensional latent representations of gene expression using a deep autoencoder, while embedding corresponding spatial information into these representations using a variational map autoencoder. SEDR is also suitable for batch ensembles, enabling the identification of heterogeneous subregions within visually homogeneous tumor regions.

[0044] However, while the aforementioned spatial transcriptomics analysis methods primarily rely on spatial location and gene expression data, they often fail to fully utilize the rich information contained within the slices, such as cell-cell interactions. This information is closely related to the tumor microenvironment and cell dynamics. Therefore, although these technologies perform well in capturing spatial and expression data, they have limitations when comprehensively considering more complex biological characteristics such as cell interactions.

[0045] In addition, some related technologies suffer from the problem of considering only gene expression data. When only gene expression data is considered, different types of cells often exist within the same tissue domain, but their gene expression patterns may differ significantly. This means that if the analytical method relies solely on gene expression data, it may be unable to accurately distinguish and interpret the complex relationships and functional differences between these cell types, thus limiting a comprehensive understanding of tissue heterogeneity.

[0046] Moreover, the algorithmic architecture in these technologies is relatively simple and insufficient to fully learn and characterize the complex information of cells in tissues. This means that when clustering and classifying tissue cells, these technologies cannot fully capture the decisive biological characteristics, thus affecting the accuracy and interpretability of the final analysis.

[0047] To address the aforementioned issues, this application provides a solution that combines various biological information such as spatial location, gene expression, cell type, cell interactions, and pathway information. Graph neural networks are used to analyze the tumor microenvironment, enabling a deeper understanding of intercellular interactions and the complexity of the microenvironment, and accurately identifying tumor microenvironment niches. Furthermore, the method provided in this application also utilizes Graph Attention Networks (GAT). By introducing feature embedding representations and gating mechanisms, it not only enhances the network's ability to process complex data but also improves the accuracy of identifying tumor microenvironment niches. Detailed explanation follows.

[0048] According to an embodiment of this application, a method embodiment for niche identification is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0049] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal for implementing a niche identification method is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0050] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as a form of processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0051] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the niche identification method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned niche identification method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0052] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0053] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.

[0054] Under the above operating environment, embodiments of this application provide a niche identification method, such as... Figure 2 As shown, the method includes the following steps:

[0055] Step S202: Obtain spatial transcriptome data of the target sample. The spatial transcriptome data includes multiple spatial points and the spatial location of each spatial point.

[0056] Step S204: Extract the biological characteristics of each spatial point in the spatial transcriptome data;

[0057] In the technical solution provided in step S204, the biological characteristics of each spatial point in the spatial transcriptome data include the cell type characteristics of each spatial point; extracting the biological characteristics of each spatial point in the spatial transcriptome data includes: collecting expression data of the target sample through a spatial transcriptomics platform, and determining gene expression characteristics based on the expression data, wherein the expression data is used to characterize the gene expression profile of the cell population at each spatial point in the tissue of the target sample; performing single-cell RNA sequencing on the target sample to obtain the reference cell type expression profile of the target sample, and determining the biological characteristics of each spatial point based on the gene expression profile of the cell population at each spatial point and the reference cell type expression profile.

[0058] In some embodiments of this application, the biological features of each spatial point in the spatial transcriptome data include the biological pathway features of each spatial point; extracting the biological features of each spatial point in the spatial transcriptome data includes: for each spatial point in the spatial transcriptome data, calculating the average value of gene expression values ​​of all biological pathways corresponding to each spatial point, and using the average value corresponding to each spatial point as the biological pathway feature of each spatial point.

[0059] As an optional implementation, the biofeatures of each spatial point in the spatial transcriptome data include the intercellular interaction features of each spatial point; extracting the biofeatures of each spatial point in the spatial transcriptome data includes: determining the receptor expression value corresponding to each spatial point, and the ligand influence value corresponding to each adjacent spatial point; determining the intercellular interaction features of each spatial point based on the receptor expression value and the ligand influence values ​​corresponding to multiple adjacent spatial points.

[0060] As an alternative implementation, the target samples that have already undergone spatial transcriptome sequencing can be used to calculate gene expression, cell type, biological pathway scores, and intercellular interaction scores. Furthermore, the Moran index is used to filter the overall scores of each spatial point, obtaining suitable spatial points to participate in subsequent model training and niche identification map generation.

[0061] Specifically, the accuracy can be obtained using the Moran index. Using the predicted result y_pred_gat from the neural network output and the corresponding spatial coordinates, with the threshold of the spatial weight matrix set to 1, the Moran index can be calculated to be 0.808, and after performing a permutation test, the P-value is 0.001.

[0062] Step S206: Construct the structure of the image to be identified based on the spatial location and biometric features of each spatial point;

[0063] In the technical solution provided in step S206, the process of constructing the graph structure to be identified includes the following steps:

[0064] The first step is to read the data and combine gene expression, cell type information, biological pathways, and intercellular interaction characteristics into a large feature matrix X.

[0065] The second step is to standardize the feature matrix.

[0066] The third step involves constructing the adjacency matrix of the graph using spatial information and the K-nearest neighbor method. Specifically, this is done by connecting each node to its eight nearest neighbors, i.e., the eight points surrounding each spot point. Here, the spot points are the spatial points mentioned above.

[0067] The fourth step is to construct a standard graph structure for identification. In this graph structure, the feature information of a point is represented by a feature matrix X, the edges represent the connection between each spot point and its eight surrounding points, and the weights are the Euclidean distances between the points.

[0068] Step S208: Process the structure of the target image through the target model to obtain the niche identification map corresponding to the target sample. The niche identification map is used to indicate the distribution of cell types, cell states and molecular characteristics in the microenvironment corresponding to the target sample.

[0069] In the technical solution provided in step S208, such as Figure 3 As shown, the target model includes an embedded neural network model, a graph neural network model, and multiple gating units. The embedded neural network model includes multiple embedded neural network layers, and the graph neural network model includes multiple graph neural network layers. Each gating unit is connected to the output of an embedded neural network layer and the output of a graph neural network layer, respectively. The output of each gating unit except the target gating unit is connected to the input of a graph neural network layer. The target gating unit is the gating unit connected to the output of the last embedded neural network layer in the embedded neural network model and the output of the last graph neural network layer in the graph neural network model.

[0070] In some embodiments of this application, an embedded neural network model can be used to pre-train nodes in the graph structure to be identified, obtaining a dimensionality-reduced representation of node features. Specifically, the embedded neural network model is implemented using an autoencoder or a self-division encoder to map high-dimensional data into a low-dimensional space, obtaining a low-dimensional representation. Then, clustering is performed on the low-dimensional features to obtain cluster labels and cluster centers. Subsequently, a niche identification map can be generated based on the cluster labels and cluster centers.

[0071] In some embodiments of this application, the input layer of the embedded neural network model can be used to receive a normalized point feature information matrix X. The embedded neural network model includes N cascaded embedded network units, wherein each unit module includes a fully connected layer, a batch normalization layer, and a LeakyReLU activation layer connected in sequence.

[0072] In some embodiments of this application, the embedded neural network model is constructed by an autoencoder or a variational autoencoder to map feature information to a low-dimensional space to obtain target latent variables.

[0073] As an optional implementation, the mean and log-variance of the latent variable z can be obtained from the output of the variational autoencoder (VAE) using two fully connected layers, and then the dependent variable z can be updated according to the formula z = μ + σ ∈ . Here, μ is the mean vector, and σ is the standard deviation vector, σ = exp(log(σ ∈ . 2 ) / 2). ∈ is a random noise vector with the same shape as σ, sampled from a standard normal distribution.

[0074] For an autoencoder (AE), the latent variable z can be obtained by sequentially connecting a fully connected layer, a batch normalization layer, and a LeakyReLU activation layer after the output of the autoencoder.

[0075] In some embodiments of this application, the embedded neural network model further includes a decoder. The input to the decoder is a latent variable z, comprising N identically structured concatenated modules, each of which includes a fully connected layer, a batch normalization layer, and a LeakyReLU activation layer connected in sequence. The output of the decoder is connected to a fully connected layer, which is used to transform the output of the last module of the decoder back to the dimension of the input data of the embedded neural network model, thereby achieving data reconstruction.

[0076] As an optional implementation, when training an embedded neural network model with a standard VAE encoder, the loss function includes reconstruction loss and KL divergence loss. The expression for the reconstruction loss is as follows:

[0077] recons_loss=MSE(x rec ,x)

[0078] In the above expression, recons_loss represents the reconstruction loss, x rec 'x' is the reconstructed data, and 'x' is the original data.

[0079] The expression for the KL divergence loss is as follows:

[0080]

[0081] In the above expression, kl_vae_loss is the KL divergence loss, and N is the number of samples in the batch. It is the square of the mean of the latent variable z of the i-th sample. Let z be the square of the variance of the latent variable z for the i-th sample.

[0082] The expression for the total loss is as follows:

[0083] loss=recons_loss+β×kl_vae_loss

[0084] Where β is a hyperparameter.

[0085] When the encoder is a Beta-VAE, the expressions for reconstruction loss and KL divergence loss are the same as those for standard VAE, and the expression for total loss is as follows:

[0086] loss=recons_loss+γ×|kl_vae_loss-c|

[0087] in γ is a hyperparameter, c max max_iter and curr_iter are parameters that control the KL divergence loss, respectively.

[0088] When the encoder is a standard AE, the total loss is the reconstruction loss described above. However, when the encoder is a contracted AE, the total loss includes the mean squared error (MSE) loss and the contraction loss. The expression for the mean squared error loss is as follows:

[0089] mse = MSE(x rec ,x)

[0090] The expression for the contraction loss is as follows:

[0091]

[0092] Where dh = h × (1 - h), h is the activation of the hidden layer. w i It is the i-th row of the weight matrix w. λ is a regularization parameter used to control the contribution of shrinkage loss to the total loss.

[0093] The expression for the total loss (total_loss) is as follows:

[0094] total_loss=mse+contractive_loss

[0095] In some embodiments of this application, when the encoder in the embedded model is a VAE, the graph neural network model obtains the reconstructed input of the embedded neural network model, the latent variable z, the mean and log-variance of the latent variable z, and the output of the encoding layer of the embedded model. When the encoder in the embedded neural network model is an AE, the graph neural network model obtains the reconstructed input of the embedded neural network model, the latent variable z, and the output of the encoding layer of the embedded neural network model.

[0096] In addition, the graph neural network model also includes a pre-module, which contains a graph attention network convolutional layer (GATv2Conv) to implement a multi-head attention mechanism for processing input features and outputting a hidden layer.

[0097] As an optional implementation method, such as Figure 3 As shown, the process of processing the target model to obtain the niche partitioning map corresponding to the target sample involves: processing the node feature information in the target graph structure through an embedded neural network model to obtain latent variables, wherein the latent variables include the intermediate latent variable Z. (i) And the target latent variable Z, which includes the reconstructed feature representation embedded in the output of the neural network model, and the intermediate latent variable Z. (i) This involves embedding intermediate processing results from the neural network model when processing feature information; processing the graph structure to be identified using a graph neural network model, and using gating units and intermediate hidden variables to control the intermediate hidden state h of the graph neural network model during the processing of the graph structure. (i) The target hidden state h is updated to obtain the output of the graph neural network model. The target hidden state is a representation of the structural features that integrate biological features and the graph structure to be identified. The niche partitioning map is determined based on the target hidden state.

[0098] In some embodiments of this application, such as Figure 3As shown, the target model includes an embedded neural network model, a graph neural network model, and multiple gating units. Each gating unit is connected to the output of an embedded neural network layer and the output of a graph neural network layer, respectively. The outputs of all gating units except the target gating unit are connected to the input of a graph neural network layer. The target gating unit is the gating unit connected to the output of the last embedded neural network layer in the embedded neural network model and the output of the last graph neural network layer in the graph neural network model. The embedded neural network model includes multiple embedded network units; the graph neural network model includes multiple graph attention layers; the gating units include update gates, reset gates, and combined gating layers. The combined gating layer includes multiple fully connected layers, used to send gating signals to the update and reset gates, and to concatenate the latent variables output by each embedded neural network layer and the hidden states output by each graph neural network layer to obtain the concatenation result. The update gate is used to determine the weights of the information provided by the hidden states retained in the concatenation result, and the weights of the information provided by the latent variables and feature vectors added, based on the gating signals. The reset gate is used to determine the contribution of the hidden states during the concatenation process based on the gating signals.

[0099] in Figure 3 In this context, z represents a latent variable, and h represents a hidden state. (i) It represents the latent variable h, which is the output of the i-th embedded neural network layer. (i) This indicates that it is a hidden variable output by the i-th graph structure network layer. Figure 3 The final layer outputs z and h represent the latent variables and hidden states of the final output. Then, based on the latent variables, the probability of belonging to a sample and its cluster center, or the soft assignment q, can be obtained. Finally, q is normalized to obtain the target distribution p.

[0100] The hidden state h of the final output is updated by using the gating unit and the hidden variables of the final output to obtain... Then based on The K-divergence loss is calculated using p, and the above is achieved through the first layer of supervision, i.e., the embedded network model loss KL(q||p), and the second layer of supervision, the graph attention loss. A two-layer self-supervised training target model is used to identify Niche maps.

[0101] in addition Figure 3 In this context, Feature matrix X represents the feature matrix X, Neighbors represents the adjacency matrix, and Adjacent matrix represents the graph structure.

[0102] In some embodiments of this application, the soft assignment (or belonging probability) q between a sample and a cluster center can be calculated based on the Deep Embedded Clustering (DEC) algorithm, with the specific calculation formula as follows. The cluster center is represented by a parameterized layer `cluster_layer`, which has a shape of (n_clusters, n_z). Here, `n_clusters` represents the total number of cluster centers, and `n_z` represents the dimension of the feature space. To initialize these cluster centers, the Xavier normal distribution initialization method can be used. This initialization strategy aims to maintain consistent variance between the input and output of each layer in the network by adjusting the scale of the weights, thereby helping to stabilize and accelerate the training process of the neural network.

[0103]

[0104] Among them, z i μ is the coordinate of data point i. j These are the coordinates of cluster center j.

[0105] As an optional implementation, q is further normalized to obtain the target distribution p. The normalization formula is as follows:

[0106]

[0107] In the above formula, i represents a data point, j represents a cluster center, N is the number of samples, and K is the total number of cluster centers.

[0108] Understandably, embedded network models can extract feature information through multi-layer nonlinear encoding, while graph neural networks can aggregate neighbor information to fully mine structural information. The embedded network model training is supervised by a target distribution P to better learn node feature representations by introducing clustering information; the target distribution P is used to improve the predicted clustering results of the graph neural network, thereby training the target model through a two-layer self-supervised module, achieving end-to-end clustering training of the target model.

[0109] As an optional implementation, the structure of the gating unit is as follows: Figure 4 As shown, it is mainly used to perform splicing operations, including updating and resetting doors and combining gate control layers, and can realize the updating of hidden states.

[0110] The update and reset gates consist of multiple modules. Each module contains a fully connected layer with input features equal to the hidden layer dimension * number of attention heads * 2 and output features equal to the hidden layer dimension * number of attention heads, along with a sigmoid activation function. Finally, it includes a fully connected layer with input features of n_z * 2 and output features of n_z, and a sigmoid activation function. The update gate is represented as z = σ(Wu [h,x]), the reset gate is represented as r=σ(W r [h,x]), where h is the hidden state at the previous time step, x is the input at the current time step, σ is the sigmoid activation function, and W... u and W r It is the weight matrix for the update gate and the reset gate.

[0111] The combined gating layer consists of multiple fully connected layers with input feature counts equal to the hidden layer dimension * number of attention heads * 2, and output feature count equal to the hidden layer dimension * number of attention heads, without using bias. Finally, a fully connected layer with input feature counts of n_z * 2 and output feature counts of n_z is added. The calculation formula is as follows: Among them, W c It is the weight matrix of the combined gating layer.

[0112] When updating the hidden state, the new hidden state h ′ Through formula Update. Afterwards, the updated hidden state and adj (the graph structure to be identified) can be processed using the current graph network layer. The current graph network layer refers to the next graph network layer after the graph network layer that outputs the hidden state before the update.

[0113] In some embodiments of this application, the loss of the target model includes embedding loss, clustering embedding loss, and clustering graph attention loss, wherein the clustering embedding loss is used to reflect the difference between the first predicted clustering distribution result of the embedded neural network model and the target clustering distribution result, and the clustering graph attention loss is used to reflect the difference between the second predicted clustering distribution result of the graph neural network model and the target clustering distribution result.

[0114] As an optional implementation, the embedding loss here is the total loss of the embedded neural network model, denoted as emb_loss. total The clustering embedding loss, cluster_emb, reflects the loss between the encoder's partially predicted distribution q and the target distribution p, and its expression is as follows:

[0115]

[0116] In the above formula, i represents a data point, j represents a cluster center, N is the total number of data points, and K is the total number of cluster centers. ij p represents the predicted probability that the i-th data point belongs to the j-th cluster center. ij This represents the target probability that the i-th data point belongs to the j-th cluster center.

[0117] The clustering graph attention loss, cluster_gat, is the node representation predicted by the graph attention network. The KL divergence loss between the target distribution p and the target distribution p is calculated in the same way as the clustering embedding loss. The total loss of the target model is expressed as follows:

[0118] loss = α1 × emb_loss total +α2×cluster_emb+α3×cluster_gat

[0119] Where α1, α2, and α3 are preset weighting coefficients.

[0120] In some embodiments of this application, spatial transcriptome slice samples LC15 and LC04 are used as examples to further illustrate the complete execution flow and technical effects of the niche identification method provided in the embodiments of this application. LC15 and LC04 are both hepatobiliary carcinoma samples.

[0121] First, the spatial transcriptome data of all samples were processed to obtain gene expression, cell type information, biological pathways, and intercellular interaction characteristics for each spot (spatial point).

[0122] Specifically, gene expression data from tissue sections can be collected using spatial transcriptomics platforms. This data represents the gene expression profiles of cell populations at various spatial locations within the tissue. For cell type information, reference cell type expression profile libraries obtained beforehand through single-cell RNA sequencing (scRNA-seq) can be used as prior information. These reference libraries contain gene expression patterns of a range of known cell types, providing the necessary gene expression context for spatial cell localization.

[0123] In Cell2location, spatial expression data and reference cell type expression profiles are integrated, and a Bayesian model is used to infer the spatial distribution of cell types. This model estimates the relative abundance of a specific cell type at each spatial location. The Bayesian framework allows the model to quantify the uncertainty of cell type distribution while considering the inherent noise and technical variability of the data. Model parameter estimation is achieved through variational inference, an efficient approximate Bayesian inference technique that approximates the posterior distribution by optimizing the variational lower bound (ELBO). This step involves optimization algorithms, such as Automatic Differential Variational Inference (ADVI), to obtain point estimates and uncertainty estimates of the model parameters. Finally, Cell2location outputs the estimated abundance of each cell type at each spatial location, representing the spatial cellular composition of the tissue.

[0124] When determining biological pathway information, for each point in space, the average expression value of all genes belonging to that pathway at that point is calculated, as shown in the following formula:

[0125]

[0126] Among them, E pathway (p) represents the average expression level of the pathway at point p. is the expression value of the k-th gene at that point, and m is the total number of genes in the pathway.

[0127] To quantify the impact of ligand-receptor interactions on specific regions in space for intercellular interaction information, ligand-receptor pair information was obtained from the CellPhoneDB data file "connectomeDB2020," and the following computational model was employed. In this model, the influence value of each target point is denoted as 'effect'. LR It is determined by the influence value of ligands within its 10 adjacent spatial points. Perform weighted aggregation and compare with the corresponding receptor expression value R j It is determined by multiplying the products, summing them, and then taking the square root of the sum. That is:

[0128]

[0129] here, Indicates a single ligand L i For receptor R j The strength of the effect is determined by the expression value L of the ligand. i Its corresponding receptor R i weight w i Multiplying them together gives:

[0130]

[0131] Weight w(L) i ,R j ) is based on ligand L i With receptor R j Spatial distance between dist(L) i ,R j The value is calculated using Gaussian function decay, and its form is:

[0132]

[0133] Among them, dist(L i ,R j ) is the ligand L i With receptor R j The Euclidean distance between the ligand and receptor is given by d(u,v), where l is the attenuation coefficient controlling the rate at which the distance's influence on the weighting decreases. The formula for calculating the distance d(u,v) between the ligand and receptor is:

[0134]

[0135] Then, based on the gene expression, cell type information, biological pathways, and intercellular interaction characteristics of each spot (spatial point), a graph structure to be processed can be constructed. The node features of this graph structure are then converted into matrix form for pre-training. The training parameters for the model during pre-training are set as follows:

[0136] Number of input features: in_features = 328

[0137] Number of hidden layers: n_hiddens = 1

[0138] Number of attention heads: n_head = 2

[0139] Number of clusters: n_cluster = number of cell types + 2

[0140] Hidden layer dimensions: dim_hidden = [128, 64, 64, 128]

[0141] Dimensions of the latent space: n_z = 64

[0142] Learning rate: learning_rate = 0.01

[0143] Maximum number of training epochs: max_epochs = 50

[0144] L2 regularization weight: l2_weight = 0.00001

[0145] Batch_size = 512

[0146] Model weight settings: alpha = 1, beta = 0.001, gamma = 0.0001, c_max = 25, lam = 1e-5

[0147] Data dimensionality reduction method: embedding='ae'

[0148] Embedded model type: model_type = 'contractive_ae'

[0149] As an optional implementation, based on the above parameters, the operations of each layer in the embedded neural network model, the size of the output feature map, and the network connectivity are shown in the table below:

[0150]

[0151]

[0152] All data in the training samples are used for training and prediction, and an unsupervised learning method is employed.

[0153] As an optional implementation, the number of training iterations for the deep neural network can be set to 50, the sample batch size to 64, and the learning rate to 0.01.

[0154] The initial latent variable z can be obtained by pre-training the target model. Then, the pre-trained embedded neural network model is saved, and cluster labels and cluster centers are obtained by clustering z.

[0155] As an optional implementation, after pre-training, the following parameters can be added to the above training parameters to continue training the graph neural network model and obtain cluster labels:

[0156] The initial cluster labels and cluster centers are defined as those obtained by clustering the latent variable z.

[0157] Model weight settings: alpha_emb_recons = 0.01, alpha_emb_cluster = 0.001, alpha_gat_cluster = 15.0, lam = 1e-5, tol = 1e-6.

[0158] Obtain the reconstructed x, z, and encoder output from the pre-trained embedded neural network model.

[0159] The input and output of the graph neural network model are shown in the table below:

[0160]

[0161]

[0162] The final output result is as follows Figure 5-7 As shown. Among them. Figure 5 This is a heatmap of the number of genes in the spatial transcriptome. Figure 6 A cell type annotation diagram based on spatial transcriptomics. Figure 7 This is a niche identification map of the tumor microenvironment, with the Moran's index as the evaluation metric being 0.808. (Comparison) Figure 6 and Figure 7 As can be seen, the method provided in the embodiments of this application can accurately identify the ecological niche of the tumor microenvironment.

[0163] The number of cells of each cell type in each ecological niche, such as Figure 8 As shown. The spatial distribution of important ecological characteristics, such as gene expression, is as follows. Figure 9 As shown. However, when using only gene expression as a feature, with other parameters unchanged, the niche identification map is as follows: Figure 10 As shown, the Moran index at this time is 0.747. (Comparison) Figure 10 and Figure 7 It can be seen that using only gene expression as a feature affects the accuracy of the final niche identification map. However, without using a gating mechanism, when the learning rate is modified to 0.001, the niche identification map is as follows: Figure 11 As shown, although the Moran index is 0.827 (slightly higher than the Moran index of 0.808 when using the gating mechanism), the number of clusters is significantly reduced, indicating that the model's portrayal of the complex tumor niche is worse without the gating mechanism than with the gating mechanism.

[0164] In summary, the niche identification method provided in this application has at least the following beneficial effects:

[0165] First, by comprehensively utilizing various biological data, including spatial location, gene expression, cell type, cell interactions, and pathway information, a more comprehensive understanding and study of the tumor microenvironment has been achieved. By fully leveraging the biological information of samples (including spatial location, gene expression, cell type information, biological pathways, and intercellular interactions), the accuracy of cancer niche identification has been significantly improved, which has important guiding significance for future early cancer screening.

[0166] Secondly, this application provides a method for representing and analyzing the aforementioned types of information using a target model that incorporates both an embedded neural network model and a graph neural network model. The embedded neural network model can be an autoencoder network or a variational autoencoder network. Furthermore, it provides a method for learning the relationships between spatial points using a graph attention network and incorporating a gating mechanism during training. In other words, the method provided in this application automatically extracts features from various biological information using an embedded model, avoiding the drawbacks of traditional manual feature extraction. The extracted features are then combined with the spatial relationships between points learned by the graph attention network through a gating mechanism, thereby fully exploiting the rich feature data information contained in each point of the sample space and ensuring the high reliability and effectiveness of the recognition results.

[0167] By acquiring spatial transcriptome data of the target sample, including multiple spatial points and the spatial location of each point, and extracting the biological characteristics of each spatial point, a target identification map structure is constructed based on the spatial location and biological characteristics of each point. The target model is then used to process the target identification map structure to obtain the niche identification map corresponding to the target sample. This niche identification map indicates the distribution of cell types, cell states, and molecular characteristics in the microenvironment corresponding to the target sample. By constructing the target identification map structure based on the spatial location information and biological characteristics in the spatial transcriptome data, the goal of comprehensively analyzing the spatial and biological characteristics of the sample is achieved. This results in obtaining a high-precision niche identification map and solves the technical problem of incomplete information on tissue heterogeneity and other characteristics of the target sample due to the inability to comprehensively analyze various feature data in the sample data in related technologies.

[0168] This application provides a niche identification device. Figure 12 This is a schematic diagram of the device. From Figure 12 As can be seen from the diagram, the device includes: a first processing module 120, used to acquire spatial transcriptome data of the target sample, the spatial transcriptome data including multiple spatial points and the spatial location of each spatial point; a second processing module 122, used to extract the biological features of each spatial point in the spatial transcriptome data; a third processing module 124, used to construct a map structure to be identified based on the spatial location and biological features of each spatial point; and a fourth processing module 126, used to process the map structure to be identified through the target model to obtain the niche identification map corresponding to the target sample, wherein the niche identification map is used to indicate the distribution of cell types, cell states and their molecular features in the microenvironment corresponding to the target sample.

[0169] In some embodiments of this application, the biological characteristics of each spatial point in the spatial transcriptome data include the cell type characteristics of each spatial point; the step of the second processing module 122 extracting the biological characteristics of each spatial point in the spatial transcriptome data includes: collecting expression data of the target sample through a spatial transcriptomics platform, and determining gene expression characteristics based on the expression data, wherein the expression data is used to characterize the gene expression profile of the cell population at each spatial point in the tissue of the target sample; performing single-cell RNA sequencing on the target sample to obtain the reference cell type expression profile of the target sample, and determining the biological characteristics of each spatial point based on the gene expression profile of the cell population at each spatial point and the reference cell type expression profile.

[0170] In some embodiments of this application, the biological features of each spatial point in the spatial transcriptome data include the biological pathway features of each spatial point; the step of the second processing module 122 extracting the biological features of each spatial point in the spatial transcriptome data includes: for each spatial point in the spatial transcriptome data, calculating the average value of gene expression values ​​of all biological pathways corresponding to each spatial point, and using the average value corresponding to each spatial point as the biological pathway feature of each spatial point.

[0171] In some embodiments of this application, the biofeatures of each spatial point in the spatial transcriptome data include the intercellular interaction features of each spatial point; the step of the second processing module 122 extracting the biofeatures of each spatial point in the spatial transcriptome data includes: determining the receptor expression value corresponding to each spatial point, and the ligand influence value corresponding to each of the multiple adjacent spatial points; and determining the intercellular interaction features of each spatial point based on the receptor expression value and the ligand influence value corresponding to each of the multiple adjacent spatial points.

[0172] In some embodiments of this application, the target model includes an embedded neural network model, a graph neural network model, and multiple gating units. The embedded neural network model includes multiple embedded neural network layers, and the graph neural network model includes multiple graph neural network layers. Each gating unit is connected to the output of an embedded neural network layer and the output of a graph neural network layer, respectively. The output of each gating unit other than the target gating unit is connected to the input of a graph neural network layer. The target gating unit is a gating unit connected to the output of the last embedded neural network layer in the embedded neural network model and the output of the last graph neural network layer in the graph neural network model.

[0173] In some embodiments of this application, the fourth processing module 126 processes the graph structure to be identified through a target model to obtain a niche partitioning map corresponding to the target sample, including: processing the node feature information in the graph structure to be identified through an embedded neural network model to obtain latent variables, wherein the latent variables include intermediate latent variables and target latent variables, the target latent variables include a low-dimensional feature representation of the feature information output by the embedded neural network model, and the intermediate latent variables are intermediate processing results of the embedded neural network model when processing feature information; processing the graph structure to be identified through a graph neural network model, and updating the intermediate hidden state of the graph neural network model through a gating unit and intermediate latent variables during the processing of the graph structure to be identified to obtain the target hidden state output by the graph neural network model, wherein the target hidden state is a representation that integrates biological features and structural features of the graph structure to be identified; and determining the niche partitioning map based on the target hidden state.

[0174] In some embodiments of this application, the embedded neural network model is constructed by an autoencoder or a variational autoencoder to map feature information to a low-dimensional space to obtain target latent variables.

[0175] In some embodiments of this application, the embedded neural network model includes multiple embedded neural network layers, the graph neural network model includes multiple graph structure network layers, and the gating unit includes an update gate, a reset gate, and a combined gating layer. The combined gating layer includes multiple fully connected layers, used to send gating signals to the update gate and the reset gate, and to concatenate the latent variables output by each embedded neural network layer and the hidden states output by each graph neural network layer to obtain a concatenated result. The update gate is used to determine, based on the gating signals, the weights of the information provided by the hidden states retained in the concatenated result, and the weights of the information provided by the latent variables and feature vectors added. The reset gate is used to determine, based on the gating signals, the contribution level of the hidden states during the concatenation process.

[0176] In some embodiments of this application, the loss of the target model includes embedding loss, clustering embedding loss, and clustering graph attention loss, wherein the clustering embedding loss is used to reflect the difference between the first predicted clustering distribution result of the embedded neural network model and the target clustering distribution result, and the clustering graph attention loss is used to reflect the difference between the second predicted clustering distribution result of the graph neural network model and the target clustering distribution result.

[0177] It should be noted that each module in the above-mentioned niche identification device can be a program module (e.g., a set of program instructions to implement a certain function) or a hardware module. For the latter, it can be manifested in the following forms, but is not limited to them: each of the above modules is manifested as a processor, or the functions of each of the above modules are implemented by a processor.

[0178] According to an embodiment of this application, a non-volatile storage medium is also provided, which stores a program. During program execution, the program controls the device containing the non-volatile storage medium to perform the following niche identification method: acquiring spatial transcriptome data of a target sample, the spatial transcriptome data including multiple spatial points and the spatial location of each spatial point; extracting biological features from each spatial point in the spatial transcriptome data; constructing a target identification map structure based on the spatial location and biological features of each spatial point; and processing the target identification map structure using a target model to obtain a niche identification map corresponding to the target sample. The niche identification map is used to indicate the distribution of cell types, cell states, and molecular characteristics in the microenvironment corresponding to the target sample.

[0179] According to an embodiment of this application, an electronic device is also provided, including: a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the following niche identification method: acquiring spatial transcriptome data of a target sample, the spatial transcriptome data including multiple spatial points and the spatial location of each spatial point; extracting biological features of each spatial point in the spatial transcriptome data; constructing a map structure to be identified based on the spatial location and biological features of each spatial point; processing the map structure to be identified through a target model to obtain a niche identification map corresponding to the target sample, wherein the niche identification map is used to indicate the distribution of cell types, cell states and their molecular features in the microenvironment corresponding to the target sample.

[0180] According to an embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the following niche identification method: acquiring spatial transcriptome data of a target sample, the spatial transcriptome data including multiple spatial points and the spatial location of each spatial point; extracting biological features of each spatial point in the spatial transcriptome data; constructing a map structure to be identified based on the spatial location and biological features of each spatial point; and processing the map structure to be identified through a target model to obtain a niche identification map corresponding to the target sample, wherein the niche identification map is used to indicate the distribution of cell types, cell states, and molecular features in the microenvironment corresponding to the target sample.

[0181] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0182] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0183] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0184] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0185] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0186] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A niche recognition method, characterized by, include: Acquire spatial transcriptome data of the target sample, wherein the spatial transcriptome data includes multiple spatial points and the spatial location of each spatial point; Extract the biological characteristics of each spatial point in the spatial transcriptome data; A graph structure to be identified is constructed based on the spatial location of each spatial point and the biometric features. The target model is used to process the structure of the target image to obtain the niche identification map corresponding to the target sample. The niche identification map is used to indicate the distribution of cell types, cell states and molecular characteristics in the microenvironment corresponding to the target sample.

2. The niche recognition method of claim 1, wherein, The target model includes an embedded neural network model, a graph neural network model, and multiple gating units, wherein... The embedded neural network model includes multiple embedded neural network layers, and the graph neural network model includes multiple graph neural network layers; Each of the gated units is connected to the output of an embedded neural network layer and the output of a graph neural network layer, respectively. The output of each gated unit, except for the target gated unit, is connected to the input of a graph neural network layer. The target gated unit is a gated unit connected to the output of the last embedded neural network layer in the embedded neural network model and the output of the last graph neural network layer in the graph neural network model.

3. The niche identification method of claim 2, wherein, The target model is used to process the structure of the target image to obtain the niche map corresponding to the target sample, including: The embedded neural network model processes the node feature information in the graph structure to be identified to obtain latent variables. The latent variables include intermediate latent variables and target latent variables. The target latent variables include the low-dimensional feature representation of the feature information output by the embedded neural network model. The intermediate latent variables are the intermediate processing results of the embedded neural network model when processing the feature information. The graph neural network model processes the graph structure to be identified, and updates the intermediate hidden state of the graph neural network model through the gating unit and the intermediate hidden variable during the processing of the graph structure to be identified, so as to obtain the target hidden state output by the graph neural network model, wherein the target hidden state is a representation that integrates the biometric features and the structural features of the graph structure to be identified; The niche map is determined based on the target's hiding state.

4. The niche identification method of claim 3, wherein, The embedded neural network model is constructed using an autoencoder or variational autoencoder to map the feature information to a low-dimensional space to obtain the target latent variable.

5. The niche identification method of claim 2, wherein, The embedded neural network model includes multiple embedded neural network layers, the graph neural network model includes multiple graph structure network layers, and the gating unit includes update gates, reset gates, and combined gating layers. The combined gating layer includes multiple fully connected layers, used to send gating signals to the update gate and the reset gate, and to concatenate the latent variables output by each of the embedded neural network layers and the hidden states output by each of the graph neural network layers to obtain a concatenation result; The update gate is used to determine the weight of the information provided by the hidden state retained in the splicing result, and the weight of the information provided by the latent variables and feature vectors added, based on the gating signal; the reset gate is used to determine the contribution of the hidden state during the splicing process based on the gating signal.

6. The niche identification method of claim 2, wherein, The loss of the target model includes embedding loss, clustering embedding loss, and clustering graph attention loss. The clustering embedding loss is used to reflect the difference between the first predicted clustering distribution result of the embedding neural network model and the target clustering distribution result. The clustering graph attention loss is used to reflect the difference between the second predicted clustering distribution of the graph neural network model and the target clustering distribution result.

7. The niche identification method of claim 1, wherein, The biometrics of each spatial point in the spatial transcriptome data include the cell type characteristics of each spatial point. Extracting biomarkers from each spatial point in the spatial transcriptome data includes: Expression data of the target sample are collected through a spatial transcriptomics platform, and gene expression characteristics are determined based on the expression data, wherein the expression data is used to characterize the gene expression profile of the cell population at each spatial point in the tissue of the target sample. Single-cell RNA sequencing is performed on the target sample to obtain the reference cell type expression profile of the target sample, and the biological characteristics of each spatial point are determined based on the gene expression profile of the cell population at each spatial point and the reference cell type expression profile.

8. The niche identification method of claim 1, wherein, The biological characteristics of each spatial point in the spatial transcriptome data include the biological pathway characteristics of each spatial point. Extracting biomarkers from each spatial point in the spatial transcriptome data includes: For each spatial point in the spatial transcriptome data, the average gene expression value of all biological pathways corresponding to each spatial point is calculated, and the average value corresponding to each spatial point is used as the biological pathway feature of each spatial point.

9. The niche identification method of claim 1, wherein, The biosignatures of each spatial point in the spatial transcriptome data include the intercellular interaction characteristics of each spatial point; extracting the biosignatures of each spatial point in the spatial transcriptome data includes: Determine the receptor expression value corresponding to each spatial point, and the ligand influence value corresponding to each of the multiple adjacent spatial points. Based on the receptor expression value, the intercellular interaction characteristics of each spatial point are determined by the ligand influence value corresponding to each of the plurality of adjacent spatial points.

10. A niche recognition apparatus, characterized by, include: The first processing module is used to acquire spatial transcriptome data of the target sample, wherein the spatial transcriptome data includes multiple spatial points and the spatial location of each spatial point. The second processing module is used to extract the biological characteristics of each spatial point in the spatial transcriptome data; The third processing module is used to construct the image structure to be identified based on the spatial location of each spatial point and the biometric features. The fourth processing module is used to process the target model to obtain the niche identification map corresponding to the target sample, wherein the niche identification map is used to indicate the distribution of cell types, cell states and molecular characteristics in the microenvironment corresponding to the target sample.

11. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a program, wherein when the program is executed, it controls the device containing the non-volatile storage medium to execute the niche identification method according to any one of claims 1 to 9.

12. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the niche identification method according to any one of claims 1 to 9.

13. A computer program product, characterised in that, It includes a computer program that, when executed by a processor, implements the niche identification method according to any one of claims 1 to 9.