Biological analysis methods, devices, equipment and media based on spatial multi-omics data

By acquiring and integrating spatial multi-omics data of target slices, constructing an adjacency graph, and using a graph attention network, the analytical stability problem of large-scale spatial multi-omics data was solved, achieving better biological analysis results.

CN120432015BActive Publication Date: 2025-09-26BGI RESEARCH SANYA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510927935.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-26
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively integrating different spatial omics data, resulting in model instability and difficulty in processing large-scale spatial multi-omics data, which affects the results of biological analysis.

Method used

By acquiring spatial multi-omics data of target slices, performing feature encoding, constructing a spatial multi-omics adjacency graph, and using a graph attention network for feature integration, spatial neighborhood relationships are dynamically modeled to capture key spatial-specific signals.

Benefits of technology

It achieves flexible integration of large-scale spatial multi-omics data, improves the accuracy and efficiency of biological analysis, and enables a better understanding of cell-to-cell interactions and spatial distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120432015B_ABST
    Figure CN120432015B_ABST
Patent Text Reader

Abstract

The present disclosure provides a biological analysis method and apparatus, equipment and medium based on spatial multi-omics data. The method includes: obtaining spatial multi-omics data of a target slice, the spatial multi-omics data including at least two spatial omics data, and each spatial omics data is used to describe the biological expression information of multiple spatial sites in the target slice; encoding the at least two spatial omics data respectively to obtain at least two spatial omics feature vectors for each spatial site; constructing a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data based on the positional relationship between each spatial site in the target slice; integrating at least two spatial omics feature vectors according to the spatial multi-omics adjacency graph and the graph attention network to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice, and performing biological analysis based on the spatial multi-omics integrated features. The embodiment of the present application can provide a biological analysis method suitable for large-scale spatial multi-omics data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of slice data integration, and in particular to a biological analysis method, apparatus, device, and medium based on spatial multi-omics data. Background Art

[0002] Currently, different spatial omics technologies measure different omics levels (e.g., genome, transcriptome, proteome, and metabolome), and single spatial omics information cannot fully reveal the complex interactions between cells and the relationship between cells and the tissue environment. Therefore, spatial multi-omics technologies have been proposed, specifically including spatial transcriptomics, spatial proteomics, and spatial metabolomics. These spatial multi-omics technologies can combine spatial information from multiple spatial omics levels with tissues to provide a more comprehensive understanding of the interactions and spatial distribution of cells within an organism.

[0003] Although related technologies have proposed methods for biological analysis of spatial multi-omics data by combining models, the distribution characteristics of different spatial multi-omics data can easily lead to model instability, making it difficult for these technologies to process large-scale spatial multi-omics data, thereby affecting their analytical results. Therefore, the development of a biological analysis method that can be applied to large-scale spatial multi-omics data has become a pressing technical problem. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a biological analysis method and apparatus, equipment and medium based on spatial multi-omics data, aiming to provide a biological analysis method that can be applied to large-scale spatial multi-omics data.

[0005] To achieve the above objectives, a first aspect of an embodiment of the present application proposes a biological analysis method based on spatial multi-omics data, the method comprising:

[0006] Acquiring spatial multi-omics data of a target slice, wherein the spatial multi-omics data includes at least two types of spatial omics data, and each type of spatial omics data is used to describe biological expression information corresponding to multiple spatial sites in the target slice;

[0007] Performing feature encoding on at least two of the spatial omics data respectively to obtain at least two spatial omics feature vectors corresponding to each of the spatial sites;

[0008] constructing a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data based on the positional relationship between the spatial sites in the target slice;

[0009] Performing feature integration on the at least two spatial omics feature vectors according to the spatial multi-omics adjacency graph and the graph attention network to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice;

[0010] Biological analysis is performed based on the spatial multi-omics integration characteristics to obtain analysis results.

[0011] In some embodiments, the step of integrating the at least two spatial omics feature vectors according to the spatial multi-omics adjacency graph and the graph attention network to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice includes:

[0012] For each of the spatial sites, obtaining a plurality of corresponding adjacent sites according to the spatial multi-omics adjacency graph; the adjacent sites are other spatial sites in the spatial multi-omics adjacency graph that have an adjacent relationship with the spatial site;

[0013] According to the correlation between the spatial site and each of the adjacent sites, feature fusion is performed on the at least two spatial omics feature vectors of each of the spatial sites to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice.

[0014] In some embodiments, the feature fusion of the at least two spatial omics feature vectors corresponding to each of the spatial sites according to the correlation between the spatial site and each of the adjacent sites to obtain the spatial multi-omics integrated feature corresponding to each spatial site in the target slice includes:

[0015] Performing feature correlation calculation on the at least two spatial omics feature vectors of the spatial site and the at least two spatial omics feature vectors of each of the adjacent sites to obtain site correlation data between the spatial site and each of the adjacent sites;

[0016] performing normalization processing on the site correlation data;

[0017] Performing feature weighting and calculation on the normalized site correlation data and the at least two spatial omics feature vectors corresponding to each of the adjacent sites to obtain a spatial multi-omics weighted feature corresponding to each of the spatial sites;

[0018] The spatial multi-omics weighted features corresponding to each of the spatial sites are activated to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice.

[0019] In some embodiments, the step of performing feature encoding on at least two spatial omics data to obtain at least two spatial omics feature vectors corresponding to each spatial site includes:

[0020] Performing feature encoding on at least two of the spatial omics data according to a preset encoder to obtain at least two spatial omics feature vectors corresponding to each of the spatial sites, wherein the preset encoder includes a feature extraction layer and a feature mapping layer;

[0021] The feature extraction layer is used to extract features from each type of spatial omics data to obtain initial omics features corresponding to each type of spatial omics data;

[0022] The feature mapping layer is used to perform feature mapping on each of the initial omics features to obtain at least two spatial omics feature vectors corresponding to each of the spatial sites.

[0023] In some embodiments, the training process of the preset encoder and the graph attention network includes the following steps:

[0024] Acquiring sample space multi-omics data of the sample slice, wherein the sample space multi-omics data includes at least two types of sample space omics data, and each type of sample space omics data is used to describe biological expression information corresponding to multiple sample space sites in the sample slice;

[0025] Performing feature encoding on at least two sample space omics data respectively to obtain at least two sample space omics feature vectors corresponding to each of the sample space sites;

[0026] constructing a sample space multi-omics adjacency graph corresponding to the sample space multi-omics data based on the positional relationship between each sample space site in the sample slice;

[0027] Performing feature integration on the at least two sample space omics feature vectors according to the sample space multi-omics adjacency graph and the graph attention network to obtain a sample space multi-omics integrated feature corresponding to each sample space site in the sample slice;

[0028] constructing a data reconstruction loss according to the sample space multi-omics data and a plurality of the sample space multi-omics integrated features;

[0029] constructing a neighborhood preservation loss based on the sample space multi-omics adjacency graph and the at least two sample space omics feature vectors;

[0030] A target loss is constructed according to the data reconstruction loss and the neighborhood preservation loss, and model parameters of the preset encoder and the graph attention network are updated based on the target loss.

[0031] In some embodiments, constructing a data reconstruction loss based on the sample space multi-omics data and a plurality of the sample space multi-omics integrated features comprises:

[0032] Performing feature decoding on a plurality of the multi-omics integration features of the sample spaces to obtain sample reconstruction data corresponding to each of the omics data of the sample spaces;

[0033] Performing data error calculation on each of the sample space omics data and the corresponding sample reconstruction data to obtain a data reconstruction error corresponding to each of the sample space omics data;

[0034] The data reconstruction loss is constructed according to a plurality of the data reconstruction errors.

[0035] In some embodiments, constructing a neighborhood preserving loss based on the sample space multi-omics adjacency graph and the at least two sample space omics feature vectors includes:

[0036] Converting the sample space multi-omics adjacency graph into a corresponding adjacency matrix;

[0037] For each of the sample space sites, a corresponding plurality of sample adjacent sites are obtained according to the sample space multi-omics adjacency graph; the sample adjacent sites are other sample space sites in the sample space multi-omics adjacency graph that have an adjacency relationship with the sample space site;

[0038] Performing feature difference calculation on the at least two sample spatial omics feature vectors of the sample spatial site and the at least two sample spatial omics feature vectors of each of the sample adjacent sites to obtain multi-omics difference features between the sample spatial site and each of the sample adjacent sites;

[0039] The neighborhood preservation loss is constructed according to the adjacency matrix and the multi-omics difference features corresponding to the plurality of sample spatial sites.

[0040] To achieve the above objectives, a second aspect of the embodiments of the present application provides a biological analysis device based on spatial multi-omics data, the device comprising:

[0041] an acquisition module, configured to acquire spatial multi-omics data of a target slice, wherein the spatial multi-omics data includes at least two types of spatial omics data, and each type of spatial omics data is used to describe biological expression information corresponding to multiple spatial sites in the target slice;

[0042] an encoding module, configured to perform feature encoding on at least two of the spatial omics data respectively to obtain at least two spatial omics feature vectors corresponding to each of the spatial sites;

[0043] A graph construction module, configured to construct a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data based on the positional relationship between the spatial sites in the target slice;

[0044] an integration module, configured to perform feature integration on the at least two spatial omics feature vectors according to the spatial multi-omics adjacency graph and the graph attention network to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice;

[0045] The analysis module is used to perform biological analysis based on the spatial multi-omics integration characteristics to obtain analysis results.

[0046] To achieve the above-mentioned purpose, a third aspect of an embodiment of the present application proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0047] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0048] To achieve the above-mentioned purpose, the fifth aspect of the embodiments of the present application proposes a computer program product, which includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the method described in the first aspect.

[0049] The embodiments of the present application propose a biological analysis method, apparatus, device and medium based on spatial multi-omics data, which obtains spatial multi-omics data of a target slice, wherein the spatial multi-omics data includes at least two types of spatial omics data, and each type of spatial omics data is used to describe the biological expression information corresponding to multiple spatial sites in the target slice; further, feature encoding is performed on the at least two types of spatial omics data respectively to obtain at least two spatial omics feature vectors corresponding to each spatial site; a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data is constructed based on the positional relationship between each spatial site in the target slice; feature integration is performed on the at least two spatial omics feature vectors according to the spatial multi-omics adjacency graph and the graph attention network to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice; biological analysis is performed based on the spatial multi-omics integrated features to obtain analysis results.

[0050] When performing biological analysis based on spatial multi-omics data, the embodiments of the present application can not only construct a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data according to each spatial site in the target slice, but also perform feature integration of multiple spatial omics feature vectors based on the spatial multi-omics adjacency graph and the graph attention network. In this way, the problem of inaccurate data integration due to the distribution characteristics of different spatial omics data can be avoided. That is, by introducing the spatial multi-omics adjacency graph and the graph attention network for data integration, the spatial neighborhood relationship between each spatial site in the target slice can be dynamically modeled, and the key spatial specific signals of each spatial site in the target slice can be captured more flexibly, thereby improving the integration capability of spatial multi-omics data, so that it can be better applied to the biological analysis method of large-scale spatial multi-omics data. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings are used to provide a further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.

[0052] Figure 1 This is a flow chart of a biological analysis method based on spatial multi-omics data provided in an embodiment of the present application;

[0053] Figure 2 yes Figure 1 A flowchart of step S104 in FIG.

[0054] Figure 3 yes Figure 1 A flowchart of step S202 in FIG.

[0055] Figure 4 This is a flowchart of training a preset encoder and a graph attention network provided by an embodiment of the present application;

[0056] Figure 5 This is a flowchart of a specific embodiment of a biological analysis method based on spatial multi-omics data provided in an embodiment of the present application;

[0057] Figure 6A This is a schematic diagram of dividing thymus tissue into different spatial regions provided in an embodiment of the present application;

[0058] Figure 6B It is a schematic diagram of the UMAP map of the embedded space after thymus tissue processing based on the biological analysis method of spatial multi-omics data;

[0059] Figure 6C It is a schematic diagram of the spatial regions after thymus tissue processing based on the biological analysis method of spatial multi-omics data;

[0060] Figure 7 This is a cluster diagram obtained by spatial clustering of thymus tissue data based on MOFA+ technology;

[0061] Figure 8 Schematic diagram of a biological analysis device based on spatial multi-omics data provided in an embodiment of the present application;

[0062] Figure 9 This is a hardware structure diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not intended to limit the present disclosure.

[0064] Before further explaining the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:

[0065] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0066] Spatial transcriptomics (ST) is an emerging technology that can analyze RNA-seq data at the tissue level to obtain transcriptional information at specific locations within a tissue. This technology allows researchers to analyze gene expression patterns in tissue sections while maintaining information about the spatial location of cells. Spatial transcriptomics can reveal cellular heterogeneity and define cell types while preserving their spatial information, which is crucial for understanding fields such as cell biology, developmental biology, neurobiology, and tumor biology.

[0067] Spatial proteomics (SP) is a technique that studies the spatial distribution, localization, and interactions of proteins in cells and tissues. This technique combines laser microdissection and high-resolution imaging to obtain information on the spatial distribution of protein molecules within cells or tissues. Spatial proteomics can analyze protein expression profiles in different cells and functional regions within a tissue, providing insights into the spatial microenvironment of tissues and cellular functions. Spatial proteomics is crucial for revealing the complex structure of the human proteome, including single-cell variation, dynamic protein translocation, altered interaction networks, and protein localization in multiple compartments.

[0068] Heterogeneous data refers to data from different technologies that differ in properties, such as data structure, measurement accuracy, and data magnitude. For example, different spatial omics technologies may generate different types of data with varying quality and characteristics, such as resolution, coverage, and signal strength. Effectively integrating this data requires finding a method that allows these data from different sources and properties to be analyzed together, enabling a more comprehensive understanding of the complexity of biological samples.

[0069] Currently, different spatial omics technologies measure different omics levels (e.g., genome, transcriptome, proteome, and metabolome). These data sources and resolutions vary, and a single spatial omics dataset cannot fully reveal the complex interactions between cells and the cell-tissue environment. Therefore, effectively integrating these heterogeneous data sets is a key challenge in further exploring the potential of spatial multi-omics data. Spatial multi-omics techniques have been proposed, specifically including spatial transcriptomics, spatial proteomics, and spatial metabolomics. These spatial multi-omics techniques combine spatial information from multiple spatial omics levels with tissues to provide a more comprehensive understanding of cellular interactions and spatial distribution within organisms. Spatial multi-omics data are often high-dimensional and noisy, and differences in data scale and distribution between different omics also complicate integration. Designing robust data integration algorithms can not only reduce data redundancy but also uncover potential biological signals, such as uncovering new cellular states and identifying key molecular pathways or markers. Therefore, integrating multi-level spatial omics data can help establish systematic cellular network models and dissect the links between biological function and spatial organization.

[0070] In related technologies, a variety of algorithms for spatial multi-omics data integration have been proposed. These methods can be summarized into the following categories: (1) Methods based on dimensionality reduction and feature extraction: For example, the key features of each omics data are extracted through methods such as principal component analysis (PCA) and non-negative matrix factorization (NMF), and then data integration is achieved by aligning or embedding them into a common low-dimensional space. (2) Methods based on deep learning: Using deep learning frameworks such as graph neural networks (GNNs) or variational autoencoders (VAEs), multi-omics data are integrated from a nonlinear perspective. (3) Probabilistic models and Bayesian frameworks: Methods based on statistical models infer the potential connections between omics. Commonly used methods include Bayesian networks and Markov chain Monte Carlo methods to provide uncertainty assessment of data integration results.

[0071] However, methods based on dimensionality reduction and feature extraction typically assume a linear data distribution, making it difficult to capture the complex nonlinear relationships in spatial multi-omics data, especially when processing signals across scales and across diverse biological pathways. This is because these methods primarily focus on the omics data itself, ignoring spatial location information, making it difficult to accurately preserve the spatial structure of tissues. The computational complexity of the graphical models employed in related technologies increases dramatically, limiting their application to large-scale datasets. Furthermore, while deep learning offers advantages in processing complex data, effectively modeling the characteristics of heterogeneous multi-omics data remains a challenge. For example, the distributional characteristics of diverse omics data can lead to model instability. Probabilistic models and Bayesian frameworks are typically computationally intensive and have slow convergence, making them difficult to process large-scale spatial multi-omics data. Heterogeneous multi-omics data refers to diverse omics data from different sequencing methods and platforms, essentially referring to the raw spatial multi-omics data. Data features include gene expression levels and protein capture levels, i.e., the measured values ​​within the diverse omics data.

[0072] Although related technologies have proposed methods for biological analysis of spatial multi-omics data by combining models, the distribution characteristics of different spatial multi-omics data can easily lead to model instability, making it difficult for related technologies to process large-scale spatial multi-omics data, thereby affecting their analytical results. Therefore, how to provide a biological analysis method that can be applied to large-scale spatial multi-omics data has become a technical problem that needs to be solved urgently.

[0073] Based on this, the embodiments of the present application propose a biological analysis method, apparatus, device and medium based on spatial multi-omics data, which can be better applied to the biological analysis of large-scale spatial multi-omics data.

[0074] The following describes the biological analysis method based on spatial multi-omics data provided in the examples of the present application.

[0075] Reference Figure 1 In some embodiments, the biological analysis method based on spatial multi-omics data provided in the embodiments of the present application includes but is not limited to steps S101 to S105.

[0076] Step S101: Acquire spatial multi-omics data of a target slice.

[0077] In step S101 of some embodiments, the target slice may refer to a tissue slice to be processed for spatial multi-omics data integration. The target slice is a biological slice obtained in compliance with relevant legal provisions, and the biological slice includes a plant slice and an animal slice. Spatial multi-omics data is used to characterize the data obtained after a plurality of spatial omics data of the target slice are collected. Spatial multi-omics data includes at least two types of spatial omics data, and each type of spatial omics data is used to describe the biological expression information corresponding to a plurality of spatial sites in the target slice. Each spatial omics data corresponds to a spatial omics technology, which may be spatial transcriptomics, spatial proteomics, spatial metabolomics and other technologies. For example, spatial multi-omics data may include spatial transcriptome data and spatial proteome data, that is, spatial transcriptome data and spatial proteome data are different spatial omics data, and exist in pairs according to the relationship of the target slice. Spatial transcriptome data may include a gene expression matrix corresponding to a plurality of spatial sites in the target slice, and spatial proteomics may include a protein expression matrix corresponding to a plurality of spatial sites in the target slice.

[0078] Among them, the spatial site refers to each spatial position in the target slice, that is, a spot or a cell. "Spot" usually refers to a specific small area on the target slice, which may contain one or more cells. For example, in spatial transcriptomics technology, each spot can be regarded as a data point, which contains the gene expression information of the area. The multiple spatial omics data in the embodiments of the present application may contain the same number of spatial sites, and form a corresponding relationship based on the spatial sites. For example, spatial multi-omics data includes spatial transcriptomics data and spatial proteomics data, then each spatial site spot can respectively include a gene expression matrix and a protein expression matrix. In this way, for each spot, it can correspond to multiple different spatial omics data.

[0079] This application can be applied to the integration of spatial multi-omics data between multiple slices of the same sample, and can also be used to process the integration of spatial multi-omics data of a single slice. For example, for multiple slices of the same sample, the target slice at this time refers to each slice obtained based on the same sample, then different spatial omics technologies can be used for each target slice to obtain multiple spatial omics data. Further, at least two spatial omics data sets of multiple target slices are combined to form spatial multi-omics data.

[0080] Step S102 : Feature encoding is performed on at least two types of spatial omics data respectively to obtain at least two spatial omics feature vectors corresponding to each spatial site.

[0081] Feature encoding refers to the process of converting raw spatial omics data into a format suitable for processing by machine learning models. Each spatial omics feature vector represents the key information extracted from the corresponding spatial omics data, converted into data suitable for model processing. In other words, each spatial omics feature vector can be a feature vector or feature matrix of the corresponding spatial omics data, reflecting the key biological information of the corresponding spatial omics data. Each spatial omics feature vector records the characteristic information of multiple spatial sites in the same feature space.

[0082] It should be noted that this application can first perform data preprocessing on each type of spatial omics data, such as standardization and normalization of each original spatial omics data, to ensure the consistency of these spatial omics data in numerical range and distribution, thereby eliminating the scale differences between different spatial omics data. In this way, the preprocessed data is updated with the corresponding spatial omics data, thereby improving the training effect of the model during the training process and the generalization ability of the use process.

[0083] In some embodiments, step S102 may include but is not limited to: performing feature encoding on at least two spatial omics data according to a preset encoder to obtain at least two spatial omics feature vectors corresponding to each spatial site.

[0084] Among them, the embodiment of the present application can use a preset encoder to perform feature encoding on each spatial omics data, that is, each pre-processed spatial omics data can be input into the preset encoder respectively, and the features can be automatically extracted through the forward propagation process of the encoder to learn complex feature representations from the spatial omics data and realize multimodal feature encoding. The preset encoder is a pre-designed and trained model or algorithm component, which is equivalent to a "converter" with specific rules and structures, and can extract and convert features of the input data. The preset encoder of the embodiment of the present application can be constructed based on structures such as Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN). In addition, the present application can select a suitable encoder according to the characteristics and analysis objectives of the spatial omics data without specific limitation.

[0085] In some specific embodiments, the preset encoder of the embodiment of the present application may specifically include a feature extraction layer and a feature mapping layer. Among them, the feature extraction layer is equivalent to the first layer of the neural network of the preset encoder, which is used to generate low-dimensional expression features. Feature extraction is a key step in data analysis, involving identifying and extracting key information from the original data that is helpful for subsequent analysis and modeling. The feature mapping layer is equivalent to the second layer of the neural network of the preset encoder, which is used to map the different feature vectors output by the feature extraction layer into the same feature space, and is also equivalent to a deep coding layer. In this way, feature encoding through the preset encoder is equivalent to converting different feature vectors into representations embedded in a shared feature space through nonlinear mapping, which is also equivalent to integrating different feature vectors.

[0086] In some specific embodiments, in the step of performing feature encoding on at least two spatial omics data according to a preset encoder to obtain at least two spatial omics feature vectors corresponding to each spatial site: the feature extraction layer is used to perform feature extraction on each spatial omics data to obtain the initial omics feature corresponding to each spatial omics data; the feature mapping layer is used to perform feature mapping on each initial omics feature to obtain at least two spatial omics feature vectors corresponding to each spatial site.

[0087] Among them, when feature encoding is performed on each spatial omics data according to a preset encoder, the present application can first perform feature extraction on each spatial omics data according to the feature extraction layer to obtain the initial omics features corresponding to each spatial omics data. For example, when the spatial omics data is spatial transcriptome data, the present application can extract features from the spatial transcriptome data through the feature extraction layer to generate a low-dimensional gene expression feature vector. The spatial proteome data can then use the same method to extract a low-dimensional protein expression feature vector. Furthermore, the present application can perform feature mapping on each initial omics feature according to the feature mapping layer to obtain a spatial omics feature vector corresponding to each spatial omics data. In other words, each initial omics feature (such as a low-dimensional gene expression feature vector and a low-dimensional protein expression feature vector) can be nonlinearly mapped through the feature mapping layer to convert it into a representation embedded in a shared feature space, i.e., a variety of spatial omics feature vectors corresponding to each spatial site are obtained.

[0088] In the above embodiment, the feature extraction layer and the feature mapping layer belong to two different links in the preset encoder processing process and adopt different network structures. The present application can effectively solve the problem of inconsistent scales and distributions of different omics data through feature encoding, that is, by introducing the feature mapping layer, compared with the linear dimensionality reduction method in the related art, it can better extract the complex information in each spatial omics data, while retaining important features of biological significance to reduce information loss. In addition, the present application can map different spatial omics data to the same shared feature space through the combined use of the feature extraction layer and the feature mapping layer, realize multimodal feature encoding, and the generated high-quality embedding representation (that is, the obtained spatial omics feature vector) not only significantly improves the integration effect and realizes high-resolution clustering of tissue regions, but also significantly reduces the impact of noise on the analysis results.

[0089] Step S103 : constructing a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data based on the positional relationship between the spatial sites in the target slice.

[0090] Among them, the spatial multi-omics adjacency graph is used to characterize the spatial graph constructed based on the positional relationship between spatial sites of multiple spatial omics data. The spatial multi-omics adjacency graph includes multiple graph nodes, and each graph node corresponds to a spatial site of the target slice, so the spatial multi-omics adjacency graph can represent the relationship between different spatial sites. These graph nodes represent different areas in the target slice, and each area may contain one or more cells. It should be noted that for the spatial multi-omics adjacency graph, the connection (edge) between the graph nodes corresponding to each spatial site represents the spatial proximity or other biological relationship between the spatial sites.

[0091] Step S104 , performing feature integration on at least two spatial omics feature vectors according to the spatial multi-omics adjacency graph and the graph attention network to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice.

[0092] The spatial multi-omics integration feature is used to describe the features obtained by integrating the corresponding multiple spatial omics data at each spatial site in the target slice. By combining the dynamic modeling of spatial neighborhood relationships through the graph attention mechanism, this application can adaptively assign weights to each graph node, avoiding the problem of excessive dependence on adjacency relationships in related technologies based on fixed graph structures, and can more flexibly capture key spatial-specific signals.

[0093] Reference Figure 2 In some embodiments, step S104 may include but is not limited to steps S201 to S202.

[0094] Step S201 : for each spatial site, obtain corresponding multiple adjacent sites according to the spatial multi-omics adjacency graph.

[0095] Among them, since spatial omics data is used to describe the biological expression information corresponding to each spatial site in the target slice, the same spatial site can correspond to multiple spatial omics data measured based on different spatial omics technologies. In this way, for spatial multi-omics data, the present application can determine the spatial omics data group corresponding to each spatial site from multiple spatial omics data, and each spatial omics data group at this time includes at least two spatial omics feature vectors. Adjacent sites are other spatial sites with an adjacency relationship with the spatial site in the spatial multi-omics adjacency graph. The embodiment of the present application can facilitate subsequent spatial relationship research based on the perspective of spatial sites by dividing the spatial omics data group corresponding to each spatial site from multiple spatial omics data.

[0096] Step S202 : Based on the correlation between the spatial site and each adjacent site, feature fusion is performed on at least two spatial omics feature vectors of each spatial site to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice.

[0097] Because each spatial site corresponds to a graph node in the spatial multi-omics adjacency graph, each graph node in the spatial multi-omics adjacency graph contains both spatial location information and a feature representation of the spatial omics feature vectors corresponding to the site. Feature fusion is a crucial step in multimodal learning, allowing models to integrate information from different data types to improve the accuracy and depth of analysis. Feature fusion can be achieved through various methods, such as feature concatenation.

[0098] Furthermore, embodiments of the present application can use a Graph Attention Network (GAT) to perform processing on a spatial multi-omics adjacency graph to obtain a final feature representation of each graph node containing spatial position information and multiple spatial omics information, which also represents the result of integrating multiple spatial omics data.

[0099] In the above steps S201 to S202, the embodiment of the present application provides a biological analysis method for spatial multi-omics data based on the graph attention mechanism, which aims to effectively integrate multiple different spatial omics data, such as spatial transcriptome, spatial proteome and other data, and through a unified embedding representation, it can better achieve noise reduction and biological analysis of the tissue spatial microenvironment.

[0100] Reference Figure 3 In some embodiments, step S202 may include but is not limited to steps S301 to S304.

[0101] Step S301 : performing feature correlation calculation on at least two spatial omics feature vectors of a spatial site and at least two spatial omics feature vectors of each adjacent site to obtain site correlation data between the spatial site and each adjacent site.

[0102] Among them, the adjacent sites are used to characterize the spatial sites adjacent to the spatial sites in the spatial multi-omics adjacency graph, and can also be called neighbor nodes. The adjacent sites corresponding to each spatial site can be 0, 1 or more, without limitation. For example, if a spatial site is at the edge of the target slice, there may not be enough adjacent sites to constitute a complete adjacency relationship, or under a specific neighborhood definition (for example, based on a specific distance threshold or number of neighbors), a spatial site may not meet the conditions for becoming a neighbor. In this case, the spatial site has no adjacent sites, that is, the number of neighbor nodes is 0. Site correlation data is used to describe the site-related information between the spatial site and the corresponding adjacent sites. The present application can first calculate the correlation coefficient between different spatial sites by the adjacency relationship expressed in the spatial multi-omics adjacency graph and the similarity of the spatial omics feature vectors. Because the adjacent sites corresponding to each spatial site can be 0, 1 or more, the site correlation data corresponding to each spatial site can be 0, 1 or more. In subsequent embodiments, the spatial site will be illustrated by including multiple adjacent sites.

[0103] It should be noted that when a spatial site contains at least one adjacent site, the process of calculating the correlation data of each site is as shown in the following formula 1:

[0104] (Formula 1)

[0105] In formula 1, Represents spatial location (Also known as node ) and spatial sites (Also known as node ), also indicates the site correlation data between nodes For nodes Correlation coefficient, spatial site is a spatial site An adjacent site. It is an activation function that helps the model learn more complex feature relationships. represents the preset first linear transformation matrix, is the preset second linear transformation matrix, Used to increase the dimension of at least two spatial omics feature vectors corresponding to the node. Indicates that the spatial location and spatial sites The increased dimensional feature vectors are concatenated to merge the features of the two nodes together, so that the information of the two nodes is considered simultaneously when calculating the correlation. Used to map the concatenated high-dimensional features into real numbers. Represents spatial location Corresponding to at least two spatial omics feature vectors, Represents spatial location Corresponding to at least two spatial omics feature vectors.

[0106] It should be noted that the spatial site in formula 1 In essence, it is all the spatial locations. Each time it is calculated, it is for the spatial locations and all its corresponding adjacent sites (i.e. spatial sites ) to calculate.

[0107] Step S302: normalize the site correlation data.

[0108] Here, each site correlation data is first normalized to obtain normalized correlation data. The normalized correlation coefficient at this time is also equivalent to the attention coefficient. The process of normalizing the site correlation data is shown in the following formula 2:

[0109] (Formula 2)

[0110] In formula 2, Represents spatial location (Also known as node ) and spatial sites (Also known as node ) after normalization of the site correlation data between the nodes, and also represents the node For nodes Attention coefficient, spatial location is a spatial site An adjacent site. Represents spatial location The set of all adjacent sites.

[0111] Step S303 , performing feature weighting and calculation on the normalized site correlation data and at least two spatial omics feature vectors corresponding to each adjacent site to obtain a spatial multi-omics weighted feature corresponding to each spatial site.

[0112] Among them, the present application can feature-weight at least two spatial omics feature vectors of each adjacent site based on the normalized site correlation data to obtain the spatial multi-omics weighted features corresponding to each adjacent site, and perform feature summation processing on the spatial omics weighted features of all adjacent sites corresponding to each spatial site to obtain the spatial multi-omics weighted features corresponding to each spatial site.

[0113] Step S304 , activating the spatial multi-omics weighted features corresponding to each spatial site to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice.

[0114] The spatial multi-omics weighted features corresponding to each spatial site are activated to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice. In this way, this application constructs a spatial multi-omics adjacency graph and uses the graph attention mechanism to effectively integrate spatial location information and omics information, thereby obtaining an integrated feature representation containing multiple spatial omics data.

[0115] It should be noted that the process of determining the spatial multi-omics integration features corresponding to each spatial site in the target slice based on the normalized correlation data in this application is shown in the following formula 3:

[0116] (Formula 3)

[0117] In formula 3, Represents spatial location (Also known as node ) can be applied to other spatial sites to obtain the spatial multi-omics integration features corresponding to each spatial site in the target slice. is the preset second linear transformation matrix, Used to increase the dimension of at least two spatial omics feature vectors corresponding to the node. Represents each adjacent site (i.e. node ) corresponding to the spatial multi-omics weighted features. Represents spatial location The set of all adjacent sites. Represents each spatial site The spatial multi-omics weighted features are obtained by summing up the spatial omics weighted features of all corresponding adjacent sites. In this way, we can obtain the spatial multi-omics integrated features of each spatial site in the target slice after processing by the graph attention mechanism.

[0118] In the above embodiment, the present application calculates the site correlation data between the graph nodes in the spatial multi-omics adjacency graph (i.e., the correlation coefficient between the spatial site and the corresponding adjacent site) through the graph attention mechanism, thereby determining the connection strength between the graph nodes. In this way, the relationship between the spatial position information and expression characteristics in the spatial transcriptome and proteome data can be better captured. In addition, the present application dynamically models the relationship between spatial neighbors through the graph attention mechanism, and can adaptively assign weights (i.e., determine the degree of weight adjustment for at least two spatial omics feature vectors based on the site correlation data), solving the problem of the related technology's excessive dependence on adjacency relationships based on fixed graph structure methods, thereby being able to more flexibly capture key spatial specific signals and better suited for the biological analysis of large-scale spatial multi-omics data.

[0119] Step S105 , performing biological analysis based on the spatial multi-omics integration characteristics to obtain analysis results.

[0120] Among them, the spatial multi-omics integration features obtained in this application can be used as the dimensionality reduction results of PCA and used for downstream analysis tasks after PCA, such as two-dimensional display, clustering, spatial region segmentation, etc. Specifically, this application can use nonlinear dimensionality reduction tools such as Uniform Manifold Approximation and Projection (UMAP) to process the spatial multi-omics integration features and reduce them to two-dimensional space display. In addition, this application can also use the results of the dimensionality reduction processing of the spatial multi-omics integration features to calculate the shared nearest neighbor similarity (SNN), and then cluster the calculation results according to different resolutions.

[0121] Reference Figure 4 In some embodiments, the training process of the preset encoder and the graph attention network may specifically include the following steps S401 to S407.

[0122] Step S401: Acquire sample space multi-omics data of a sample slice.

[0123] The sample space multi-omics data includes at least two types of sample space omics data, each type of sample space omics data is used to describe the biological expression information corresponding to multiple sample space sites in the sample slice. The sample space multi-omics data has the same meaning as the spatial multi-omics data in the above embodiment, except that it is used for model training here and is not further described.

[0124] Step S402 : Feature encoding is performed on at least two types of sample space omics data respectively to obtain at least two sample space omics feature vectors corresponding to each sample space site.

[0125] Among them, each sample space omics data is feature encoded according to the preset encoder to obtain the sample space omics feature vector corresponding to each sample space omics data, and the sample space omics feature vector has the same meaning as the space omics feature vector in the above embodiment, but it is used for model training here and will not be repeated.

[0126] Step S403 : constructing a sample space multi-omics adjacency graph corresponding to the sample space multi-omics data based on the positional relationship between the sample space sites in the sample slice.

[0127] The sample space multi-omics adjacency graph has the same meaning as the spatial multi-omics adjacency graph in the above embodiment, except that it is used for model training and will not be described in detail.

[0128] Step S404 , performing feature integration on at least two sample space omics feature vectors according to the sample space multi-omics adjacency graph and the graph attention network to obtain the sample space multi-omics integrated features corresponding to each sample space site in the sample slice.

[0129] The sample space multi-omics integration feature has the same meaning as the spatial multi-omics integration feature in the above embodiment, except that it is used for model training and will not be described in detail here.

[0130] Step S405 : constructing a data reconstruction loss based on the sample space multi-omics data and multiple sample space multi-omics integration features.

[0131] Among them, in order to ensure that the spatial multi-omics integration features after embedding representation can accurately retain the key information of the input spatial multi-omics data, this application can construct data reconstruction loss based on sample spatial multi-omics data and multiple sample spatial multi-omics integration features.

[0132] In some embodiments, the step of constructing a data reconstruction loss based on the sample space multi-omics data and multiple sample space multi-omics integration features may specifically include:

[0133] Decode the multi-omics integration features of multiple sample spaces to obtain the sample reconstruction data corresponding to the omics data of each sample space;

[0134] The data error calculation is performed on each sample space omics data and the corresponding sample reconstruction data to obtain the data reconstruction error corresponding to each sample space omics data;

[0135] The data reconstruction loss is constructed based on multiple data reconstruction errors.

[0136] Multiple sample spatial multi-omics integration features are used to indicate the spatial multi-omics integration features corresponding to each spatial site in the target slice. Feature decoding is the process of restoring each spatial multi-omics integration feature to the original spatial omics data or a form close to the original spatial omics data. Sample reconstruction data refers to the reconstruction results corresponding to the original spatial omics data obtained through feature decoding, such as spatial transcriptome reconstruction data corresponding to spatial transcriptome data, spatial proteome reconstruction data corresponding to spatial proteome data, etc.

[0137] It should be noted that, for example, the sample space multi-omics data includes spatial transcriptome data and spatial proteome data. By decoding and reconstructing the feature of multiple sample space multi-omics integration features, spatial transcriptome reconstruction data and spatial proteome reconstruction data can be obtained. The process of constructing the data reconstruction loss can be shown in the following formula 4:

[0138] (Formula 4)

[0139] In formula 4, represents the data reconstruction loss, represents the input spatial transcriptome data, Represents the spatial transcriptome reconstruction data corresponding to the input spatial transcriptome data (i.e., sample reconstruction data), represents the input spatial transcriptome data, represents the spatial proteome reconstruction data (i.e., sample reconstruction data) corresponding to the input spatial proteome data. Represents the data reconstruction error corresponding to the input spatial transcriptome data, Represents the data reconstruction error corresponding to the input spatial proteomic data.

[0140] Step S406 : constructing a neighborhood preservation loss based on the sample space multi-omics adjacency graph and at least two sample space omics feature vectors.

[0141] In order to preserve the local topological relationship of spatial data, the present application also designs a neighborhood preservation loss. In some embodiments, the step of constructing a neighborhood preservation loss based on the sample space multi-omics adjacency graph and multiple sample space omics feature vectors may specifically include:

[0142] Convert the sample space multi-omics adjacency graph into the corresponding adjacency matrix;

[0143] For each sample space site, obtain the corresponding multiple sample adjacent sites according to the sample space multi-omics adjacency graph;

[0144] Performing feature difference calculation on at least two sample spatial omics feature vectors of the sample spatial site and at least two sample spatial omics feature vectors of each sample adjacent site to obtain multi-omics difference features of the sample spatial site and each sample adjacent site;

[0145] A neighborhood preservation loss is constructed based on the adjacency matrix and multi-omics difference features corresponding to multiple sample spatial sites.

[0146] It should be noted that the sample adjacent sites are other sample space sites that have an adjacent relationship with the sample space site in the sample space multi-omics adjacency graph. The process of constructing the neighborhood preservation loss (Spatial Consistency Loss) can be shown in the following formula 5:

[0147] (Formula 5)

[0148] In formula 5, represents the constructed neighborhood preservation loss, Represents the graph matrix form of the constructed sample space multi-omics adjacency graph transformation, Represents spatial location (Also known as node ) and spatial sites (Also known as node ) in the figure matrix A corresponding to the data, Represents spatial location The corresponding first fusion feature vector (equivalent to at least two sample spatial omics feature vectors of the spatial site in the above embodiment, and the determination method has been described in the above embodiment and will not be repeated here), Represents spatial location The corresponding second fusion feature vector (equivalent to at least two sample space omics feature vectors of adjacent sites in the above embodiment, which will not be described in detail).

[0149] Step S407: construct a target loss based on the data reconstruction loss and the neighborhood preservation loss, and update the model parameters of the preset encoder and the graph attention network based on the target loss.

[0150] Among them, this application can ensure that the embedded representation can accurately retain the key information of the input data by constructing the data reconstruction loss, and can encourage the embedded representations of adjacent nodes to be more similar by constructing the neighborhood preservation loss. Furthermore, the target loss is constructed based on the data reconstruction loss and the neighborhood preservation loss, that is, the final optimization target of the model is determined. The process of constructing the target loss is shown in the following formula 6:

[0151] (Formula 6)

[0152] In formula 6, represents the target loss, represents the data reconstruction loss, represents the neighborhood preservation loss.

[0153] In other embodiments, the present application may first determine the weights corresponding to the data reconstruction loss and the neighborhood preservation loss respectively, and then perform weighted calculation on the data reconstruction loss and the neighborhood preservation loss based on the determined weights to obtain the target loss. The set weights are used to reflect the degree of influence of different training requirements on the model training process to improve the flexibility of determining the target loss. Furthermore, the target loss obtained in the present application can synchronously update the model parameters of the preset encoder and the graph attention network.

[0154] For example, in some embodiments, the present application may also utilize a contrastive learning framework to further enhance the biological significance of the embedded representation by constructing positive and negative sample pairs during the encoding phase. Contrastive learning can construct positive and negative sample pairs during model training and set a learning contrast loss to ensure more accurate embedding results.

[0155] In the above embodiment, the present application can construct a target loss by self-supervised decoding and reconstruction of data, and use the constructed target loss to train the preset encoder and graph attention network. In this way, the obtained preset encoder and graph attention network can effectively integrate a variety of different spatial omics data into a shared embedding space, solving the integration difficulty problem caused by the distribution and scale differences of different omics data in related technologies, so that the generated embedding representation more comprehensively retains the biological information of multimodal data.

[0156] The biological analysis method based on spatial multi-omics data provided in the embodiment of the present application, when performing biological analysis based on spatial multi-omics data, can not only construct a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data according to each spatial site in the target slice, but also perform feature integration of multiple spatial omics feature vectors according to the spatial multi-omics adjacency graph and the graph attention network, so as to avoid the problem of inaccurate data integration due to the distribution characteristics of different spatial omics data. That is, by introducing the spatial multi-omics adjacency graph and the graph attention network, the embodiment of the present application can adaptively model the spatial neighborhood relationship and dynamically allocate weights between nodes, overcome the limitation of the related technology that relies too much on the fixed adjacency graph, and can more accurately capture the complex interactions between cells or tissue regions in the spatial microenvironment, improve the accuracy and robustness of spatial relationship modeling, and improve the integration ability of spatial multi-omics data, so that it can be better applied to the biological analysis method of large-scale spatial multi-omics data.

[0157] Reference Figure 5 In a specific embodiment, the method provided in the embodiment of the present application may include the following steps:

[0158] Step S501 : acquiring spatial multi-omics data of a target slice, and performing feature encoding on each type of spatial multi-omics data according to a preset encoder to obtain a spatial multi-omics feature vector corresponding to each type of spatial multi-omics data.

[0159] Among them, the target slice at this time can be the thymus slice of the experimental mouse, and the spatial multi-omics data obtained can be Stereo-CITE-seq data. Among them, Stereo-CITE-seq is a high-resolution spatial multi-omics sequencing technology that combines the two technologies of Cellular Indexing of Transcriptomes and Epitopes by Sequencing (CITE-seq) and Spatial Enhanced REsolution Omics-sequencing (Stereo-seq). It can achieve co-detection of RNA and protein in the same tissue slice with high spatial resolution. Figures 6A-6C As shown in FIG, the spatial multi-omics integrated features corresponding to the thymus tissue obtained by combining the biological analysis method of this application are clustered. Figure 6A As shown, this application can identify different spatial regions in Stereo-CITE-seq data and identify their corresponding functional structures in combination with cell type markers. Figure 6BAs shown in Figure 2, after the Stereo-CITE-seq data is processed by the preset encoder and graph attention network of this application, the corresponding UMAP visualization results can be obtained. Figure 6C As shown, the present application can name the divided spatial regions as Capsule region, Cortex region 1-3, Cortico-medullary Junction (CMJ region 1-3) and Medullaregion 1-2 in sequence. These spatial regions are consistent with the functional structure of the thymus.

[0160] It should be noted that this application can also simultaneously input other multi-scale information (such as long-range interactions between cells) when extracting features from each spatial omics data. In this case, tools such as CellChat can be used to calculate the interactions between different regions to increase the accuracy of data integration.

[0161] Step S502 : constructing a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data based on the positional relationship between the spatial sites in the target slice.

[0162] It should be noted that this application also introduces a spatial graph nesting structure, utilizing it to improve the graph attention network to better represent information in the slice space. In deep learning, this spatial graph nesting structure generally refers to a hierarchical structure used to handle problems with multiple levels of spatial information. This structure is commonly found in graph neural networks (Graph Neural Networks) and some vision tasks. Specifically, a spatial graph nesting structure refers to the existence of different levels of relationships and connections between nodes in a graph-like data structure (such as social networks, chemical molecular structures, and transportation networks), and the model needs to be able to effectively utilize this hierarchical information for learning and reasoning. In graph neural networks, the spatial graph nesting structure can be expressed as multi-layered graph convolutional networks (GCNs) or models with hierarchical attention mechanisms. These structures allow the model to transfer and aggregate information between nodes and edges in the graph at multiple levels, thereby better capturing the complexity and hierarchical nature of the graph. In general, a spatial graph nesting structure refers to a structure used in deep learning models to handle multi-level spatial information. Through hierarchical information transfer and aggregation, the model can better understand and utilize the hierarchical information present in the data.

[0163] Step S503 , performing feature integration on at least two spatial omics feature vectors according to the spatial multi-omics adjacency graph and the graph attention network to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice.

[0164] Among them, because each spatial site can correspond to a graph node of the spatial multi-omics adjacency graph, each graph node in the spatial multi-omics adjacency graph contains spatial position information and feature representations of at least two spatial omics feature vectors.

[0165] Step S504: determine the adjacent sites corresponding to each spatial site based on the spatial multi-omics adjacency graph, and perform feature correlation calculation based on at least two spatial omics feature vectors of the spatial site and at least two spatial omics feature vectors of the corresponding adjacent sites to determine the site correlation data between the spatial site and the corresponding adjacent sites.

[0166] Among them, site correlation data is used to describe the site-related information between a spatial site and its corresponding adjacent sites. This application can first calculate the correlation coefficient between different spatial sites through the adjacency relationship expressed in the spatial multi-omics adjacency graph and the similarity of the spatial omics feature vectors.

[0167] Step S505 : updating the features of at least two spatial omics feature vectors according to the site correlation data of each spatial site, and obtaining the spatial multi-omics integrated features corresponding to each spatial site in the target slice.

[0168] Among them, the present application can perform feature weighting on at least two spatial omics features of each adjacent site based on normalized related data to obtain the spatial multi-omics weighted features corresponding to each adjacent site, and sum the spatial omics weighted features of all adjacent sites corresponding to each spatial site to obtain the spatial multi-omics weighted features corresponding to each spatial site. Afterwards, the spatial multi-omics weighted features corresponding to each spatial site are activated to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice. In this way, the present application constructs a spatial multi-omics adjacency graph and uses the graph attention mechanism to effectively integrate spatial position information and omics information, thereby obtaining an integrated feature representation containing multiple spatial omics data.

[0169] It can be understood that this embodiment effectively integrates spatial transcriptome and spatial proteome data into a shared embedding space through an encoding-decoding self-supervision method. It solves the integration difficulties caused by the distribution and scale differences of different omics data in traditional methods, and the generated embedding representation more comprehensively retains the biological information of multimodal data. It is a method for analyzing cell interactions at single-cell precision in spatial transcriptome research, which can simultaneously utilize cell group annotation results and spatial information, and the results are biologically interpretable. Furthermore, by introducing the Graph Attention Mechanism (GAT), spatial neighborhood relationships are adaptively modeled and weights between nodes are dynamically assigned. It overcomes the limitations of traditional graph models that rely on fixed adjacency graphs, can more accurately capture the complex interactions between cells or tissue regions in spatial microenvironments, and improves the accuracy and robustness of spatial relationship modeling.

[0170] Step S506 , performing biological analysis based on the spatial multi-omics integration characteristics to obtain analysis results.

[0171] Among them, such as Figures 6A-6C As shown in FIG, the spatial multi-omics integrated features corresponding to the thymus tissue obtained by combining the biological analysis method of this application are clustered. Figure 6A As shown, this application can identify different spatial regions in Stereo-CITE-seq data and identify their corresponding functional structures in combination with cell type markers. Figure 6B As shown in Figure 2, after the Stereo-CITE-seq data is processed by the preset encoder and graph attention network trained in this application, the UMAP visualization results corresponding to the embedding space of the spatial multi-omics integrated features are obtained. Figure 6C As shown, the present application can name the divided spatial regions as capsule region, cortex region 1-3, cortico-medullary junction region 1-3 (CMJ region 1-3) and medulla region 1-2 in sequence. These spatial regions are consistent with the functional structure of the thymus. Figure 7 The figure below shows a schematic diagram of clusters obtained by spatially clustering the same data using MOFA+ technology. The different colors represent different regional categories after clustering. Compared with the present application, the information clustered by MOFA+ technology is more chaotic and lacks intuitive spatial connections, which can easily affect the results of biological analysis.

[0172] Reference Figure 8, the embodiment of the present application further provides a biological analysis device based on spatial multi-omics data, the device comprising:

[0173] An acquisition module 801 is configured to acquire spatial multi-omics data of a target slice, where the spatial multi-omics data includes at least two types of spatial omics data, and each type of spatial omics data is used to describe biological expression information corresponding to multiple spatial sites in the target slice;

[0174] Extraction module 802, for performing feature encoding on at least two spatial omics data respectively to obtain at least two spatial omics feature vectors corresponding to each spatial site;

[0175] A graph construction module 803 is used to construct a spatial multi-omics adjacency graph corresponding to the spatial multi-omics data based on the positional relationship between the spatial sites in the target slice;

[0176] An integration module 804 is configured to perform feature integration on at least two spatial omics feature vectors based on a spatial multi-omics adjacency graph and a graph attention network to obtain spatial multi-omics integrated features corresponding to each spatial site in the target slice;

[0177] The analysis module 805 is used to perform biological analysis based on the spatial multi-omics integration characteristics to obtain analysis results.

[0178] It can be seen that the contents of the above-mentioned embodiments of the biological analysis method based on spatial multi-omics data are all applicable to the embodiments of the biological analysis device based on spatial multi-omics data. The functions specifically implemented by the embodiments of the biological analysis device based on spatial multi-omics data in this application are the same as those in the above-mentioned embodiments of the biological analysis method based on spatial multi-omics data, and the beneficial effects achieved are also the same as the beneficial effects achieved by the above-mentioned embodiments of the biological analysis method based on spatial multi-omics data.

[0179] Reference Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0180] The processor 901 may be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0181] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the biological analysis method based on spatial multi-omics data in the embodiments of this application.

[0182] Input / output interface 903, used to implement information input and output;

[0183] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0184] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );

[0185] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0186] The present application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device implements the above-mentioned biological analysis method based on spatial multi-omics data.

[0187] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "comprises" and "comprising," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.

[0188] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0189] It should be understood that in the description of the embodiments of the present application, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.

[0190] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0191] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0192] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0194] It should also be understood that the various implementation methods provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0195] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A biological analysis method based on spatial multi-omics data, characterized in that: The method comprises: Acquire spatial multi-omics data of a target slice, wherein the spatial multi-omics data includes at least spatial transcriptome data and spatial proteome data, wherein the spatial transcriptome data includes gene expression data of each spatial site of the target slice, and the spatial proteome data includes protein expression data of each spatial site of the target slice; Performing feature encoding on the gene expression data and the protein expression data respectively, and mapping the encoded features to the same shared feature space to obtain a gene expression feature vector and a protein expression feature vector corresponding to each of the spatial sites; Based on the positional relationship between the spatial sites in the target slice and the gene expression feature vector and protein expression feature vector corresponding to each spatial site, a spatial multi-omics adjacency graph with the spatial sites as nodes is constructed; The node features of each node in the spatial multi-omics adjacency graph are updated according to the graph attention network to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice; wherein, the node features of each node in the spatial multi-omics adjacency graph are updated according to the graph attention network to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice, including: for each of the spatial sites, according to the spatial multi-omics adjacency graph, obtaining a corresponding plurality of adjacent sites; the adjacent sites are other spatial sites in the spatial multi-omics adjacency graph that have an adjacent relationship with the spatial site; according to the graph attention network, the correlation between the spatial site and each of the adjacent sites is determined, and the node features corresponding to each of the spatial sites are subjected to feature fusion to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice; Biological analysis is performed based on the spatial multi-omics integration characteristics to obtain analysis results.

2. The method according to claim 1, characterized in that The determining of the correlation between the spatial site and each of the adjacent sites based on the graph attention network, performing feature fusion on the node features corresponding to each of the spatial sites, and obtaining the spatial multi-omics integrated features corresponding to each spatial site in the target slice include: Performing feature correlation calculation on the node feature corresponding to the spatial site and the node feature corresponding to each of the adjacent sites to obtain site correlation data between the spatial site and each of the adjacent sites; performing normalization processing on the site correlation data; Performing feature weighting and calculation on the normalized site correlation data and the node features corresponding to each adjacent site to obtain the spatial multi-omics weighted features corresponding to each spatial site; The spatial multi-omics weighted features corresponding to each of the spatial sites are activated to obtain the spatial multi-omics integrated features corresponding to each spatial site in the target slice.

3. The method according to claim 1, characterized in that The feature encoding of the gene expression data and the protein expression data is performed separately, and the encoded features are mapped to the same shared feature space to obtain the gene expression feature vector and the protein expression feature vector corresponding to each of the spatial sites, including: Performing feature encoding on the gene expression data and the protein expression data respectively according to a preset encoder, and mapping the encoded features to the same shared feature space to obtain a gene expression feature vector and a protein expression feature vector corresponding to each of the spatial sites, wherein the preset encoder includes a feature extraction layer and a feature mapping layer; The feature extraction layer is used to perform feature extraction on the gene expression data and the protein expression data respectively to obtain a low-dimensional gene expression feature vector and a low-dimensional protein expression feature vector corresponding to each of the spatial sites; The feature mapping layer is used to perform feature mapping on the low-dimensional gene expression feature vector and the low-dimensional protein expression feature vector corresponding to each of the spatial sites, and obtain the gene expression feature vector and the protein expression feature vector corresponding to each of the spatial sites in the shared feature space.

4. The method according to claim 3, characterized in that The training process of the preset encoder and the graph attention network includes the following steps: Acquire sample space multi-omics data of the sample slice, wherein the sample space multi-omics data at least includes sample space transcriptome data and sample space proteome data, wherein the sample space transcriptome data includes sample gene expression data of each sample space site of the sample slice, and the sample space proteome data includes sample protein expression data corresponding to each sample space site of the sample slice; Performing feature encoding on the sample gene expression data and the sample protein expression data respectively, and mapping the encoded features to the same shared feature space to obtain a sample gene expression feature vector and a sample protein expression feature vector corresponding to each sample space site; Based on the positional relationship between each sample spatial site in the sample slice, and the sample gene expression feature vector and the sample protein expression feature vector corresponding to each sample spatial site, a sample space multi-omics adjacency graph with the sample spatial site as a node is constructed; updating the sample node features of each sample node in the sample space multi-omics adjacency graph according to the graph attention network to obtain the sample space multi-omics integrated features corresponding to each sample space site in the sample slice; constructing a data reconstruction loss based on the multi-omics data of the sample space and sample node features of the plurality of sample nodes; constructing a neighborhood preservation loss according to sample node features of a plurality of sample nodes in the multi-omics adjacency graph of the sample space; A target loss is constructed according to the data reconstruction loss and the neighborhood preservation loss, and model parameters of the preset encoder and the graph attention network are updated based on the target loss.

5. The method according to claim 4, characterized in that The constructing of data reconstruction loss according to the multi-omics data of the sample space and the sample node features of the plurality of sample nodes includes: Performing feature decoding on sample node features of the plurality of sample nodes to obtain sample reconstruction data corresponding to the sample space transcriptome data and sample reconstruction data corresponding to the sample space proteome data; Performing data error calculation on the sample space transcriptome data and the corresponding sample reconstruction data, and on the sample space proteome data and the corresponding sample reconstruction data, respectively, to obtain a data reconstruction error corresponding to the sample space transcriptome data and a data reconstruction error corresponding to the sample space proteome data; The data reconstruction loss is constructed according to the data reconstruction errors of multiple sample nodes.

6. The method according to claim 5, characterized in that The constructing a neighborhood preservation loss according to the sample node features of the plurality of sample nodes in the multi-omics adjacency graph of the sample space includes: Converting the sample space multi-omics adjacency graph into a corresponding adjacency matrix; For each of the sample space sites, a corresponding plurality of sample adjacent sites are obtained according to the sample space multi-omics adjacency graph; the sample adjacent sites are other sample space sites in the sample space multi-omics adjacency graph that have an adjacency relationship with the sample space site; Performing feature difference calculation on the sample node feature corresponding to the sample spatial site and the sample node feature corresponding to each of the sample adjacent sites to obtain multi-omics difference features between the sample spatial site and each of the sample adjacent sites; The neighborhood preservation loss is constructed according to the adjacency matrix and the multi-omics difference features corresponding to the plurality of sample spatial sites.

7. A biological analysis device based on spatial multi-omics data, characterized in that: The device comprises: an acquisition module, configured to acquire spatial multi-omics data of a target slice, wherein the spatial multi-omics data comprises at least spatial transcriptomic data and spatial proteomic data, wherein the spatial transcriptomic data comprises gene expression data at each spatial site of the target slice, and the spatial proteomic data comprises protein expression data at each spatial site of the target slice; An encoding module is used to perform feature encoding on the gene expression data and the protein expression data respectively, and map the encoded features to the same shared feature space to obtain a gene expression feature vector and a protein expression feature vector corresponding to each spatial site; A graph construction module is used to construct a spatial multi-omics adjacency graph with the spatial sites as nodes based on the positional relationship between the spatial sites in the target slice and the gene expression feature vector and protein expression feature vector corresponding to each spatial site; An integration module is used to update the node features of each node in the spatial multi-omics adjacency graph according to a graph attention network to obtain the spatial multi-omics integration features corresponding to each spatial site in the target slice; wherein, updating the node features of each node in the spatial multi-omics adjacency graph according to the graph attention network to obtain the spatial multi-omics integration features corresponding to each spatial site in the target slice includes: for each of the spatial sites, obtaining a corresponding plurality of adjacent sites according to the spatial multi-omics adjacency graph; the adjacent sites are other spatial sites in the spatial multi-omics adjacency graph that have an adjacent relationship with the spatial site; determining the correlation between the spatial site and each of the adjacent sites according to the graph attention network, performing feature fusion on the node features corresponding to each of the spatial sites, and obtaining the spatial multi-omics integration features corresponding to each spatial site in the target slice; The analysis module is used to perform biological analysis based on the spatial multi-omics integration characteristics to obtain analysis results.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the biological analysis method based on spatial multi-omics data according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the biological analysis method based on spatial multi-omics data according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the biological analysis method based on spatial multi-omics data according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Spatial multi-omics data integration method based on graph attention and multivariate loss function

    CN120260688A