Spatial multi-omics data integration method based on graph attention and multivariate loss function

By constructing graph structure and multivariate loss function, combining graph attention encoder and decoder to integrate spatial multi-omic data, the problem of insufficient de-batch and spatial smoothness in the existing technology is solved, and more efficient spatial microscopy data analysis is achieved.

CN120260688APending Publication Date: 2025-07-04SHENZHEN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510234007.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has problems of insufficient debatch, omics alignment and spatial smoothness in spatial omics data integration, which affects the accuracy of the analysis.

Method used

Using a method based on graph attention and multivariate loss function, data fusion is carried out by constructing a graph structure, and data reconstruction is carried out using graph attention encoder and decoder, and iterative optimization is carried out in combination with multiple loss functions to improve the smoothness of spatial multi-omic data and the accuracy of omic alignment.

Benefits of technology

It realizes efficient integration of spatial multi-omic data, improves the accuracy of spatial analysis and the accuracy of omic alignment, and solves the problems of batch effect and information loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260688A_ABST
    Figure CN120260688A_ABST
Patent Text Reader

Abstract

The invention discloses a spatial multi-omics data integration method based on graph attention and a multivariate loss function. Comprising the following steps that spatial multi-omics data and spatial position coordinates corresponding to the spatial multi-omics data are obtained, a preset spatial multi-omics data model is input, and the model comprises a graph attention encoder and a decoder; the spatial multi-omics data model constructs a graph structure according to spatial multi-omics data and spatial position coordinates, a graph attention encoder performs cell spatial multi-omics data fusion based on the constructed graph structure, and a decoder performs data reconstruction based on fused data to complete cell spatial multi-omics data integration; and constructing a multivariate loss function to carry out iterative optimization on the cell space multi-omics data integration process of the space multi-omics data model, and outputting an optimal cell space multi-omics data integration result. According to the method, the batch effect is solved, the smoothness of a final result in space and the accuracy of spatial analysis are improved, and the integration of multi-omics data of the cell space is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computational biology, and in particular, to a method for integrating spatial omics data based on graph attention and a multi - loss function. Background Art

[0002] Spatial omics technology is an emerging method for cell omics measurement. Compared with traditional bulk cell sequencing and single - cell sequencing technologies, it can additionally provide the spatial location information of samples, thus providing an important way for studying the spatial distribution and physiological phenomena of cells. The analysis of spatial omics data has high complexity, mainly because there are significant heterogeneities between different omics, significant batch effects between different slices, and the omics composition of different slices may be different, forming "mosaic" data. The integration of spatial mosaic data involves aligning the information of different omics, removing the batch effects between different batches of data, solving the problem of missing modalities, and finally achieving a unified representation of the data. By fusing different omics, more information can be provided to help comprehensively understand cells, and unifying the data of different slices helps to compare and analyze the spatial distribution and physiological behavior of cells.

[0003] In recent years, with the rapid development of deep learning technology, deep learning has shown great potential in the field of data mining and has become an efficient and high - performance solution for spatial mosaic data analysis. By integrating spatial omics data through deep learning, high - dimensional and sparse data can be transformed into a unified low - dimensional representation, and clustering, label transfer and other tasks can be achieved using this low - dimensional representation, becoming a reliable and effective method.

[0004] In past research, many integration methods for spatial omics data have been proposed. For example, in terms of multi - batch spatial multi - omics integration, Chen et al. proposed the SpaMosaic method. This method first performs de - batch processing on single - omics data through a third - party de - batch method, then uses a multi - layer graph neural network to extract omics data and spatial information, and realizes modality alignment through contrastive learning, and finally achieves the mosaic integration of spatial multi - batch multi - omics.

[0005] However, the de - batch processing of the SpaMosaic method completely depends on the third - party Harmony de - batch method. When the third - party method is unstable, it may have a negative effect on the results, and information loss will occur during the dimensionality reduction process. In addition, although traditional single - cell multi - omics methods have been able to achieve the integration of mosaic data at the single - cell level, due to the lack of full consideration of spatial information, the final results are often not smooth enough spatially, affecting the accuracy of spatial analysis. Summary of the Invention

[0006] There are still certain limitations in the prior art in aspects such as batch removal, omics alignment, and spatial smoothing during the integration of spatial omics data. Therefore, the purpose of the present invention is to propose a spatial multi-omics data integration method based on graph attention and a multi-loss function, which can effectively solve batch effects by combining multiple losses and fully consider spatial information to improve the spatial smoothness of the final result and the accuracy of omics alignment, thereby realizing the integration of spatial multi-omics data.

[0007] To achieve the purpose of the present invention, the present invention is implemented by the following technical solutions:

[0008] A spatial multi-omics data integration method based on graph attention and a multi-loss function, the method comprising the following steps:

[0009] Obtain spatial multi-omics data and its corresponding spatial position coordinates, and input them into a preset spatial multi-omics data model, the model including a graph attention encoder and a decoder;

[0010] The spatial multi-omics data model constructs a graph structure according to the spatial multi-omics data and spatial position coordinates, the graph attention encoder performs cell spatial multi-omics data fusion based on the constructed graph structure, and the decoder performs data reconstruction based on the fused data to complete cell spatial multi-omics data integration;

[0011] Construct a multi-loss function to iteratively optimize the cell spatial multi-omics data integration process of the spatial multi-omics data model, and output the optimal cell spatial multi-omics data integration result.

[0012] In the above technical solution, by constructing a graph structure according to the spatial multi-omics data and spatial position coordinates through a preset spatial multi-omics data model, the representation ability of omics data can be improved. The graph attention encoder is used to perform cell spatial multi-omics data fusion based on the constructed graph structure, and the decoder is used to perform data reconstruction based on the fused data to complete cell spatial multi-omics data integration, which can fully consider spatial information, so that the model can efficiently master the spatial multi-omics data and spatial position coordinate information; in addition, during the iterative optimization of the cell spatial multi-omics data integration process of the model by the constructed multi-loss function, it can effectively solve the challenges of batch removal effect and multi-omics data fusion through multiple losses and in combination with the graph attention network, improve the spatial smoothness, the accuracy of omics alignment and the accuracy of spatial analysis of the final result, thereby realizing the integration of spatial mosaic omics data (i.e., spatial multi-omics data).

[0013] Further, the spatial multi-omics data model preprocesses the input spatial multi-omics data, and the process includes:

[0014] The spatial multi-omics data includes transcriptomics data and epigenetics data;

[0015] Perform a log1p transformation on the transcriptomic data, and perform standardization and dimensionality reduction on the transformed transcriptomic data;

[0016] Convert the format of the epigenetic data into peak form, and perform high-variable feature screening and dimensionality reduction on the epigenetic data in peak form.

[0017] In the above technical solution, performing a log1p transformation on the transcriptomic data can select highly variable genes in the data. Performing standardization and dimensionality reduction on the transformed transcriptomic data can effectively reduce redundant features in the data. Converting the format of the epigenetic data into peak form and performing high-variable feature screening and dimensionality reduction on the epigenetic data in peak form can improve the practicality and reliability of the data.

[0018] Furthermore, the process of constructing a graph structure for the spatial multi-omics data model based on the spatial multi-omics data and spatial position coordinates includes:

[0019] According to the spatial position coordinates corresponding to each spatial omics data, use the K-nearest neighbor algorithm to calculate the corresponding adjacency matrix, and the expression is:

[0020] A s ∈{0,1} N×N

[0021] Construct a graph structure based on the adjacency matrix and its corresponding spatial omics data The expression is:

[0022]

[0023] where s represents the s-th slice in the spatial omics data, N represents a natural number, represents the set of m-omics data of all points on the s-th slice.

[0024] In the above technical solution, using the K-nearest neighbor algorithm can effectively calculate the corresponding adjacency matrix according to the spatial position coordinates corresponding to each spatial omics data, thereby constructing a graph structure based on the adjacency matrix and its corresponding spatial omics data, and further improving the representation ability of the omics data.

[0025] Furthermore, the process of completing the integration of cell spatial multi-omics data includes:

[0026] During the process of fusing cell spatial multi-omics data based on the constructed graph structure by the graph attention encoder, an independent graph attention encoder is set for the graph structure of each omics data; the graph attention encoder includes a graph attention network with an attention mechanism and a multi-layer perceptron, and the decoder includes a multi-layer perceptron and a graph attention network;

[0027] The graph attention encoder encodes the graph structure through the graph attention network with an attention mechanism and a multi-layer perceptron to extract the low-dimensional representation corresponding to the omics data in the graph structure, and fuses the obtained low-dimensional representations;

[0028] The decoder reconstructs the data from the fused low-dimensional representation through a multi-layer perceptron and a graph attention network, and integrates the cell spatial multi-omics data according to the reconstructed data, and outputs the integration result.

[0029] Furthermore, during the process of the graph attention encoder extracting the low-dimensional representation corresponding to the omics data in the graph structure, the graph attention network is used to map the omics data and the adjacency matrix to the low-dimensional space, and a multi-layer perceptron is set for low-dimensional representation learning to extract the low-dimensional representation z corresponding to the omics data in the graph structure m 。

[0030] Furthermore, the average pooling method is used to fuse the obtained low-dimensional representation z m to obtain the joint representation z joint , and the expression is:

[0031]

[0032] where represents the set of spatial omics data.

[0033] Furthermore, the process of setting the graph attention and the multi-layer perceptron for low-dimensional representation learning includes:

[0034] For any omics, the latent variable at the 0th layer is the original data, and the expression is:

[0035]

[0036] The latent variable at the k∈{1,...,K - 1} layer is represented as:

[0037]

[0038] The last layer does not use the attention layer, and the expression is:

[0039]

[0040] In the kth layer, the attention intensity between point n and point j is:

[0041]

[0042] The attention intensity vector is normalized to obtain the attention of the final graph attention network, and the expression is:

[0043]

[0044] The graph attention network adopts attention Extract the low-dimensional representation z corresponding to the omics data in the graph structure m ;

[0045] Among them, W k represents a trainable weight matrix, σ represents a non-linear activation function, and S i represents the neighbor set of point i, and represent trainable vectors.

[0046] In the above technical solution, setting an independent graph attention encoder for encoding the graph structure of each omics data can effectively consider different spatial information for different omics data. The graph attention network is used to map the omics data and the adjacency matrix to a low-dimensional space, and a multi-layer perceptron is set for low-dimensional representation learning, so as to improve the smoothness of the final result in space, the accuracy of omics alignment, and the accuracy of spatial analysis, so as to extract the low-dimensional representations corresponding to different omics data in the graph structure.

[0047] Furthermore, the process of constructing a multi-variable loss function to iteratively optimize the process of integrating cell spatial omics data in the spatial omics data model includes:

[0048] Set several rounds of training to iteratively optimize the parameters of the spatial omics data model, specifically:

[0049] During the training process, set the mean squared error loss function to calculate the loss during the data reconstruction process of the fused low-dimensional representation, and the expression is:

[0050]

[0051] Set the low-dimensional representation consistency constraint loss function to constrain the consistency between the low-dimensional representations of omics data, and the expression is:

[0052]

[0053] For the low-dimensional representations of any two slices and Set the slice alignment loss function to constrain the consistency of the expressions of different slices, and the expression is:

[0054]

[0055] Among them, represents the data of the nth point on the s-th slice in the m-th omics data, represents the data of the nth point on the s-th slice in the m-th omics data obtained by reconstruction.

[0056] Furthermore, a total loss function is constructed according to the mean square error loss function, the low-dimensional representation consistency constraint loss function, and the slice alignment loss function, and the expression is:

[0057] l total = λ1l recon + λ2l consist + λ3l mmd

[0058] Among them, λ1, λ2, and λ3 represent preset hyperparameters used to balance various losses;

[0059] Furthermore, the optimizer Adam is set to optimize the parameters of the spatial multi-omics data model during the training process, and the weight decay is set to 0.0005.

[0060] In the above technical solution, the set mean square error loss function can effectively reduce the loss during the low-dimensional representation reconstruction process, thereby improving the accuracy of spatial analysis. The set low-dimensional representation consistency constraint loss function can ensure the consistency between the low-dimensional representations of omics data. The set slice alignment loss function can improve the accuracy of omics alignment. Constructing a total loss function according to the mean square error loss function, the low-dimensional representation consistency constraint loss function, and the slice alignment loss function, and setting corresponding hyperparameters for each loss can effectively balance the contributions of different losses. On the basis of the multi-loss function, setting the optimizer Adam to optimize the parameters of the model during the training process can overall improve the smoothness of the final result in space, the accuracy of omics alignment, and the accuracy of spatial analysis, thereby realizing the integration of spatial mosaic omics data (i.e., spatial multi-omics data).

[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0062] The present invention proposes a method for integrating spatial multi-omics data based on graph attention and a multi-loss function. By constructing a graph structure according to spatial multi-omics data and spatial position coordinates through a preset spatial multi-omics data model, the representation ability of omics data can be improved. A graph attention encoder is used to fuse cell spatial multi-omics data based on the constructed graph structure, and a decoder is used to reconstruct data based on the fused data to complete the integration of cell spatial multi-omics data, which can fully consider spatial information, enabling the model to efficiently master spatial multi-omics data and spatial position coordinate information. In addition, during the iterative optimization process of the cell spatial multi-omics data integration process of the spatial multi-omics data model, the constructed multi-loss function can effectively solve the challenges of batch effect removal and multi-omics data fusion by combining multiple losses with a graph attention network, improving the smoothness of the final result in space, the accuracy of omics alignment, and the accuracy of spatial analysis, thereby realizing the integration of spatial mosaic omics data (i.e., spatial multi-omics data). BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 FIG. is a flowchart of the steps of a method for integrating spatial multi-omics data based on graph attention and a multi-loss function provided in this embodiment;

[0064] Figure 2 FIG. is a schematic diagram of the data integration principle of the spatial multi-omics data model provided in this embodiment;

[0065] Figure 3 FIG. is an effect diagram of the integration of spatial multi-omics data in the experiment provided in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the understanding of the disclosure of the present invention more thorough and comprehensive.

[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0068] Embodiment 1:

[0069] This embodiment provides a method for integrating spatial multi-omics data based on graph attention and a multi-loss function. Refer to Figure 1, the method includes the following steps:

[0070] Step S1: Obtain multiple groups of spatial omics data and their corresponding spatial position coordinates, and input them into a preset spatial omics data model, where the model includes a graph attention encoder and a decoder;

[0071] Step S2: The spatial omics data model constructs a graph structure based on the spatial omics data and spatial position coordinates. The graph attention encoder performs cell spatial omics data fusion based on the constructed graph structure, and the decoder performs data reconstruction based on the fused data to complete cell spatial omics data integration;

[0072] Step S3: Construct a multivariate loss function to iteratively optimize the cell spatial omics data integration process of the spatial omics data model, and output the optimal cell spatial omics data integration result.

[0073] As a preferred embodiment, in step S1, refer to Figure 2 , the spatial omics data model preprocesses the input spatial omics data, and the process includes:

[0074] The spatial omics data includes transcriptomics data and epigenetics data;

[0075] Perform log1p transformation on the transcriptomics data, and perform standardization and dimensionality reduction processing on the transformed transcriptomics data;

[0076] Convert the format of the epigenetics data into peak form, and perform high-variable feature screening and dimensionality reduction processing on the epigenetics data in peak form.

[0077] It can be understood that performing log1p transformation on the transcriptomics data is to select high-variable genes in the data. Performing standardization and dimensionality reduction processing on the transformed transcriptomics data can effectively reduce redundant features in the data. Converting the format of the epigenetics data into peak form and performing high-variable feature screening and dimensionality reduction processing on the epigenetics data in peak form can improve the practicability and reliability of the data and reduce the data processing time.

[0078] In this embodiment, a graph structure is constructed according to spatial omics data and spatial position coordinates through a preset spatial omics data model, which can improve the representation ability of omics data. The graph attention encoder is used to fuse cell spatial omics data based on the constructed graph structure, and the decoder is used to reconstruct data based on the fused data to complete the integration of cell spatial omics data, fully considering spatial information, so that the model can efficiently master spatial omics data and spatial position coordinate information. In addition, during the iterative optimization of the process of integrating cell spatial omics data by the constructed multi-loss function, it can effectively solve the challenges of batch effect removal and multi-omics data fusion through multiple losses combined with the graph attention network, improving the smoothness of the final result in space, the accuracy of omics alignment, and the accuracy of spatial analysis, thus realizing the integration of spatial mosaic omics data (i.e., spatial omics data).

[0079] Embodiment 2:

[0080] This embodiment further elaborates on steps S2 - S3 in Embodiment 1, specifically as follows:

[0081] As a preferred embodiment, in step S2, the process of the spatial omics data model constructing a graph structure according to the spatial omics data X and spatial position coordinates includes:

[0082] According to the spatial position coordinates corresponding to each spatial omics data, the K-nearest neighbor algorithm is used to calculate the corresponding adjacency matrix, and the expression is:

[0083] A s ∈{0,1} N×N

[0084] Construct a graph structure based on the adjacency matrix and its corresponding spatial omics data The expression is:

[0085]

[0086] where s represents the s-th slice in the spatial omics data, N represents a natural number, represents the set of m-omics data of all points on the s-th slice.

[0087] It can be understood that using the K-nearest neighbor algorithm can effectively calculate the corresponding adjacency matrix according to the spatial position coordinates corresponding to each spatial omics data, and then construct a graph structure based on the adjacency matrix and its corresponding spatial omics data, thereby improving the representation ability of omics data.

[0088] As a preferred embodiment, in step S2, the process of completing the integration of cell spatial omics data includes:

[0089] During the process of the graph attention encoder performing cell spatial multi-omics data fusion based on the constructed graph structure, an independent graph attention encoder is set for the graph structure of each omics data; the graph attention encoder includes a graph attention network with an attention mechanism and a multi-layer perceptron, and the decoder includes a multi-layer perceptron and a graph attention network;

[0090] The graph attention encoder performs auto-encoding on the graph structure through the graph attention network with an attention mechanism and a multi-layer perceptron to extract the low-dimensional representation corresponding to the omics data in the graph structure, and fuses the obtained low-dimensional representations;

[0091] The decoder performs data reconstruction on the fused low-dimensional representation through a multi-layer perceptron and a graph attention network, and integrates the cell spatial multi-omics data according to the reconstructed data, and outputs the integration result.

[0092] As a preferred embodiment, during the process of the graph attention encoder extracting the low-dimensional representation corresponding to the omics data in the graph structure, the graph attention network is used to map the omics data and the adjacency matrix to the low-dimensional space, and a multi-layer perceptron is set for low-dimensional representation learning to extract the low-dimensional representation z corresponding to the omics data in the graph structure m 。

[0093] Specifically, the process of setting the graph attention and the multi-layer perceptron for low-dimensional representation learning includes:

[0094] For any omics, the latent variable of the 0th layer is the original data, and the expression is:

[0095]

[0096] The latent variable of the k∈{1,...,K - 1} layer is expressed as:

[0097]

[0098] The last layer does not use the attention layer, and the expression is:

[0099]

[0100] In the kth layer, the attention intensity between point n and point j is:

[0101]

[0102] The attention intensity vector is normalized to obtain the final attention of the graph attention network, and the expression is:

[0103]

[0104] The graph attention network uses the attention Extract the low-dimensional representation z corresponding to the omics data in the graph structure m ;

[0105] Among them, W k represents a trainable weight matrix, σ represents a non-linear activation function, S i represents the set of neighbors of point i, and represent trainable vectors.

[0106] Specifically, in the process of fusing the obtained low-dimensional representation z m , the average pooling method is used to fuse the obtained low-dimensional representation z m to obtain a joint representation z joint , and the expression is:

[0107]

[0108] Among them, represents the set of spatial omics data.

[0109] It can be understood that setting independent graph attention encoders for the graph structures of each omics data can effectively consider different spatial information for different omics data. The graph attention network is used to map the omics data and the adjacency matrix to a low-dimensional space, and a multi-layer perceptron is set for low-dimensional representation learning, so as to improve the spatial smoothness of the final result, the accuracy of omics alignment, and the accuracy of spatial analysis, in order to extract the low-dimensional representations corresponding to different omics data in the graph structure.

[0110] As a preferred embodiment, in step S3, the process of constructing a multi-variate loss function to iteratively optimize the cell spatial omics data integration process of the spatial multi-omics data model includes:

[0111] Set several rounds of training to iteratively optimize the parameters of the spatial multi-omics data model, specifically:

[0112] During the training process, after obtaining the low-dimensional representation z m , set the mean square error loss function to calculate the loss during the data reconstruction process of the fused low-dimensional representation, and the expression:

[0113]

[0114] In order to ensure the consistency between the low-dimensional representations of the omics data, set the low-dimensional representation consistency constraint loss function to constrain the consistency between the low-dimensional representations of the omics data, and the expression is:

[0115]

[0116] For the low-dimensional representations of any two slices and Set a slice alignment loss function to constrain the consistency of the expressions of different slices. The expression is as follows:

[0117]

[0118] Among them, represents the data of the nth point on the sth slice in the mth omics data, represents the data of the nth point on the sth slice in the mth omics data obtained by reconstruction.

[0119] Furthermore, construct a total loss function according to the mean square error loss function, the low-dimensional representation consistency constraint loss function, and the slice alignment loss function. The expression is as follows:

[0120] l total = λ1l recon + λ2l consist + λ3l mmd

[0121] Among them, λ1, λ2, and λ3 represent preset hyperparameters used to balance various losses. Set λ1 = 1, λ2 = 100, and λ3 = 1000;

[0122] As a preferred embodiment, set the number of rounds of training the spatial multi-omics data model to 1500 rounds. Set the optimizer Adam to optimize the parameters of the spatial multi-omics data model during the integration training process of the spatial multi-combination data, and set the weight decay to 0.0005.

[0123] In this embodiment, the set mean square error loss function can effectively reduce the loss during the low-dimensional representation reconstruction process, thereby improving the accuracy of spatial analysis. The set low-dimensional representation consistency constraint loss function can ensure the consistency between the low-dimensional representations of the omics data. The set slice alignment loss function can improve the accuracy of omics alignment. Constructing a total loss function according to the mean square error loss function, the low-dimensional representation consistency constraint loss function, and the slice alignment loss function, and setting corresponding hyperparameters for each loss can effectively balance the contributions between different losses. On the basis of the multi-loss function, setting the optimizer Adam to optimize the parameters of the model during the training process can overall improve the smoothness of the final result in space, the accuracy of omics alignment, and the accuracy of spatial analysis, thereby realizing the integration of spatial mosaic omics data (i.e., spatial multi-omics data).

[0124] Example Three:

[0125] Based on the methods described in Example One and Example Two, the corresponding experimental data are given as follows:

[0126] Data collection: The experiments of this invention used a mouse brain slice dataset, which consists of four slices and includes transcriptomics data, ATAC dataset, and measurement data of multiple histone modifications (including H3K27ac, H3K27me3, and H3K4me3). The preprocessing of the dataset and the process of constructing the graph were carried out according to the methods described above. In particular, in the k-nearest neighbor algorithm, k = 10 was selected, that is, each data point was connected to its 10 nearest neighbors to construct an adjacency matrix.

[0127] Model construction: In the graph attention encoder of the graph structure, a graph attention network containing one layer of attention mechanism was adopted, and a multi-layer perceptron was combined for feature learning. In the decoder part, one layer of multi-layer perceptron and one layer of graph attention network structure were used for reconstruction.

[0128] Model training: The number of training rounds was set to 1500 rounds. The Adam optimizer was adopted, the learning rate was 0.001, and the weight decay was set to 0.0005. For the hyperparameters, λ1 = 1, λ2 = 100, and λ3 = 1000 were set.

[0129] To verify the effect of the model, clustering analysis was performed on the low-dimensional representation results. For the joint representation, the low-dimensional representations of all slices were visualized after UMAP dimensionality reduction to observe the alignment between slices; at the same time, the Mclust clustering method was used for automatic clustering, and the clustering results were visually displayed together with the spatial information. For the low-dimensional representation of omics, the low-dimensional representations of multiple omics for each slice were subjected to UMAP dimensionality reduction and then visualized to observe their alignment. The experimental results are as Figure 3 shown, further demonstrating the effectiveness of the method of this invention in the integration of spatial omics data.

[0130] Figure 3 In the figure, the first column of graphs is the UMAP visualization result of the joint representation; the second column is the UMAP visualization result of the omics representation of each slice: from top to bottom are the UMAP visualization results of different omics of slices 1-4; the third column is the visualization result of the spatial distribution of the clustering results of each slice, from top to bottom are slices 1-4.

[0131] The above are only the embodiments of this invention, and do not limit the patent scope of this invention accordingly. Any equivalent structural or equivalent process transformation made using the content of the specification and drawings of this invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this invention.

Claims

1. A spatial multi-omics data integration method based on graph attention and multi-loss functions, characterized in that The method includes the following steps: Obtain multiple sets of spatial omics data and their corresponding spatial position coordinates, and input them into a preset spatial omics data model, where the model includes a graph attention encoder and a decoder; The spatial omics data model constructs a graph structure based on the spatial omics data and spatial position coordinates. The graph attention encoder performs cell spatial omics data fusion based on the constructed graph structure, and the decoder performs data reconstruction based on the fused data to complete cell spatial omics data integration; Construct a multi - variable loss function to iteratively optimize the cell spatial omics data integration process of the spatial omics data model, and output the optimal cell spatial omics data integration result.

2. The spatial omics data integration method based on graph attention and multi - loss function according to claim 1, wherein The spatial omics data model pre - processes the input spatial omics data. The process includes: The spatial omics data includes transcriptomics data and epigenetics data; Perform log1p transformation on the transcriptomics data, and perform normalization and dimensionality reduction on the transformed transcriptomics data; Convert the format of the epigenetics data into peak form, and perform high - variable feature screening and dimensionality reduction on the epigenetics data in peak form.

3. The spatial multi-omics data integration method based on graph attention and multi-loss function according to claim 2, wherein, The process by which the spatial omics data model constructs a graph structure based on the spatial omics data and spatial position coordinates includes: According to the spatial position coordinates corresponding to each spatial omics data, use the K - nearest neighbor algorithm to calculate the corresponding adjacency matrix, and the expression is: A s ∈{0,1} N×N Constructing a graph structure based on the adjacency matrix and its corresponding spatial omics data The expression is as follows: Among them, s represents the s-th slice in the spatial omics data, N represents natural numbers, represents the set of m-omics data of all points on the s-th slice.

4. The spatial omics data integration method based on graph attention and multi - loss function according to claim 3, characterized in that The process of completing cell spatial omics data integration includes: During the process of the graph attention encoder performing cell spatial omics data fusion based on the constructed graph structure, set an independent graph attention encoder for the graph structure of each omics data; the graph attention encoder includes a graph attention network with an attention mechanism and a multi - layer perceptron, and the decoder includes a multi - layer perceptron; The graph attention encoder encodes the graph structure through the graph attention network with an attention mechanism and a multi - layer perceptron to extract the low - dimensional representation corresponding to the omics data in the graph structure, and fuses the obtained low - dimensional representations; The decoder reconstructs the data through the multi - layer perceptron for the fused low - dimensional representation, and performs cell spatial omics data integration according to the reconstructed data, and outputs the integration result.

5. The spatial omics data integration method based on graph attention and multi - loss function according to claim 4, wherein In the process of the graph attention encoder extracting the low-dimensional representation corresponding to the omics data in the graph structure, the graph attention network is used to map the omics data and the adjacency matrix into the low-dimensional space, and combined with a multi-layer perceptron for low-dimensional representation learning to extract the low-dimensional representation z corresponding to the omics data in the graph structure m 。 6. The spatial multi-omics data integration method based on graph attention and multi-loss function according to claim 5, characterized in that The obtained low-dimensional representation z is fused by using the average pooling method m to obtain the joint representation z joint , and the expression is as follows: Among them, represents a set of spatial omics data.

7. The spatial multi-omics data integration method based on graph attention and multi-loss function according to claim 5, wherein The process of setting the graph attention and multi - layer perceptron for low - dimensional representation learning includes: For any omics, there is a latent variable at the 0th layer which is the raw data, and the expression is: The latent variable representation of the k - th layer, where k ∈ {1,..., K - 1}, is: The last layer does not use an attention layer, and the expression is: In the k - th layer, the attention intensity between point n and point j is: Perform normalization on the attention intensity vector to obtain the attention of the final graph attention network, and the expression is: The graph attention network adopts attention to extract the low-dimensional representation z corresponding to the omics data in the graph structure m ; where, W k represents a trainable weight matrix, σ represents a non-linear activation function, S i represents the set of neighbors of point i, and represents a trainable vector.

8. The spatial multi-omics data integration method based on graph attention and multi-loss function according to claim 7, characterized in that The process of constructing a multi - variable loss function to iteratively optimize the cell spatial omics data integration process of the spatial omics data model includes: Set several rounds of training to iteratively optimize the parameters of the spatial omics data model. Specifically: During the training process, set the mean squared error loss function to calculate the loss during the data reconstruction process for the fused low - dimensional representation, and the expression is: Set the low - dimensional representation consistency constraint loss function to constrain the consistency between the low - dimensional representations of the omics data, and the expression is: Low-dimensional representations of any two slices and Set a slice alignment loss function to constrain the consistency of the expressions of different slices. The expression is as follows: Among them, represents the data of the nth point on the s-th slice of the m-th omics data, represents the data of the nth point on the s-th slice of the m-th omics data obtained by reconstruction, and k(·,·) represents the Gaussian kernel function.

9. The method for integrating spatial omics data based on graph attention and multi - loss function according to claim 8, characterized in that, Construct a total loss function based on the mean squared error loss function, the low-dimensional representation consistency constraint loss function, and the slice alignment loss function. The expression is as follows: l total = λ1l recon + λ2l consist + λ3l mmd Among them, λ1, λ2, and λ3 represent preset hyperparameters used to balance various losses.

10. The spatial multi-omics data integration method based on graph attention and multi-loss function according to claim 9, wherein Set the Adam optimizer to optimize the parameters of the spatial omics data model during training, and set the weight decay to 0.0005.

Citation Information

Cited By

  • Biological analysis method and device based on spatial multi-omics data, equipment and medium

    CN120432015A

  • Biological analysis methods, devices, equipment and media based on spatial multi-omics data

    CN120432015B

  • Logging lithology identification method based on physical information constraint

    CN121115166A