Method and device for carrying out multi-scale modeling on spatial transcriptomics data

By using multi-scale information extraction and two-dimensional rotation and translation-variable attention network training, the difficulties in modeling spatial transcriptomics data in existing technologies have been solved, realizing a high-quality data representation learning and analysis tool, and improving the efficiency and accuracy of spatial transcriptomics data analysis.

CN120895101APending Publication Date: 2025-11-04TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510863188.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies struggle to fully utilize unsupervised spatial transcriptomics data to capture multi-scale information for modeling, resulting in insufficient efficiency and accuracy in spatial transcriptomics data analysis.

Method used

A multi-scale information extraction method is adopted, including macroscopic, microscopic and gene-scale information extraction. A two-dimensional rotation and translation isovariant attention network is used to model spatial transcriptomics data. The cell representation is optimized by training a pre-trained cell encoder and a two-dimensional rotation and translation isovariant attention network, combined with masked cell modeling and pairwise distance recovery tasks.

Benefits of technology

It enables high-quality, general representation learning for spatial transcriptomics data, provides effective artificial intelligence tools, and improves the efficiency and accuracy of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895101A_ABST
    Figure CN120895101A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for performing multi-scale modeling on spatial transcriptomics data, and relates to the technical field of biological information.The method comprises the steps that a to-be-processed spatial transcriptomics ST slice is obtained, multi-scale information extraction is performed on the ST slice, and a plurality of to-be-processed sub slices are obtained; sequentially inputting the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices into the two-dimensional rotation translation equivariant attention network to obtain a cell representation of each cell in each slice; and obtaining the cell expression of each cell in the ST slice based on the cell expression of each cell in each slice. According to the method and device for carrying out multi-scale modeling on the spatial transcriptomics data, by capturing and fusing complex multi-scale information in spatial transcriptomics, high-quality general representation learning on the spatial transcriptomics data is realized, and an effective artificial intelligence tool is provided for spatial transcriptomics data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and in particular, to a method and device for multi-scale modeling of spatial transcriptomics data. BACKGROUND

[0002] Spatial transcriptomics technology provides biologists with rich single-cell biology insights by preserving the spatial context of cells. Building a basic model of spatial transcriptomics can significantly enhance the analysis of large and complex data sources and open up new perspectives on the complexity of biological tissues.

[0003] However, it is difficult to model spatial transcriptomics data because it is necessary to extract multi-scale information from tissue sections containing a large number of cells. This process requires the integration of macro-scale tissue morphology, micro-scale cell microenvironment, and gene-scale gene expression information. Related technologies do not have a model that can fully utilize a large amount of unsupervised spatial transcriptomics data and effectively capture multi-scale information for modeling. SUMMARY

[0004] The purpose of the present application is to provide a method and device for multi-scale modeling of spatial transcriptomics data, which captures and fuses the complex multi-scale information in spatial transcriptomics, realizes high-quality general representation learning of spatial transcriptomics data, and provides an effective artificial intelligence tool for spatial transcriptomics data analysis.

[0005] The present application provides a method for multi-scale modeling of spatial transcriptomics data, comprising: obtaining a spatial transcriptomics ST section to be processed, and performing multi-scale information extraction on the ST section to obtain a plurality of to-be-processed sub-sections; inputting to-be-processed information corresponding to each to-be-processed sub-section in the plurality of to-be-processed sub-sections into a two-dimensional rotation translation equivariant attention network in turn to obtain cell representation of each cell in each section; obtaining cell representation of each cell in the ST section based on the cell representation of each cell in each section; wherein the multi-scale information extraction comprises macro-scale information extraction, micro-scale information extraction, and gene-scale information extraction; the to-be-processed information contained in the to-be-processed sub-section includes cell feature information for representing a spatial distance matrix of cell spacing.

[0006] Optionally, the multi-scale information extraction on the ST slice to obtain a plurality of to-be-processed sub-slices comprises: dividing the cells contained in the ST slice based on the two-dimensional position coordinate information of each cell in the ST slice to obtain a plurality of sub-slices; each sub-slice contains a subset of the cell set contained in the ST slice; encoding the gene expression information of each cell in the ST slice by using a pre-trained cell encoder to obtain cell feature information of each cell, and clustering and dividing the cells in the ST slice based on the cell feature information of each cell and the two-dimensional position coordinate information of each cell to obtain a plurality of clusters; generating a virtual cell corresponding to each cluster based on the cell feature mean and the two-dimensional position coordinate mean of the cells contained in each cluster to obtain a plurality of virtual cells, and merging the plurality of virtual cells into each sub-slice respectively to obtain a plurality of to-be-processed sub-slices.

[0007] Optionally, before the cell representation of each cell in each slice is obtained by inputting the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices into the two-dimensional rotation translation equivariant attention network in turn, the method further comprises: obtaining a plurality of ST slice samples; each ST slice sample contains: cell feature information of each cell, two-dimensional position coordinate information of each cell, and gene expression information of each cell; clustering each ST slice sample in the plurality of ST slice samples respectively, and obtaining a plurality of sub-slice samples corresponding to each ST slice sample according to the clustering result; using each sub-slice sample as a training sample to perform different stage training on the pre-trained cell encoder and the two-dimensional rotation translation equivariant attention network, to obtain a trained cell encoder and a trained two-dimensional rotation translation equivariant attention network; wherein, the different stage training comprises: first stage training and second stage training; the first stage training comprises: freezing the parameters of the cell encoder, and training the two-dimensional rotation translation equivariant attention network using the cell feature information contained in the sub-slice sample until the two-dimensional rotation translation equivariant attention network meets a first preset convergence condition; the second stage training comprises: randomly selecting part of the cells to calculate the cell feature information by the unfreezing part of the parameter cell encoder in each forward propagation, so as to synchronously update the cell encoder and the two-dimensional rotation translation equivariant attention network by back propagation; the first stage training does not use an iterative strategy, and the second stage training optimizes the clustering effect by using an iterative strategy.

[0008] Optionally, the training task of the second stage training comprises a masked cell modeling task; the masked cell modeling task comprises: randomly masking target feature information in a first preset proportion of to-be-input cell feature information to obtain masked to-be-input cell feature information; inputting the masked to-be-input cell feature information into the two-dimensional rotation translation equivariant attention network, and predicting the masked target feature information based on cell representation output by the two-dimensional rotation translation equivariant attention network through a regression head; wherein the masked target feature information is feature information of a non-virtual cell; and the masked cell modeling task uses mean square error as a loss function.

[0009] Optionally, the loss function of the masked cell modeling task is: wherein, is a loss function value, is a set of masked cells, is feature information of a jth cell in the masked cells, is feature information of the jth cell predicted by the regression head. j

[0010] Optionally, the training task of the second stage training comprises a pair distance recovery task; the pair distance recovery task comprises: randomly selecting target cells in a second preset proportion of sub-slices, and adding noise interference to two-dimensional position coordinate information of the target cells to obtain an adjusted distance matrix; inputting the adjusted distance matrix into the two-dimensional rotation translation equivariant attention network, and reconstructing the distance matrix disturbed by noise based on pair representation output by the two-dimensional rotation translation equivariant attention network through a regression head; wherein the pair distance recovery task uses mean square error as a loss function.

[0011] Optionally, the loss function of the pair distance recovery task is: wherein, is a loss function value, is a set of elements disturbed by noise, is a distance between a cell j and a cell k after being disturbed by noise, is a distance between the cell j and the cell k predicted by the regression head.

[0012] The application further provides a device for multi-scale modeling of spatial transcriptomic data, comprising: ​The acquisition module is configured to acquire a spatial transcriptomics (ST) slice to be processed. The multi-scale information extraction module is configured to perform multi-scale information extraction on the ST slice to obtain a plurality of to-be-processed sub-slices. The cell representation generation module is configured to input to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices into a two-dimensional rotation translation equiform attention network in sequence to obtain cell representation of each cell in each slice. The cell representation generation module is further configured to obtain cell representation of each cell in the ST slice based on the cell representation of each cell in each slice. The multi-scale information extraction includes macro-scale information extraction, micro-scale information extraction, and gene-scale information extraction. The to-be-processed information included in the to-be-processed sub-slice includes cell feature information used to represent a spatial distance matrix of cell spacing.

[0013] Optionally, the multi-scale information extraction module is specifically configured to divide cells included in the ST slice based on two-dimensional position coordinate information of each cell in the ST slice to obtain a plurality of sub-slices. A cell set included in each sub-slice is a subset of a cell set included in the ST slice. The multi-scale information extraction module is specifically further configured to encode gene expression information of each cell in the ST slice by using a pre-trained cell encoder to obtain cell feature information of each cell, and perform clustering and division on cells in the ST slice based on the cell feature information of each cell and the two-dimensional position coordinate information of each cell to obtain a plurality of clusters. The multi-scale information extraction module is specifically further configured to generate a virtual cell corresponding to each cluster based on a cell feature mean value and a two-dimensional position coordinate mean value of cells included in each cluster to obtain a plurality of virtual cells, and merge the plurality of virtual cells into each sub-slice respectively to obtain a plurality of to-be-processed sub-slices.

[0014] Optionally, the apparatus further comprises a sample generation module and a training module; the acquisition module is further configured to acquire a plurality of ST slice samples; each ST slice sample comprises cell feature information of each cell, two-dimensional position coordinate information of each cell, and gene expression information of each cell; the sample generation module is configured to cluster each ST slice sample in the plurality of ST slice samples respectively, and obtain a plurality of sub-slice samples corresponding to each ST slice sample according to a clustering result; the training module is configured to train a pre-trained cell encoder and a two-dimensional rotation translation equivariant attention network in different stages by taking each sub-slice sample as a training sample, to obtain a trained cell encoder and a trained two-dimensional rotation translation equivariant attention network; wherein the training in different stages comprises first stage training and second stage training; the first stage training comprises freezing parameters of the cell encoder, and training the two-dimensional rotation translation equivariant attention network using the cell feature information contained in the sub-slice sample until the two-dimensional rotation translation equivariant attention network meets a first preset convergence condition; the second stage training comprises randomly selecting part of cells to calculate cell feature information by the cell encoder with unfrozen part of parameters in each forward propagation, so as to synchronously update the cell encoder and the two-dimensional rotation translation equivariant attention network through back propagation; the first stage training does not use an iterative strategy, and the second stage training optimizes clustering effect by using an iterative strategy.

[0015] Optionally, the training task of the second stage training comprises a masked cell modeling task; the training module is specifically configured to randomly mask target feature information of a first preset proportion in input cell feature information to obtain masked input cell feature information; the training module is specifically further configured to input the masked input cell feature information into the two-dimensional rotation translation equivariant attention network, and predict the masked target feature information based on cell representation output by the two-dimensional rotation translation equivariant attention network through a regression head; wherein the masked target feature information is feature information of a non-virtual cell; the masked cell modeling task uses mean square error as a loss function.

[0016] Optionally, the training task of the second stage training comprises a pair distance recovery task; the training module is specifically configured to randomly select target cells of a second preset proportion in a sub-slice, and add noise interference to two-dimensional position coordinate information of the target cells to obtain an adjusted distance matrix; the training module is specifically further configured to input the adjusted distance matrix into the two-dimensional rotation translation equivariant attention network, and reconstruct the distance matrix disturbed by noise based on pair expression output by the two-dimensional rotation translation equivariant attention network through a regression head; wherein the pair distance recovery task uses mean square error as a loss function.

[0017] The application also provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the method for multi-scale modeling of spatial transcriptomic data according to any one of the above.

[0018] The application also provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the method for multi-scale modeling of spatial transcriptomic data according to any one of the above when executing the program.

[0019] The application also provides a computer-readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the method for multi-scale modeling of spatial transcriptomic data according to any one of the above.

[0020] The application provides a method and device for multi-scale modeling of spatial transcriptomic data. Firstly, a spatial transcriptomic ST slice to be processed is obtained, and multi-scale information extraction is performed on the ST slice to obtain a plurality of to-be-processed sub-slices. Then, the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices is sequentially input into a two-dimensional rotation translation equivariant attention network to obtain a cell representation of each cell in each slice. Based on the cell representation of each cell in each slice, a cell representation of each cell in the ST slice is obtained. The multi-scale information extraction includes macro-scale information extraction, micro-scale information extraction, and gene-scale information extraction. The to-be-processed information included in the to-be-processed sub-slice includes cell feature information for representing a spatial distance matrix of cell spacing. In this way, by capturing and fusing the complex multi-scale information in spatial transcriptomics, high-quality general representation learning of spatial transcriptomic data is achieved, and an effective artificial intelligence tool is provided for spatial transcriptomic data analysis. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0022] Figure 1 is one of the flowcharts of the method for multi-scale modeling of spatial transcriptomic data provided by the present application; Figure 2 is another flowchart of the method for multi-scale modeling of spatial transcriptomic data provided by the present application; Figure 3Figure 3 is a schematic diagram of a third process of a method for multi-scale modeling of spatial transcriptomics data provided in the present application; Figure 4 Figure 4 is a structural schematic diagram of an apparatus for multi-scale modeling of spatial transcriptomics data provided in the present application; Figure 5 Figure 5 is a structural schematic diagram of an electronic device provided in the present application. DETAILED DESCRIPTION

[0023] For the purposes of the present application, the technical solutions and advantages thereof will be more clearly understood from the following description of the drawings in the present application, which will clearly and completely describe the technical solutions in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.

[0024] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a class, and are not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally indicates that the objects before and after are in an "or" relationship.

[0025] In the related art, there is no model that can fully utilize a large amount of unsupervised spatial transcriptomics data and effectively capture multi-scale information. Although some research work using machine learning to analyze spatial transcriptomics has achieved remarkable results, most of the research is designed only for specific analysis tasks, and does not have scalability in data quantity and parameter quantity. The existing transcriptomics basic model only focuses on scRNA-seq data and cannot effectively utilize complex spatial location information, and the ability to mine and aggregate multi-scale information is insufficient.

[0026] In view of the above technical problems existing in the related art, the embodiments of the present application provide a method for multi-scale modeling of spatial transcriptomics data, which captures and fuses complex multi-scale information in spatial transcriptomics, realizes high-quality general representation learning of spatial transcriptomics data, and provides an effective artificial intelligence tool for spatial transcriptomics data analysis.

[0027] The method for multi-scale modeling of spatial transcriptomics data provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0028] like Figure 1 As shown in the embodiments of this application, a method for multi-scale modeling of spatial transcriptomics data is provided, which may include the following steps 101 to 103: Step 101: Obtain the spatial transcriptomics ST slices to be processed, and perform multi-scale information extraction on the ST slices to obtain multiple sub-slices to be processed.

[0029] The multi-scale information extraction includes: macro-scale information extraction, micro-scale information extraction, and gene-scale information extraction; the information to be processed contained in the sub-slice includes: cell feature information, which is used to characterize the spatial distance matrix between each cell.

[0030] For example, ST data is obtained in the form of slices, which can be described in computation as a two-dimensional point cloud data structure with high-dimensional features. The user-provided ST slices obtained above... S A slice contains multiple cells, along with gene expression information for each cell. E (Gene Expression), two-dimensional position coordinates of each cell P ,Right now S ={ E , P}

[0031] For example, after obtaining the ST slice, it is necessary to extract multi-scale information, including: macro-scale tissue morphology information extraction, micro-scale intercellular communication information extraction, and gene-scale single-cell gene expression information extraction.

[0032] Specifically, step 101 above may also include steps 101a1 to 101a3: Step 101a1: Based on the two-dimensional position coordinate information of each cell in the ST slice, the cells contained in the ST slice are divided to obtain multiple sub-slices.

[0033] Each sub-slice contains a subset of the cell set contained in the ST slice.

[0034] Step 101a2: Encode the gene expression information of each cell in the ST slice using a pre-trained cell encoder to obtain the cell feature information of each cell, and cluster the cells in the ST slice based on the cell feature information and the two-dimensional position coordinate information of each cell to obtain multiple clusters.

[0035] Step 101a3, generating a virtual cell corresponding to each cluster based on the cell feature mean value and two-dimensional position coordinate mean value of the cells contained in each cluster, obtaining a plurality of virtual cells, and merging the plurality of virtual cells into each sub-slice respectively to obtain a plurality of to-be-processed sub-slices.

[0036] Exemplarily, after obtaining the ST slice input by the user, the gene expression information of each cell can be encoded into cell feature information by using a cell encoder F (that is, S={E,F, P} is obtained), and the two-dimensional position coordinate information P The ST slice is divided into a plurality of sub-slices. At the same time, Leiden algorithm is used to cluster all cells according to the cell feature information F and the two-dimensional position coordinate information P of the cells, and each cluster is summarized into a virtual cell. These virtual cells retain the main morphology and partition of the slice, serving as a compression of macro-scale information. Finally, the virtual cells are merged into each sub-slice, so that the model can learn micro-scale information while maintaining the ability to perceive macro-structure organization.

[0037] Exemplarily, by the above steps, the ST slice can be divided into a plurality of sub-slices, and the cell feature information corresponding to each sub-slice and the spatial distance matrix for representing the distance between each cell are respectively input into the two-dimensional rotation translation equivariant attention network to obtain the cell representation of each cell in each slice.

[0038] Step 102, inputting the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices into the two-dimensional rotation translation equivariant attention network in turn to obtain the cell representation of each cell in each slice.

[0039] Exemplarily, the input of the two-dimensional rotation translation equivariant attention network is the cell feature information F (Cell Embedding) of the cells contained in the sub-slice and the spatial distance matrix D (Spatial Distance Matrix) for representing the distance between each cell, and the output is the cell representation Y (Cell Representation) of each cell in the sub-slice.

[0040] Step 103, obtaining the cell representation of each cell in the ST slice based on the cell representation of each cell in each slice.

[0041] Exemplarily, after obtaining the cell representation of each cell in each slice, the cell representation of each cell in the ST slice can be obtained, and the multi-scale modeling of the spatial transcriptomic data is completed.

[0042] For example, as shown in the following figure, after obtaining the ST slice input by the user, the ST slice is divided into a plurality of sub-slices through multi-scale information extraction, and then each sub-slice is processed respectively, that is, the cell feature information of each cell in each sub-slice and the spatial distance matrix are input into the two-dimensional rotation translation equivariant attention network respectively to obtain the cell representation of each cell in each sub-slice. Finally, after the cell representation of each cell in each sub-slice is summarized, the cell representation of each cell in the ST slice is obtained. Figure 2

[0043] Optionally, in the embodiment of the present application, the above cell encoder and two-dimensional rotation translation equivariant attention network (SE(2) Transformer) can be trained in the following manner.

[0044] Exemplarily, before the above step 102, the method for multi-scale modeling of spatial transcriptomic data provided by the embodiment of the present application can further include the following steps 104 to 106: Step 104, a plurality of ST slice samples are obtained. Each ST slice sample includes cell feature information of each cell, two-dimensional position coordinate information of each cell, and gene expression information of each cell.

[0045] It should be noted that the ST slice sample containing N cells in the embodiment of the present application can be represented by the following formula one: (Formula one) Wherein, is an n-dimensional gene expression matrix, is used to represent the two-dimensional position coordinates.

[0046] Exemplarily, after multi-scale information extraction is performed on each ST slice sample, a group of sub-slices collecting multi-scale information can be obtained, which is defined as: Wherein, is used as the model input of the next process.

[0047] ​It can be understood that, given the high dimensionality and sparsity of gene expression profiles, we need to obtain their reduced dimension expression. Due to technical limitations, spatial transcriptomic data often exhibits defects such as limited gene coverage and a high proportion of "missing zero values". Therefore, compared with single-cell RNA sequencing data, the quality of spatial transcriptomic data is much lower. Training single-cell RNA sequencing-based models step by step on spatial transcriptomic data helps to transfer the knowledge learned from single-cell RNA sequencing data to spatial transcriptomic data with different distributions, thereby achieving higher quality cell feature information. In the embodiments of the present application, Geneformer is used to initialize our cell encoder, and Geneformer is one of the most advanced single-cell base models based on Transformer, which realizes high-quality representation by encoding the gene sequence arranged in descending order of relative expression sequence of cells.

[0048] Exemplarily, in the embodiments of the present application, a part of the ST slice samples can be used to train the cell encoder, such as Figure 3 As shown in (a) of FIG. 1, the gene expression information of the cells in the ST slice samples is used to train the cell encoder, and domain adaptation is realized. After that, the cell encoder subjected to domain adaptation can be used to process the gene expression information of the cells E , and obtain h dimensional features F , which can be represented by the following formula two: (Formula two) Wherein, After that, the F is added to the as additional information of the cells, and S ( E , F , P ) is obtained.

[0049] Step 105, respectively clustering each ST slice sample in the plurality of ST slice samples, and obtaining a plurality of sub-slice samples corresponding to each ST slice sample according to the clustering result.

[0050] Exemplarily, in the embodiments of the present application, a multi-scale information extraction method is designed, which can effectively integrate information from microscale and macroscale. Specifically, first, each ST slice is divided into a plurality of sub-slices according to the spatial position P (that is, the above two-dimensional position coordinate information), and each sub-slice contains a manageable number of cells (about 1,000). This division realizes the trade-off between maintaining computational efficiency and preserving sufficient local intercellular interactions from experience. Then, in order to maintain the macroscale information, the Leiden algorithm can be used to obtain the cell feature information F and cell positionP For all cells S Clustering is performed, and each cluster is aggregated into a virtual cell, whose feature information and location coordinates are obtained by averaging all cells in that cluster. These virtual cells retain the main morphology and partitioning of the slice, serving as a compression of macro-scale information. We merge virtual cells into each sub-slice, enabling the model to learn micro-scale information while maintaining its ability to perceive macro-structural organization.

[0051] As shown in the following micro- and macro-scale integration algorithm, S =( E , F , P Divide into a set of sub-slices containing virtual cells. These sub-slices encapsulate microscopic-scale information and integrate macroscopic and genetic-scale information through virtual cells and cell feature information, respectively. It is worth noting that in this embodiment, only the feature information of virtual cells is calculated. Instead of gene expression, the average gene expression profile of multiple cells is out of range and cannot be effectively encoded by the cell encoder.

[0052] For example, such as Figure 3 As shown in (b), in the multi-scale ST representation learning process, given the sub-slices obtained from the above steps, a cell representation containing transcriptomics and multi-scale spatial information can be obtained.

[0053] Step 106: Use each sub-slice sample as a training sample to train the pre-trained cell encoder and the two-dimensional rotation and translation equivariant attention network at different stages to obtain the trained cell encoder and the trained two-dimensional rotation and translation equivariant attention network.

[0054] The training at different stages includes: a first stage training and a second stage training; the first stage training includes: freezing the parameters of the cell encoder and training the two-dimensional rotation-translation equivariant attention network using the cell feature information contained in the sub-slice samples until the two-dimensional rotation-translation equivariant attention network satisfies a first preset convergence condition; the second stage training includes: randomly selecting a portion of cells in each forward propagation to calculate cell feature information from the cell encoder with partially unfrozen parameters, so as to synchronously update the cell encoder and the two-dimensional rotation-translation equivariant attention network through backpropagation; the first stage training does not use an iterative strategy, while the second stage training uses an iterative strategy to optimize the clustering effect.

[0055] It should be noted that the target of the two-dimensional rotation translation equivariant attention network is to jointly encode cell feature information F and cell positions P to obtain a spatial-aware representation and capture the interaction between cells. In addition, the output representation should be invariant to the two-dimensional translation and rotation of cell positions P , i.e., SE(2) invariance. To this end, the SE(2) Transformer architecture widely adopted in the embodiments of the present application is used, which is originally an SE(3) Transformer but can be applied to 2D scenes with little architectural modification. This simple and efficient architecture performs well in representing proteins and small molecules.

[0056] Specifically, the distance matrix is used as the input of the position information, where: The distance matrix is first processed by a Gaussian module to obtain an initial pair-wise representation. In each Transformer layer, the pair-wise representation is updated by adding an attention matrix calculated from the cell representation to it. Then, the updated pair-wise representation is used as the actual attention score to update the cell representation. This method of using pair-wise representation as attention bias has been proven to be effective in various important works, such as Graphformer and Alphafold series. In addition, since the interaction between cells has strong locality, the distance information as the input of the pair-wise representation is also suitable for processing biological scenes of spatial transcription data.

[0057] Exemplarily, in the embodiments of the present application, the two-dimensional rotation translation equivariant attention network is trained to encode each sub-slice to obtain a cell representation and a pair-wise representation , which can be represented by the following formula two: (Formula three) During the training process, the cell representation and the pair-wise representation can be used to calculate the loss function.

[0058] Exemplarily, during the domain adaptation process before training, the masking gene modeling objective in Geneformer can be adopted, i.e., predicting the random masking genes arranged in the input gene sequence according to the relative expression level. In addition, inspired by LangCell, the embodiments of the present application also introduce a self-supervised contrastive learning task to enhance the representation quality of the cell encoder.

[0059] Exemplarily, in order to better capture multi-scale information from the ST slice, it is necessary to design a training target to obtain a supervised signal from gene expression and cell coordinates. To this end, the training scheme of the embodiment of the present application includes two pre-training tasks: Masked Cell Modeling and Pairwise Distance Recovery, to capture the gene expression and spatial features of spatial transcriptomics data.

[0060] Specifically, for the masked cell modeling task, the above step 106 can further include the following step 106a1 and step 106a2: Step 106a1, randomly mask target feature information in a first preset proportion of the to-be-input cell feature information to obtain the masked to-be-input cell feature information.

[0061] Step 106a2, input the masked to-be-input cell feature information into the two-dimensional rotation translation equivariant attention network, and predict the masked target feature information through the regression head based on the cell representation output by the two-dimensional rotation translation equivariant attention network.

[0062] Wherein, the masked target feature information is the feature information of a non-virtual cell; the masked cell modeling task uses mean square error as a loss function.

[0063] Exemplarily, in the masked cell modeling (MCM) task, 10% of the gene expression feature information in the sub-slice can be randomly masked , and the output representation of the two-dimensional rotation translation equivariant attention network is used to predict the masked embedding through the regression head. Considering the unfeasibility of using microscopic information to reconstruct macroscopic information, the embodiment of the present application does not mask the virtual cell embedding. The embodiment of the present application can use a mean square error (MSE) loss function. The target is defined as the following formula four: (Formula four) Wherein, is the loss function value, is the set of masked cells, is the feature information of the jth cell in the masked cell, is the feature information of the jth cell predicted by the regression head. j

[0064] Specifically, for the pairwise distance recovery task, the above step 106 can further include the following step 106b1 and step 106b2: ​Step 106b1, randomly select a second preset proportion of target cells in the sub-slice, and add noise interference to the two-dimensional position coordinate information of the target cells to obtain an adjusted distance matrix.

[0065] Step 106b2, input the adjusted distance matrix into the two-dimensional rotation translation equivariant attention network, and reconstruct the distance matrix disturbed by noise based on the pair representation output by the two-dimensional rotation translation equivariant attention network through a regression head.

[0066] Wherein, the pair distance recovery task uses mean square error as a loss function.

[0067] Exemplarily, in the pair distance recovery (PDR), 10% of the cells in the sub-slice can be randomly selected, and Gaussian noise can be added to their 2D coordinates , which modifies the corresponding rows and columns of the distance matrix . Then, the two-dimensional rotation translation equivariant attention network attempts to reconstruct the undisturbed distance matrix by using the pair representation . For the same cells, it is ensured that their embeddings are not masked and their coordinates are disturbed at the same time. In the embodiment of the application, the MSE loss function can be used. The target is defined as the following formula five: (Formula five) Wherein, is the loss function value, is a set of elements disturbed by noise, is the distance between cell j and cell k after being disturbed by noise, is the distance between cell j and cell k predicted by the regression head.

[0068] It should be noted that in the training stage of the model, the cell encoder and the two-dimensional rotation translation equivariant attention network need to be trained at the same time, while in the application stage, the cell encoder is only used to encode the gene expression into cell feature information, and then the cell feature information and the spatial distance matrix are input into the two-dimensional rotation translation equivariant attention network to obtain the cell representation.

[0069] Exemplarily, in the embodiment of the application, some strategies can be used to help stabilize the training and accelerate the convergence. In the early stage of pre-training, since the cell encoder that has undergone domain adaptation has been pre-trained, and the SE(2) Transformer is trained from scratch, the cell encoder can be frozen and the Synchronization as input to SE(2) Transformer. Once the training of SE(2) Transformer is basically converged, we adjust the parameters of cell encoder and SE(2) Transformer. Considering the high computational cost of cell encoder, we only perform the second cell encoding on some randomly sampled cells within the sub-slices, which makes us update the cell encoder and SE(2) Transformer simultaneously by backpropagation, improving the learning ability of the model.

[0070] It should be noted that in the embodiments of the present application, the spatial transcriptome public data of multiple sources is sorted. For sequencing methods that can obtain more public data, considering that the sequencing point diameter of some sequencing methods is much larger than the cell diameter (such as 10x Visium), such data cannot reflect the gene expression spatial distribution at the single cell scale, therefore, we only select sequencing methods with single cell resolution (such as MERFISH) or sequencing methods with a sequencing point diameter close to the cell diameter (such as Stereo-seq). Finally, we selected six sequencing methods and sorted the pre-training corpus SToCorpus-80M containing about 2.5k slices and 80M cells. SToCorpus-80M is the largest spatial transcriptomics pre-training corpus at present (previously, the corpus containing 55M spatial transcriptome cells constructed by Nicheformer).

[0071] For the pre-training setting, first, the cell encoder is fine-tuned for an epoch to adapt to the cells in the spatial transcriptome data. Then we pre-train the complete SToFM model for three epochs. Since the graph encoder is trained from scratch, in the first two epochs, we freeze all the parameters of the cell encoder and do not use the iterative strategy. In the third epoch, when the graph encoder is basically converged, we randomly select some cells in each forward propagation, and the cell embedding is calculated by the unfrozen part of the cell encoder, which makes us update the cell encoder and the graph encoder simultaneously by gradient backpropagation, improving the learning ability of the model. And in the third epoch, we use a one-step iterative strategy to optimize the clustering effect. See the appendix for detailed pre-training hyperparameter settings.

[0072] The method for multi-scale modeling of spatial transcriptomics data provided in the embodiments of the present application first acquires a spatial transcriptomics ST slice to be processed, and multi-scale information extraction is performed on the ST slice to obtain a plurality of to-be-processed sub-slices; then, the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices is sequentially input into a two-dimensional rotation translation equivariant attention network to obtain cell representation of each cell in each slice; finally, based on the cell representation of each cell in each slice, cell representation of each cell in the ST slice is obtained; wherein the multi-scale information extraction includes macro-scale information extraction, micro-scale information extraction and gene-scale information extraction; the to-be-processed information included in the to-be-processed sub-slice includes cell feature information for representing a spatial distance matrix of cell spacing. In this way, by capturing and fusing complex multi-scale information in spatial transcriptomics, high-quality general representation learning of spatial transcriptomics data is achieved, and an effective artificial intelligence tool is provided for spatial transcriptomics data analysis.

[0073] It should be noted that the method for multi-scale modeling of spatial transcriptomics data provided in the embodiments of the present application can be a device for multi-scale modeling of spatial transcriptomics data, or a control module in the device for multi-scale modeling of spatial transcriptomics data for executing the method for multi-scale modeling of spatial transcriptomics data. In the embodiments of the present application, the device for multi-scale modeling of spatial transcriptomics data is taken as an example to illustrate the device for multi-scale modeling of spatial transcriptomics data provided in the embodiments of the present application.

[0074] It should be noted that in the embodiments of the present application, the method for multi-scale modeling of spatial transcriptomics data shown in each of the above method figures is illustratively described by taking one of the figures in the embodiments of the present application as an example. In specific implementation, the method for multi-scale modeling of spatial transcriptomics data shown in each of the above method figures can also be implemented in combination with any other figure that can be combined as illustrated in the above embodiments, which will not be described herein again.

[0075] The device for multi-scale modeling of spatial transcriptomics data provided in the present application will be described below, and the method for multi-scale modeling of spatial transcriptomics data described below can be mutually correspondingly referred to the method for multi-scale modeling of spatial transcriptomics data described above.

[0076] Figure 4 The structure diagram of the device for multi-scale modeling of spatial transcriptomics data provided in the embodiments of the present application is shown in FIG. 1, and specifically includes: Figure 4 ​The acquisition module 401 is configured to acquire a spatial transcriptomics ST slice to be processed. The multi-scale information extraction module 402 is configured to perform multi-scale information extraction on the ST slice to obtain a plurality of to-be-processed sub-slices. The to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices is sequentially input into a two-dimensional rotation translation equivariable attention network to obtain cell representation of each cell in each slice. The cell representation generation module 403 is further configured to obtain cell representation of each cell in the ST slice based on the cell representation of each cell in each slice. The multi-scale information extraction includes macro-scale information extraction, micro-scale information extraction and gene-scale information extraction. The to-be-processed information included in the to-be-processed sub-slice includes cell feature information used to represent a spatial distance matrix of cell spacing.

[0077] Optionally, the multi-scale information extraction module 402 is specifically configured to divide the cells included in the ST slice based on two-dimensional position coordinate information of each cell in the ST slice to obtain a plurality of sub-slices. A cell set included in each sub-slice is a subset of a cell set included in the ST slice. The multi-scale information extraction module 402 is specifically further configured to encode gene expression information of each cell in the ST slice by using a pre-trained cell encoder to obtain cell feature information of each cell, and perform clustering division on the cells in the ST slice based on the cell feature information of each cell and the two-dimensional position coordinate information of each cell to obtain a plurality of clusters. The multi-scale information extraction module 402 is specifically further configured to generate a virtual cell corresponding to each cluster based on a cell feature mean and a two-dimensional position coordinate mean of the cells included in each cluster to obtain a plurality of virtual cells, and merge the plurality of virtual cells into each sub-slice respectively to obtain a plurality of to-be-processed sub-slices.

[0078] Optionally, the apparatus further comprises a sample generation module and a training module; the acquisition module 401 is further configured to acquire a plurality of ST slice samples; each ST slice sample comprises cell feature information of each cell, two-dimensional position coordinate information of each cell, and gene expression information of each cell; the sample generation module is configured to cluster each ST slice sample in the plurality of ST slice samples respectively, and obtain a plurality of sub-slice samples corresponding to each ST slice sample according to a clustering result; the training module is configured to perform different stage training on the pre-trained cell encoder and the two-dimensional rotation translation equivariant attention network by taking each sub-slice sample as a training sample, to obtain a trained cell encoder and a trained two-dimensional rotation translation equivariant attention network; wherein the different stage training comprises first stage training and second stage training; the first stage training comprises freezing parameters of the cell encoder, and training the two-dimensional rotation translation equivariant attention network using the cell feature information contained in the sub-slice sample until the two-dimensional rotation translation equivariant attention network meets a first preset convergence condition; the second stage training comprises randomly selecting part of cells to calculate cell feature information by the unfreezing part of the parameters of the cell encoder in each forward propagation, so as to synchronously update the cell encoder and the two-dimensional rotation translation equivariant attention network through back propagation; the first stage training does not use an iterative strategy, and the second stage training optimizes the clustering effect by using an iterative strategy.

[0079] Optionally, the training task of the second stage training comprises a masked cell modeling task; the training module is specifically configured to randomly mask target feature information of a first preset proportion in the to-be-input cell feature information to obtain masked to-be-input cell feature information; the training module is specifically further configured to input the masked to-be-input cell feature information into the two-dimensional rotation translation equivariant attention network, and predict the masked target feature information based on the cell representation output by the two-dimensional rotation translation equivariant attention network through a regression head; wherein the masked target feature information is feature information of a non-virtual cell; the masked cell modeling task uses mean square error as a loss function.

[0080] Optionally, the training task of the second stage training comprises a pair distance recovery task; the training module is specifically configured to randomly select target cells of a second preset proportion in the sub-slice, and add noise interference to the two-dimensional position coordinate information of the target cells to obtain an adjusted distance matrix; the training module is specifically further configured to input the adjusted distance matrix into the two-dimensional rotation translation equivariant attention network, and reconstruct the distance matrix disturbed by noise based on the pair expression output by the two-dimensional rotation translation equivariant attention network through a regression head; wherein the pair distance recovery task uses mean square error as a loss function.

[0081] The device for multi-scale modeling of spatial transcriptomics data provided by the application first acquires a spatial transcriptomics ST slice to be processed, and multi-scale information extraction is performed on the ST slice to obtain a plurality of to-be-processed sub-slices; then, the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices is sequentially input into a two-dimensional rotation translation equivariant attention network to obtain cell representation of each cell in each slice; finally, based on the cell representation of each cell in each slice, cell representation of each cell in the ST slice is obtained; wherein the multi-scale information extraction includes macro-scale information extraction, micro-scale information extraction and gene-scale information extraction; the to-be-processed information included in the to-be-processed sub-slice includes cell feature information for representing a spatial distance matrix of cell spacing. In this way, by capturing and fusing complex multi-scale information in spatial transcriptomics, high-quality general representation learning of spatial transcriptomics data is achieved, and an effective artificial intelligence tool is provided for spatial transcriptomics data analysis.

[0082] Figure 5 An example of a schematic diagram of a physical structure of an electronic device is shown as Figure 5 The electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can invoke a logical instruction in the memory 530 to execute a method for multi-scale modeling of spatial transcriptomics data, which includes: first, acquiring a spatial transcriptomics ST slice to be processed, and performing multi-scale information extraction on the ST slice to obtain a plurality of to-be-processed sub-slices; then, the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices is sequentially input into a two-dimensional rotation translation equivariant attention network to obtain cell representation of each cell in each slice; finally, based on the cell representation of each cell in each slice, cell representation of each cell in the ST slice is obtained; wherein the multi-scale information extraction includes macro-scale information extraction, micro-scale information extraction and gene-scale information extraction; the to-be-processed information included in the to-be-processed sub-slice includes cell feature information for representing a spatial distance matrix of cell spacing. In this way, by capturing and fusing complex multi-scale information in spatial transcriptomics, high-quality general representation learning of spatial transcriptomics data is achieved, and an effective artificial intelligence tool is provided for spatial transcriptomics data analysis.

[0083] Further, the logic instructions in the memory 530 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0084] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a computer readable storage medium, and the computer program comprises program instructions, when the program instructions are executed by a computer, the computer can execute the method for multi-scale modeling of spatial transcriptomics data provided by the above-mentioned methods, which comprises: first, obtaining a spatial transcriptomics ST slice to be processed, and performing multi-scale information extraction on the ST slice to obtain a plurality of to-be-processed sub-slices; then, inputting the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices into a two-dimensional rotation translation equivariant attention network in turn to obtain the cell representation of each cell in each slice; finally, based on the cell representation of each cell in each slice, obtaining the cell representation of each cell in the ST slice; wherein the multi-scale information extraction comprises macro-scale information extraction, micro-scale information extraction and gene-scale information extraction; the to-be-processed information contained in the to-be-processed sub-slice includes cell feature information for representing a spatial distance matrix of cell spacing. In this way, by capturing and fusing the complex multi-scale information in spatial transcriptomics, high-quality general representation learning of spatial transcriptomics data is realized, and an effective artificial intelligence tool is provided for spatial transcriptomics data analysis.

[0085] In yet another aspect, the present application also provides a computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the method for multi-scale modeling of spatial transcriptomic data provided above, which comprises: first, obtaining a spatial transcriptomic ST slice to be processed, and performing multi-scale information extraction on the ST slice to obtain a plurality of to-be-processed sub-slices; then, inputting the to-be-processed information corresponding to each to-be-processed sub-slice in the plurality of to-be-processed sub-slices into a two-dimensional rotation translation equivariant attention network in turn to obtain the cell representation of each cell in each slice; finally, based on the cell representation of each cell in each slice, obtaining the cell representation of each cell in the ST slice; wherein the multi-scale information extraction comprises macro-scale information extraction, micro-scale information extraction and gene-scale information extraction; the to-be-processed information contained in the to-be-processed sub-slice includes cell feature information for representing a spatial distance matrix of cell spacing. In this way, by capturing and fusing the complex multi-scale information in spatial transcriptomics, high-quality general representation learning of spatial transcriptomic data is achieved, providing an effective artificial intelligence tool for spatial transcriptomic data analysis.

[0086] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement without creative labor.

[0087] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus necessary universal hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in terms of the contribution to the prior art, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0088] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the same; although the present application has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for multi-scale modeling of spatial transcriptomics data, characterized in that, include: Obtain spatial transcriptomics ST slices to be processed, and perform multi-scale information extraction on the ST slices to obtain multiple sub-slices to be processed; The information to be processed corresponding to each of the multiple sub-slices to be processed is sequentially input into a two-dimensional rotation and translation isovariant attention network to obtain the cell representation of each cell in each slice; Based on the cell representation of each cell in each slice, the cell representation of each cell in the ST slice is obtained; The multi-scale information extraction includes: macro-scale information extraction, micro-scale information extraction, and gene-scale information extraction; the information to be processed contained in the sub-slice includes: cell feature information, which is used to characterize the spatial distance matrix between each cell.

2. The method for multi-scale modeling of spatial transcriptomics data according to claim 1, characterized in that, The process of extracting multi-scale information from the ST slice yields multiple sub-slices to be processed, including: Based on the two-dimensional position coordinate information of each cell in the ST slice, the cells contained in the ST slice are divided into multiple sub-slices; the set of cells contained in each sub-slice is a subset of the set of cells contained in the ST slice; The gene expression information of each cell in the ST slice is encoded using a pre-trained cell encoder to obtain the cell feature information of each cell. Based on the cell feature information and the two-dimensional position coordinate information of each cell, the cells in the ST slice are clustered to obtain multiple clusters. Based on the mean cell feature value and the mean two-dimensional position coordinate value of the cells contained in each cluster, virtual cells are generated for each cluster, resulting in multiple virtual cells. These multiple virtual cells are then merged into each sub-slice to obtain multiple sub-slices to be processed.

3. The method for multi-scale modeling of spatial transcriptomics data according to claim 1 or 2, characterized in that, Before sequentially inputting the information to be processed corresponding to each of the plurality of sub-slices into a two-dimensional rotation-translation isovariant attention network to obtain the cell representation of each cell in each slice, the method further includes: Multiple ST slice samples were obtained; each ST slice sample contained: cell feature information of each cell, two-dimensional position coordinate information of each cell, and gene expression information of each cell. Each of the multiple ST slice samples is clustered, and multiple sub-slice samples corresponding to each ST slice sample are obtained based on the clustering results. Each sub-slice sample is used as a training sample to train the pre-trained cellular encoder and the two-dimensional rotation and translation equivariant attention network at different stages, resulting in the trained cellular encoder and the trained two-dimensional rotation and translation equivariant attention network. The training at different stages includes: a first stage training and a second stage training; the first stage training includes: freezing the parameters of the cell encoder and training the two-dimensional rotation-translation equivariant attention network using the cell feature information contained in the sub-slice samples until the two-dimensional rotation-translation equivariant attention network satisfies a first preset convergence condition; the second stage training includes: randomly selecting a portion of cells in each forward propagation to calculate cell feature information from the cell encoder with partially unfrozen parameters, so as to synchronously update the cell encoder and the two-dimensional rotation-translation equivariant attention network through backpropagation; the first stage training does not use an iterative strategy, while the second stage training uses an iterative strategy to optimize the clustering effect.

4. The method for multi-scale modeling of spatial transcriptomics data according to claim 3, characterized in that, The training tasks in the second phase of training include: masked cell modeling task; The masked cell modeling task includes: Randomly mask a first preset proportion of target feature information in the cell feature information to be input, and obtain the masked cell feature information to be input. The masked cell feature information is input into a two-dimensional rotation and translation isovariant attention network, and the masked target feature information is predicted by a regression head based on the cell representation output by the two-dimensional rotation and translation isovariant attention network. The masked target feature information is the feature information of non-virtual cells; the masked cell modeling task uses mean squared error as the loss function.

5. The method for multi-scale modeling of spatial transcriptomics data according to claim 4, characterized in that, The loss function for the masked cell modeling task is: in, The value of the loss function. For the masked collection of cells, For the characteristic information of the j-th cell in the masked cells, For the first predicted by regression head j The characteristic information of each cell.

6. The method for multi-scale modeling of spatial transcriptomics data according to claim 3, characterized in that, The training tasks in the second phase of training include: paired distance recovery tasks; The paired distance recovery task includes: Randomly select target cells in a sub-slice at a second preset ratio, and add noise interference to the two-dimensional position coordinate information of the target cells to obtain an adjusted distance matrix; The adjusted distance matrix is ​​input into a two-dimensional rotation and translation equal-variable attention network, and the distance matrix affected by noise is reconstructed by the regression head based on the pairwise representations output by the two-dimensional rotation and translation equal-variable attention network. The pairwise distance recovery task uses the mean squared error as the loss function.

7. The method for multi-scale modeling of spatial transcriptomics data according to claim 6, characterized in that, The loss function for the pairwise distance recovery task is: in, The value of the loss function. For the set of elements affected by noise, For cells j and cells k Distance between them after being affected by noise For cells j and cells k The distance between them predicted by the regression head.

8. An apparatus for multi-scale modeling of spatial transcriptomics data, characterized in that, The device includes: The acquisition module is used to acquire the spatial transcriptomics ST slices to be processed; A multi-scale information extraction module is used to extract multi-scale information from the ST slice to obtain multiple sub-slices to be processed. The cell representation generation module is used to sequentially input the information to be processed corresponding to each of the multiple sub-slices to be processed into a two-dimensional rotation and translation isovariant attention network to obtain the cell representation of each cell in each slice. The cell representation generation module is also used to obtain the cell representation of each cell in the ST slice based on the cell representation of each cell in each slice; The multi-scale information extraction includes: macro-scale information extraction, micro-scale information extraction, and gene-scale information extraction; the information to be processed contained in the sub-slice includes: cell feature information, which is used to characterize the spatial distance matrix between each cell.

9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the method for multi-scale modeling of spatial transcriptomics data as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method for multi-scale modeling of spatial transcriptomics data as described in any one of claims 1 to 7.