Spatial Transcriptome Analysis Method, Device and Readable Medium Based on Deep Learning

The deep learning-based method for space transcriptome analysis efficiently segments cells and allocates RNA molecules, addressing the challenge of accurate cell segmentation in complex tissues, enabling precise cell type identification and tissue reconstruction.

CN116994245BActive Publication Date: 2025-07-15XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310961931.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-02
Publication Date
2025-07-15
Estimated Expiration
2043-08-02

AI Technical Summary

Technical Problem

The existing spatial transcriptomic analysis methods have insufficient cell segmentation speed and accuracy, especially in complex tissues, with slow and insufficient analysis speed.

Method used

Using a deep learning-based method, an instance segmentation neural network is constructed, including a hierarchical transformation encoder and a fully connected decoder, and nuclear segmentation is performed through effective self-attention layer, hybrid feedforward network layer and overlapping image block merging layer, and three-dimensional visual reconstruction is carried out in combination with point cloud technology.

Benefits of technology

Fast and precise nucleus segmentation is achieved, enabling accurate identification of nuclei and reconstructing the surface profile of tissues or organs, suitable for next-generation sequencing and imaging spatial transcriptomic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994245B_ABST
    Figure CN116994245B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and readable medium for spatial transcriptome analysis based on deep learning. By constructing and training an instance segmentation neural network, a cell nucleus segmentation model is obtained. The instance segmentation neural network includes a hierarchical transformation encoder and a fully connected decoder. The cell nucleus image is input into the cell nucleus segmentation model to obtain the cell nucleus segmentation result and determine the position information of the cell. The coordinates of ribonucleic acid molecules generated by spatial transcriptomics are obtained, and each ribonucleic acid molecule is assigned to the cell with the shortest distance from it according to the cell nucleus segmentation result and the coordinates of the ribonucleic acid molecules to form a single-cell expression matrix. According to the single-cell expression matrix, cell type annotation is performed on the cells to obtain the cell type, and the anatomical region is identified according to the cell type and the position information of the cell. The point cloud technology is used for three-dimensional visualization to reconstruct the surface contour of the tissue or organ. This method can perform cell segmentation quickly and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of spatial transcriptome data analysis, and in particular to a spatial transcriptome analysis method, device and readable medium based on deep learning. Background Art

[0002] Spatial transcriptomics can measure ribonucleic acid molecules while retaining spatial information, enabling researchers to study the cell gene regulatory network and the impact of the extracellular microenvironment on cell expression and function. In spatial transcriptome analysis, converting subcellular resolution molecular profiles into cell-level cell segmentation has become a major challenge. Several techniques for cell segmentation have been developed in this field, but their analysis speed is slow and inaccurate on complex tissues. To better understand the impact of the cell gene regulatory network and the extracellular microenvironment on cell expression and function, a fast and accurate cell segmentation method is urgently needed. Summary of the Invention

[0003] Aiming at the above-mentioned technical problems, the purpose of the embodiments of the present application is to propose a spatial transcriptome analysis method, device and readable medium based on deep learning to solve the technical problems mentioned in the above background art section.

[0004] In a first aspect, the present invention provides a spatial transcriptome analysis method based on deep learning, including the following steps:

[0005] Obtain a plurality of consecutive nuclear images;

[0006] Construct an instance segmentation neural network and train it to obtain a nuclear segmentation model. The instance segmentation neural network includes a hierarchical transformation encoder and a fully connected decoder. The hierarchical transformation encoding layer includes a first encoder block, a second encoder block, a third encoder block, and a fourth encoder block connected in sequence. The first encoder block, the second encoder block, the third encoder block, and the fourth encoder block all include an efficient self-attention layer, a hybrid feed-forward network layer, and an overlapping image patch merging layer. Each image patch in the nuclear image is input into the efficient self-attention layer to extract discriminative features. The discriminative features are input into the hybrid feed-forward network layer to inject local information and obtain a synthesized image patch. The synthesized image patch is input into the overlapping image patch merging layer for merging to obtain hierarchical features. The four hierarchical features corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the fully connected decoder to obtain a nuclear segmentation result and determine the position information of the cells;

[0007] Obtain the coordinates of ribonucleic acid molecules generated by spatial transcriptomics, input the nuclear image into the nuclear segmentation model to obtain a nuclear segmentation result, and assign each ribonucleic acid molecule to the cell with the shortest distance to it according to the nuclear segmentation result and the coordinates of the ribonucleic acid molecules, and form a single-cell expression matrix;

[0008] Based on the single-cell expression matrix, cell type annotation is performed on cells to obtain cell types, and anatomical regions are identified based on the cell types and the position information of the cells.

[0009] Based on the position information of the cells, a number of cell nucleus images are aligned to obtain the aligned cell nucleus images. Based on the aligned cell nucleus images, point cloud technology is used for three-dimensional visualization to reconstruct the surface contour of the tissue or organ.

[0010] Preferably, in the effective self-attention layer, the image patches are respectively multiplied by the keyword weight matrix W Q and the query weight matrix W K and the key-value weight matrix W V to obtain a keyword matrix K, a query matrix Q, and a key-value matrix V of size N×C. The keyword matrix K is input into a linear layer for dimensionality reduction to obtain the dimensionality-reduced keyword matrix

[0011]

[0012] where R is the reduction coefficient, and the linear layers in the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block respectively correspond to reduction coefficients that decrease in sequence;

[0013] According to the dimensionality-reduced keyword matrix the query matrix Q, and the key-value matrix V, self-attention operation is performed to obtain the discriminant feature formula as follows:

[0014]

[0015] where Attention(Q,K,V) is the discriminant feature, and d head represents the dimension of K.

[0016] Preferably, the hybrid feed-forward network layer includes a fourth fully connected layer, a 3×3 convolutional layer, a GELU activation function layer, and a fifth fully connected layer connected in sequence. The discriminant matrix is input into the fourth fully connected layer and sequentially passes through the 3×3 convolutional layer, the GELU activation function layer, and the fifth fully connected layer to obtain local features. The local features are added to the discriminant matrix to obtain a composite image patch. In the overlapping image patch merging layer, the stride S and the padding size P between two adjacent composite image patches are respectively defined, and a number of cooperative image patches are overlapped and merged according to the stride S and the padding size P to obtain hierarchical features.

[0017] Preferably, the fully connected decoder includes a first fully connected layer, a second fully connected layer, and a third fully connected layer. The four hierarchical features corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the first fully connected layer to output a first feature. The formula is as follows:

[0018]

[0019] Among them, F i is the hierarchical feature extracted by the i-th encoder block, is the first feature output by the i-th encoder block input into the first fully connected layer, and C i and C respectively refer to the input dimension and output dimension of the first fully connected layer;

[0020] The first feature undergoes upsampling and concatenation to obtain a connection feature. The formula is as follows:

[0021]

[0022]

[0023] Among them, is the output after upsampling , W represents the width of the nucleus image, and W / 4×W / 4 represents the target upsampling size;

[0024] The connection feature is input into the second fully connected layer to obtain a second feature. The formula is as follows:

[0025]

[0026] The second feature is input into the third fully connected layer to predict the nucleus segmentation result. The formula is as follows:

[0027]

[0028] Among them, M contains three channels, corresponding to the segmentation mask, vertical gradient, and horizontal gradient respectively. The intersection of the three is taken as the nucleus segmentation result.

[0029] Preferably, according to the nucleus segmentation result and the coordinates of ribonucleic acid molecules, each ribonucleic acid molecule is assigned to the cell with the shortest distance to it, and a single-cell expression matrix is formed. Specifically, it includes:

[0030] A unique gray value is assigned to each nucleus in the nucleus segmentation result. Each nucleus forms an independent region. According to the principle of connected components, the center of the nucleus is determined, and the center of the nucleus is used as the assumed cell center;

[0031] Calculate the distance between each ribonucleic acid molecule and each cell center, and assign each ribonucleic acid molecule to the cell with the shortest distance:

[0032] Cell RNA i = Cell j , if distance cell j = min(d cell 1, d cell 2, d cell 3, …, d cell n );

[0033] where i = 1, 2, …, M, Cell RNAi represents the cell to which the RNA i belongs, d cell j represents the distance between the RNA i and cell j , distance cell j represents the cell with the smallest distance from the RNA i , M represents the total amount of ribonucleic acid molecules, and n represents the total amount of cells;

[0034] Count the ribonucleic acid molecules corresponding to each cell to obtain a single-cell expression matrix.

[0035] Preferably, annotate the cells according to the single-cell expression matrix to obtain cell types, and identify the anatomical regions based on the cell types and the location information of the cells. Specifically, it includes:

[0036] Annotate the cell types according to the single-cell expression matrix using single-cell analysis techniques or Seurat, or annotate the cell types using the Tangram algorithm based on the single-cell data and the single-cell expression matrix to determine the cell types;

[0037] Generate a neighborhood cell type composition vector of the cell type distribution of K cells around each cell based on the cell type and the location information of the cells;

[0038] Use the kmeans clustering algorithm to cluster the neighborhood cell type composition vectors to obtain anatomical regions with different cell types.

[0039] Preferably, align several nuclear images according to the location information of the cells to obtain the aligned nuclear images, and perform three-dimensional visualization based on the aligned nuclear images using point cloud technology to reconstruct the surface contour of the tissue or organ. Specifically, it includes:

[0040] Determine the centroid of each nuclear image based on the cell center. The formula is as follows:

[0041] Centroid section r = ((r cell 1 + r cell 2 + r cell 3 + … + r cell n ) / n, (c cell 1 + c cell 2 + c cell 3 + … + c cell n ) / n);

[0042] Among them, Centroid section r represents the centroid of the nucleus image r, r cell i represents the row coordinate of the cell center of cell i, c cell i represents the column coordinate of the cell center of cell i;

[0043] Calculate the difference in pixel values between the aligned nucleus images, and select the angle that minimizes the difference as the final rotation angle for each nucleus image. The formula is as follows:

[0044] Rotate section r = min(d angle 0, d angle 1, d angle 2, …, d angle i ) {angle = 0:359};

[0045] Use the Alpha Shapes scatter contour algorithm to extract the external contour points of the cells and construct a convex hull, and reprocess the cell contours using an average filter and grid refinement to create a smooth and continuous surface to obtain the surface contour of the tissue or organ.

[0046] In a second aspect, the present invention provides a spatial transcriptome analysis device based on deep learning, including:

[0047] A data acquisition module configured to acquire a plurality of consecutive nucleus images;

[0048] The model construction module is configured to construct and train an instance segmentation neural network to obtain a cell nucleus segmentation model. The instance segmentation neural network includes a hierarchical transformation encoder and a fully connected decoder. The hierarchical transformation encoding layer includes a first encoder block, a second encoder block, a third encoder block, and a fourth encoder block connected in sequence. The first encoder block, the second encoder block, the third encoder block, and the fourth encoder block all include an efficient self-attention layer, a hybrid feed-forward network layer, and an overlapping image patch merging layer. Each image patch in the cell nucleus image is input into the efficient self-attention layer to extract discriminative features. The discriminative features are input into the hybrid feed-forward network layer to inject local information and obtain a synthesized image patch. The synthesized image patch is input into the overlapping image patch merging layer for merging to obtain hierarchical features. The four hierarchical features corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the fully connected decoder to obtain the cell nucleus segmentation result and determine the position information of the cells.

[0049] The RNA allocation module is configured to obtain the coordinates of ribonucleic acid molecules generated by spatial transcriptomics, input the cell nucleus image into the cell nucleus segmentation model to obtain the cell nucleus segmentation result, and allocate each ribonucleic acid molecule to the cell with the shortest distance thereto according to the cell nucleus segmentation result and the coordinates of the ribonucleic acid molecules, and form a single-cell expression matrix.

[0050] The annotation recognition module is configured to perform cell type annotation on cells according to the single-cell expression matrix to obtain cell types, and identify anatomical regions according to the cell types and the position information of the cells.

[0051] The three-dimensional reconstruction module is configured to align a plurality of cell nucleus images according to the position information of the cells to obtain the aligned cell nucleus images, and perform three-dimensional visualization based on the aligned cell nucleus images using point cloud technology to reconstruct the surface contour of the tissue or organ.

[0052] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the first aspect.

[0053] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] (1) The present invention inputs the pictures of 4',6-diamidino-2-phenylindole stained cell nuclei into a cell nucleus segmentation model for cell nucleus recognition, which can quickly and accurately segment cells and accurately identify the cell nuclei in the cell nucleus pictures.

[0056] (2) The present invention can use the connected domain principle to identify the cell nucleus center and use it as the cell center, calculate the distance between the ribonucleic acid molecule coordinates and the cell center coordinates, assign the ribonucleic acid molecules to the nearest cells to obtain a single-cell expression matrix, can accurately divide the ribonucleic acid molecules into cells, determine the cell types according to the single-cell expression matrix, and finally perform three-dimensional visualization based on the position information, cell types and anatomical regions of the cells and reconstruct the surface of the tissue / organ.

[0057] (3) The present invention has a relatively fast analysis speed and is applicable to next-generation sequencing-based spatial transcriptomics and imaging-based spatial transcriptomics, with wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0059] Figure 1 is an exemplary device architecture diagram to which an embodiment of the present application can be applied;

[0060] Figure 2 is a schematic flowchart of a deep learning-based spatial transcriptome analysis method according to an embodiment of the present application;

[0061] Figure 3 is a flowchart of a deep learning-based spatial transcriptome analysis method according to an embodiment of the present application;

[0062] Figure 4 is a schematic structural diagram of an instance segmentation neural network of a deep learning-based spatial transcriptome analysis method according to an embodiment of the present application;

[0063] Figure 5 is a schematic diagram of a deep learning-based spatial transcriptome analysis device according to an embodiment of the present application;

[0064] Figure 6 is a schematic structural diagram of a computer device of an electronic device suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0066] Figure 1 An exemplary device architecture 100 is shown that can apply the deep learning-based spatial transcriptome analysis method or the deep learning-based spatial transcriptome analysis device according to the embodiments of the present application.

[0067] As Figure 1 shown, the device architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0068] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications may be installed on the terminal devices 101, 102, 103, such as data processing applications, file processing applications, etc.

[0069] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.

[0070] The server 105 may be a server that provides various services, such as a background data processing server that processes files or data uploaded by the terminal devices 101, 102, 103. The background data processing server may process the obtained files or data to generate a processing result.

[0071] It should be noted that the deep learning-based spatial transcriptome analysis method provided by the embodiments of the present application may be executed by the server 105, or may be executed by the terminal devices 101, 102, 103. Correspondingly, the deep learning-based spatial transcriptome analysis device may be set in the server 105, or may be set in the terminal devices 101, 102, 103.

[0072] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0073] Figure 2 A spatial transcriptome analysis method based on deep learning provided by an embodiment of the present application is shown, including the following steps:

[0074] S1. Obtain a plurality of consecutive cell nucleus images.

[0075] Specifically, the cell nucleus images stained with 4',6-diamidino-2-phenylindole are used in the embodiments of the present application. Referring to Figure 3 this cell nucleus image is used as the input image of the cell nucleus segmentation model based on the instance segmentation neural network.

[0076] S2. Construct an instance segmentation neural network and train it to obtain a cell nucleus segmentation model. The instance segmentation neural network includes a hierarchical transformation encoder and a fully connected decoder. The hierarchical transformation encoding layer includes a first encoder block, a second encoder block, a third encoder block, and a fourth encoder block connected in sequence. The first encoder block, the second encoder block, the third encoder block, and the fourth encoder block all include an efficient self-attention layer, a hybrid feed-forward network layer, and an overlapping image patch merging layer. Each image patch in the cell nucleus image is input into the efficient self-attention layer to extract discriminant features. The discriminant features are input into the hybrid feed-forward network layer to inject local information to obtain a synthesized image patch. The synthesized image patch is input into the overlapping image patch merging layer for merging to obtain hierarchical features. The four hierarchical features respectively corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the fully connected decoder to obtain the cell nucleus segmentation result and determine the position information of the cells.

[0077] In a specific embodiment, in the efficient self-attention layer, the image patches are respectively multiplied by the keyword weight matrix W Q , the query weight matrix W K , and the key-value weight matrix W V to obtain a keyword matrix K, a query matrix Q, and a key-value matrix V with a size of N×C. The keyword matrix K is input into a linear layer for dimensionality reduction to obtain the dimensionality-reduced keyword matrix

[0078]

[0079] Wherein, R is a reduction coefficient, and the linear layers in the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block respectively correspond to reduction coefficients that decrease in sequence;

[0080] According to the keyword matrix after dimensionality reduction Query matrix Q and key-value matrix V are used for self-attention operation to obtain discriminative features. The formula is as follows:

[0081]

[0082] Wherein, Attention(Q, K, V) is the discriminative feature, and d head represents the dimension of K.

[0083] In a specific embodiment, the hybrid feed-forward network layer includes a fourth fully-connected layer, a 3×3 convolutional layer, a GELU activation function layer, and a fifth fully-connected layer connected in sequence. The discriminant matrix is input into the fourth fully-connected layer and sequentially passes through the 3×3 convolutional layer, the GELU activation function layer, and the fifth fully-connected layer to obtain local features. The local features are added to the discriminant matrix to obtain a composite image patch. In the overlapping image patch merging layer, the stride S and padding size P between two adjacent composite image patches are respectively defined, and a plurality of cooperative image patches are overlapped and merged according to the stride S and padding size P to obtain hierarchical features.

[0084] Specifically, referring to Figure 4 , the instance segmentation neural network consists of two modules, namely a hierarchical transformation encoder and a lightweight fully-connected decoder. Specifically, given a nucleus image of size H×W×3, the nucleus image is segmented into several image patches, and the image patches are input into the hierarchical transformation encoder module to obtain multiple hierarchical features, including high-resolution features and low-resolution features, while the lightweight fully-connected decoder module is used to obtain the vertical gradient and horizontal gradient of the segmentation result map, as well as a segmentation mask indicating whether a given pixel is inside or outside the region of interest. To obtain the final nucleus recognition result, a gradient tracking method is used to construct the nucleus and its shape by aggregating the corresponding central pixels in a given nucleus and the surrounding pixels belonging to the same central pixel.

[0085] The hierarchical transformation encoder module includes four encoder blocks, and each encoder block respectively generates four hierarchical features with resolutions of {1 / 1, 1 / 2, 1 / 4, 1 / 8} of the nucleus image resolution. Each encoder block includes an effective self-attention layer, a hybrid feed-forward network layer, and an overlapping image patch merging layer. The effective self-attention layer is used to extract discriminative features, the hybrid feed-forward network layer is used to inject position information, and the overlapping image patch merging layer is used to obtain hierarchical features.

[0086] Specifically, the effective self-attention layer is a multi-head self-attention process. To reduce the computational complexity, a series of dimensionality reduction processes are used. By forwarding K into the linear layer, the dimension of K is reduced from N to N / R×C. Therefore, the complexity of the multi-head self-attention process can be reduced from O(N^2) to O(N^2 / R). In the embodiments of the present application, from the first encoder block to the fourth encoder block, R is sequentially set to [64, 16, 4, 1].

[0087] Local information is crucial for visual tasks, especially for semantic segmentation. In the embodiments of the present application, local information is injected into the features by applying a 3×3 convolutional layer in the hybrid feed-forward network layer to obtain local features, as shown in the following formula:

[0088] x out = MLP(GELU(Conv 3×3 (MLP(x in )))) + x in ;

[0089] where x in is the discriminative feature output by the effective self-attention layer, and x out is the local feature. In addition to the 3×3 convolutional layer, it also includes the fourth fully connected layer, the GELU activation function layer, the fifth fully connected layer and addition. A 3×3 convolutional layer is also applied to reduce the number of parameters, which can effectively improve the efficiency.

[0090] The image patch merging technique is used to unify the N×N×3 image patches into a 1×1×C vector. In the embodiments of the present application, the image patch merging technique is extended to overlapping image patch merging, which can generate hierarchical features while maintaining local continuity around the image patches. Specifically, T is defined as the image patch size, S is defined as the stride between two adjacent image patches, and P is defined as the padding size. Among them, the stride S means that the second image patch moves S pixel values to the right and down compared to the first image patch. The padding size P means that P pixel widths are added to the top, bottom, left, and right of the image patch. In the overlapping image patch merging layer, {T = 7, S = 4, P = 3} and {T = 3, S = 2, P = 1} can be set so that the size of the feature map can be halved. For example, given a feature map of size H / 1×W / 1×C_1, after overlapping image patch merging, it is reduced to the size of {H / 2×W / 2×C_2}. The overlapping image patch merging layer is added to each encoder block to generate hierarchical features.

[0091] In a specific embodiment, the fully connected decoder includes the first fully connected layer, the second fully connected layer and the third fully connected layer. The four hierarchical features corresponding to the first encoder block, the second encoder block, the third encoder block and the fourth encoder block are respectively input into the first fully connected layer, and the first feature is output, as shown in the following formula:

[0092]

[0093] Among them, F i is the hierarchical feature extracted by the i-th encoder block, is the first feature output by the first fully connected layer for the input of the i-th encoder block, and C i and C respectively refer to the input dimension and output dimension of the first fully connected layer;

[0094] The first feature undergoes upsampling and concatenation to obtain a concatenated feature, and the formula is as follows:

[0095]

[0096]

[0097] Among them, is the output after upsampling , W represents the width of the nucleus image, and W / 4×W / 4 represents the target upsampling size;

[0098] The concatenated feature is input into the second fully connected layer to obtain a second feature, and the formula is as follows:

[0099]

[0100] The second feature is input into the third fully connected layer, and the nucleus segmentation result is predicted, and the formula is as follows:

[0101]

[0102] Among them, M contains three channels, corresponding to the segmentation mask, vertical gradient, and horizontal gradient respectively, and the intersection of the three is taken as the nucleus segmentation result.

[0103] Specifically, the fully connected decoder, as an upsampling model, consists of three fully connected layers. Multiple hierarchical features obtained by the hierarchical transformation encoder are input into the first fully connected layer to convert the channel dimensions of the multiple hierarchical features into the same dimension, obtaining the first feature. Then, the first feature is upsampled to 1 / 4 of the original image resolution and concatenated together, and then the concatenated feature is input into the second fully connected layer for fusion to obtain the second feature. The second feature is input into the third fully connected layer to predict the segmentation mask, vertical gradient, and horizontal gradient. M has three channels, and the intersection of the three is taken as the final nucleus segmentation result.

[0104] Specifically, the embodiments of the present application collect 2 known datasets as training data to train a nucleus segmentation model for nucleus recognition, and the nucleus segmentation model can be directly loaded during use to predict each nucleus in the nucleus image.

[0105] S3. Obtain the coordinates of ribonucleic acid molecules generated by spatial transcriptomics, input the nuclear image into the nuclear segmentation model to obtain the nuclear segmentation result, and assign each ribonucleic acid molecule to the cell with the shortest distance to it according to the nuclear segmentation result and the coordinates of the ribonucleic acid molecules, and form a single-cell expression matrix.

[0106] In a specific embodiment, assigning each ribonucleic acid molecule to the cell with the shortest distance to it according to the nuclear segmentation result and the coordinates of the ribonucleic acid molecules, and forming a single-cell expression matrix specifically includes:

[0107] Assign a unique gray value to each nucleus in the nuclear segmentation result. Each nucleus constitutes an independent region. Determine the center of the nucleus according to the principle of connected components, and use the center of the nucleus as the assumed cell center;

[0108] Calculate the distance between each ribonucleic acid molecule and each cell center, and assign each ribonucleic acid molecule to the cell with the shortest distance:

[0109] Cell RNA i =Cell j ,if distance cell j =min(d cell 1,d cell 2,d cell 3,…,d cell n );

[0110] where i = 1, 2, …, M, Cell RNAi represents the cell to which the RNA i belongs, d cell j represents the distance between the RNA i and cell j , distance cell j represents the cell with the smallest distance from the RNA i , M represents the total amount of ribonucleic acid molecules, and n represents the total amount of cells;

[0111] Count the ribonucleic acid molecules corresponding to each cell to obtain a single-cell expression matrix.

[0112] Specifically, obtain the coordinate matrix of ribonucleic acid molecules generated by spatial transcriptomics. Spatial transcriptomics can measure ribonucleic acid molecules while retaining spatial information, enabling researchers to study the cell gene regulatory network and the impact of the extracellular microenvironment on cell expression and function. By inputting the collected nuclear images into the trained nuclear segmentation model, the nuclear segmentation results are output, and further, the cell position information is determined based on the nuclear segmentation results. Specifically, after using the nuclear segmentation model to identify the nuclei, in the nuclear segmentation results, each nucleus is assigned a unique grayscale value on the image, and the pixels with the same grayscale value correspond to the same nucleus. Each nucleus constitutes an independent region. In the field of computer vision, a connected component is defined as a set of adjacent pixels with the same grayscale value. Therefore, the principle of connected components can be used to determine the number of pixels occupied by the center of the nucleus (the assumed cell center), and the cell position information can also be determined, including the coordinates of the cell center.

[0113] Since the expression of cells shows aggregation in space, calculate the distance between each ribonucleic acid molecule and each cell center, and assign each ribonucleic acid molecule to the cell with the shortest distance, thereby forming a cell expression matrix with single-cell resolution.

[0114] S4. Perform cell type annotation on the cells according to the single-cell expression matrix to obtain the cell type, and identify the anatomical region based on the cell type and the cell position information.

[0115] In a specific embodiment, performing cell type annotation on the cells according to the single-cell expression matrix to obtain the cell type, and identifying the anatomical region based on the cell type and the cell position information specifically includes:

[0116] Perform cell type annotation according to the single-cell expression matrix using single-cell analysis techniques or Seurat, or perform cell type annotation using the Tangram algorithm according to the single-cell data and the single-cell expression matrix to determine the cell type;

[0117] Generate a neighborhood cell type composition vector of the cell type distribution of K cells around each cell based on the cell type and the cell position information;

[0118] Use the kmeans clustering algorithm to cluster the neighborhood cell type composition vectors to obtain anatomical regions with different cell types.

[0119] Specifically, different cell type annotation methods can be selected for different data. Usually, cell segmentation generates a cell expression matrix with single-cell resolution, which can be annotated using standard single-cell analysis techniques. However, in situ sequencing technology sometimes only measures a limited number of gene types, so traditional single-cell analysis techniques cannot be used to annotate cell types. To connect single-cell data (the single-cell expression matrix generated by single-cell sequencing technology) and spatial transcriptomics data, the Tangram algorithm can be used to annotate the cell types of spatial transcriptomics. The Tangram algorithm takes the single-cell data and the cell expression matrix generated in step S3 as input, uses the KL divergence and cosine similarity to simulate the gene similarity between the single-cell data and the cell expression matrix, and generates a probability mapping matrix, that is, the probability that each cell in the cell expression matrix corresponds to a cell in the single-cell data. As the last step, in spatial transcriptomics, the cell type with the highest probability in the single-cell data is selected as the cell type of the cell.

[0120] When the number of genes is sufficient, Seurat is used to cluster and annotate cells. The cell expression matrix is used as the input of Seurat. After a series of processes, including quality control, dimensionality reduction, clustering, and finding marker genes, the cluster to which each cell belongs and the marker genes of each cluster are determined. Then, by comparing these marker genes with the markers of known cell types, the cell type corresponding to each cluster can be determined.

[0121] After determining the cell type of the cell, the anatomical region is identified using the cell type and the position information of the cell. First, a neighborhood cell type composition vector (NCCV) representing the cell type distribution of K cells around each cell is generated. Then, the kmeans clustering algorithm is used to cluster the neighborhood cell type composition vector (NCCV) to obtain anatomical regions (different regions in the tissue) with different cell type compositions, that is, an anatomical region label is assigned to each cell.

[0122] S5. Align several nuclear images according to the position information of the cells to obtain the aligned nuclear images, and use point cloud technology for three-dimensional visualization based on the aligned nuclear images to reconstruct the surface contour of the tissue or organ.

[0123] In a specific embodiment, step S5 specifically includes:

[0124] Determine the centroid of each nuclear image based on the cell center, and the formula is as follows:

[0125] Centroid section r =((r cell 1+r cell 2+r cell 3+…+rcell n ) / n,(c cell 1 + c cell 2 + c cell 3 + … + c cell n ) / n);

[0126] where, Centroid section r represents the centroid of the nucleus image r, and r cell i represents the row coordinate of the cell center of cell i, and c cell i represents the column coordinate of the cell center of cell i;

[0127] Calculate the difference in pixel values between the aligned nucleus images, and select the angle that minimizes the difference as the final rotation angle for each nucleus image. The formula is as follows:

[0128] Rotate section r = min(d angle 0, d angle 1, d angle 2, …, d angle i ){angle = 0:359};

[0129] Use the Alpha Shapes scatter contour algorithm to extract the external contour points of the cells and construct the convex hull, and reprocess the cell contours using an average filter and grid refinement to create a smooth and continuous surface to obtain the surface contour of the tissue or organ.

[0130] Specifically, since the organ is divided into continuous nucleus images, the contours of adjacent nucleus images are very similar. Therefore, the contours of multiple nucleus images can be aligned. First, use the erosion and dilation processes to completely fill the inside of the nucleus image and eliminate environmental noise. In addition, since a finite number of points always have a geometric center (centroid), the centroid can be obtained by calculating the arithmetic mean of each coordinate component of these points. Therefore, the centroid of each nucleus image is determined based on the cell center. Before rotation alignment, the centroids of the aligned adjacent nucleus images are already aligned. Then, calculate the difference by rotating one part and subtracting the pixel values at the aligned positions of the other parts. To achieve the automatic alignment of multiple nucleus images, the angle that minimizes the difference is selected as the final rotation angle for each nucleus image. After translational alignment and rotation, the multiple nucleus images will be aligned.

[0131] After alignment, point cloud technology is used to reconstruct the three-dimensional surface of the aligned nucleus image. A point cloud is a collection of three-dimensional coordinate points and can be used to describe the spatial contour and precise position of an object. In the embodiments of the present application, three-dimensional point cloud visualization and three-dimensional point cloud reconstruction are used. Three-dimensional point cloud visualization uses the position information of cells (row coordinates and column coordinates in two-dimensional space), cell types, and anatomical regions for three-dimensional visualization in three-dimensional space. Three-dimensional point cloud reconstruction is to reconstruct the surface contour of an organ based on cell coordinates. The three-dimensional point cloud reconstruction technology will first use the Alpha Shapes scatter contour algorithm to extract external contour points and construct a convex hull. At this time, the obtained external contour is relatively rough. Then, it is reprocessed using an average filter and mesh refinement to create a smooth and continuous surface.

[0132] Among them, the idea of the Alpha Shapes algorithm is as follows:

[0133] (1) Set a threshold radius R. This value determines the fineness of the boundary. The smaller it is, the finer the boundary;

[0134] (2) Assume that the data set has n unordered points. Draw a circle with a radius of R through any two points N1 and N2 (excluding the case where the distance between the two points is 2R. Obviously, there are usually two circles that meet the requirements). If there are no other data points inside any circle, it is considered that the points N1 and N2 belong to the boundary points, and P1P2 is the boundary line segment.

[0135] The above steps S1 - S5 do not represent the order between the steps, but are only step symbol representations.

[0136] For further reference Figure 5 , as an implementation of the methods shown in the above figures, the present application provides an embodiment of a spatial transcriptome analysis device based on deep learning. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0137] The embodiments of the present application provide a spatial transcriptome analysis device based on deep learning, which is characterized by including:

[0138] A data acquisition module 1, configured to acquire a number of consecutive nucleus images;

[0139] The model construction module 2 is configured to construct and train an instance segmentation neural network to obtain a nucleus segmentation model. The instance segmentation neural network includes a hierarchical transformation encoder and a fully connected decoder. The hierarchical transformation encoding layer includes a first encoder block, a second encoder block, a third encoder block, and a fourth encoder block connected in sequence. The first encoder block, the second encoder block, the third encoder block, and the fourth encoder block all include an efficient self-attention layer, a hybrid feed-forward network layer, and an overlapping image patch merging layer. Each image patch in the nucleus image is input into the efficient self-attention layer to extract discriminative features. The discriminative features are input into the hybrid feed-forward network layer to inject local information and obtain a synthesized image patch. The synthesized image patch is input into the overlapping image patch merging layer for merging to obtain hierarchical features. The four hierarchical features corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the fully connected decoder to obtain the nucleus segmentation result and determine the position information of the cells.

[0140] The RNA allocation module 3 is configured to obtain the coordinates of ribonucleic acid molecules generated by spatial transcriptomics, input the nucleus image into the nucleus segmentation model to obtain the nucleus segmentation result, and allocate each ribonucleic acid molecule to the cell with the shortest distance from it according to the nucleus segmentation result and the coordinates of the ribonucleic acid molecules, and form a single-cell expression matrix.

[0141] The annotation recognition module 4 is configured to perform cell type annotation on the cells according to the single-cell expression matrix to obtain the cell type, and identify the anatomical region according to the cell type and the position information of the cells.

[0142] The three-dimensional reconstruction module 5 is configured to align a plurality of nucleus images according to the position information of the cells to obtain the aligned nucleus images, and perform three-dimensional visualization based on the aligned nucleus images using point cloud technology to reconstruct the surface contour of the tissue or organ.

[0143] Next, refer to Figure 6 , which shows a schematic structural diagram of a computer device 600 of an electronic device (such as Figure 1 the server or terminal device shown) suitable for implementing the embodiments of the present application. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0144] As Figure 6As shown, the computer device 600 includes a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 603 or programs loaded from a storage section 609 into a random access memory (RAM) 604. In the RAM 604, various programs and data required for the operation of the device 600 are also stored. The CPU 601, GPU 602, ROM 603, and RAM 604 are connected to each other via a bus 605. An input / output (I / O) interface 606 is also connected to the bus 605.

[0145] The following components are connected to the I / O interface 606: an input section 607 including a keyboard, a mouse, etc.; an output section 608 including, for example, a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 609 including a hard disk, etc.; and a communication section 610 including a network interface card such as a LAN card, a modem, etc. The communication section 610 performs communication processing via a network such as the Internet. A drive 611 may also be connected to the I / O interface 606 as needed. A removable medium 612, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 611 as needed so that a computer program read therefrom can be installed into the storage section 609 as needed.

[0146] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 610, and / or installed from the removable medium 612. When the computer program is executed by the central processing unit (CPU) 601 and the graphics processing unit (GPU) 602, the above functions defined in the method of the present application are executed.

[0147] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable medium or any combination of the two. The computer-readable medium can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or components, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution device, apparatus, or component. And in this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0148] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0150] The modules described in the embodiments of the present application can be implemented in software or in hardware. The described modules can also be provided in a processor.

[0151] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a plurality of consecutive nuclear images; construct and train an instance segmentation neural network to obtain a nuclear segmentation model, the instance segmentation neural network includes a hierarchical transformation encoder and a fully connected decoder, the hierarchical transformation encoding layer includes a first encoder block, a second encoder block, a third encoder block, and a fourth encoder block connected in sequence, the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block each include an effective self-attention layer, a hybrid feed-forward network layer, and an overlapping image patch merging layer, each image patch in the nuclear image is input into the effective self-attention layer to extract discriminative features, the discriminative features are input into the hybrid feed-forward network layer to inject local information to obtain a synthesized image patch, and the synthesized image patch is input into the overlapping image patch merging layer for merging to obtain hierarchical features; the four hierarchical features respectively corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the fully connected decoder to obtain a nuclear segmentation result and determine the position information of the cells; obtain the coordinates of ribonucleic acid molecules generated by spatial transcriptomics, input the nuclear image into the nuclear segmentation model to obtain a nuclear segmentation result, and assign each ribonucleic acid molecule to the cell with the shortest distance thereto according to the nuclear segmentation result and the coordinates of the ribonucleic acid molecules to form a single-cell expression matrix; perform cell type annotation on the cells according to the single-cell expression matrix to obtain cell types, and identify the anatomical regions according to the cell types and the position information of the cells; align a plurality of nuclear images according to the position information of the cells to obtain aligned nuclear images, and perform three-dimensional visualization based on the aligned nuclear images using point cloud technology to reconstruct the surface contour of the tissue or organ.

[0152] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present application.

Claims

1. A method for spatial transcriptome analysis based on deep learning, characterized in that, It includes the following steps: Obtain several consecutive nucleus images; Construct and train an instance segmentation neural network to obtain a nucleus segmentation model. The instance segmentation neural network includes a hierarchical transformation encoder and a fully connected decoder. The hierarchical transformation encoder includes a first encoder block, a second encoder block, a third encoder block, and a fourth encoder block connected in sequence. The first encoder block, the second encoder block, the third encoder block, and the fourth encoder block each include an efficient self-attention layer, a hybrid feed-forward network layer, and an overlapping image patch merging layer. Each image patch in the nucleus image is input into the efficient self-attention layer to extract discriminative features. The discriminative features are input into the hybrid feed-forward network layer to inject local information and obtain a synthesized image patch. The hybrid feed-forward network layer includes a fourth fully connected layer, a 3×3 convolutional layer, a GELU activation function layer, and a fifth fully connected layer connected in sequence. The discriminative features are input into the fourth fully connected layer and sequentially pass through the 3×3 convolutional layer, the GELU activation function layer, and the fifth fully connected layer to obtain local features. The local features are added to the discriminative features to obtain the synthesized image patch. The synthesized image patch is input into the overlapping image patch merging layer for merging to obtain hierarchical features. In the overlapping image patch merging layer, the stride S and padding size P between two adjacent synthesized image patches are respectively defined, and several synthesized image patches are overlapped and merged according to the stride S and padding size P to obtain hierarchical features; The four hierarchical features respectively corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the fully connected decoder to obtain the nucleus segmentation result and determine the position information of the cells; Obtain the coordinates of ribonucleic acid molecules generated by spatial transcriptomics. Input the nucleus image into the nucleus segmentation model to obtain the nucleus segmentation result. According to the nucleus segmentation result and the coordinates of the ribonucleic acid molecules, each ribonucleic acid molecule is assigned to the cell with the shortest distance to it, and a single-cell expression matrix is formed; Perform cell type annotation on the cells according to the single-cell expression matrix to obtain the cell type, and identify the anatomical region according to the cell type and the position information of the cells; Align several nucleus images according to the position information of the cells to obtain the aligned nucleus images, and perform three-dimensional visualization based on the aligned nucleus images using point cloud technology to reconstruct the surface contour of the tissue or organ.

2. The spatial transcriptome analysis method based on deep learning according to claim 1, wherein In the effective self-attention layer, multiply the image patches by the keyword weight matrix W Q , the query weight matrix W K , the key-value weight matrix W V to obtain a keyword matrix K, a query matrix Q, and a key-value matrix V of size N×C. Input the keyword matrix K into a linear layer for dimensionality reduction to obtain the dimensionality-reduced keyword matrix : Wherein, R is a reduction coefficient, and the linear layers in the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block respectively correspond to reduction coefficients that decrease in sequence; According to the keyword matrix after dimensionality reduction Query matrix Q and key matrix V are used for self-attention operation to obtain the discriminant feature formula as follows: Among them, Attention(Q, K, V) is the discriminative feature, and d head represents the dimension of K.

3. The spatial transcriptome analysis method based on deep learning according to claim 1, characterized in that The fully connected decoder includes a first fully connected layer, a second fully connected layer, and a third fully connected layer. The four hierarchical features respectively corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the first fully connected layer, and the first feature is output. The formula is as follows: Among them, F i is the hierarchical feature extracted by the i-th encoder block, is the first feature output by the input of the i-th encoder block to the first fully connected layer, P i and P respectively refer to the input dimension and output dimension of the first fully connected layer; The first feature is upsampled and concatenated to obtain a connection feature The formula is as follows: Among them, is the output after upsampling , W represents the width of the nucleus image, and W / 4×W / 4 represents the target upsampling size. The connection feature Input the second fully-connected layer to obtain the second feature The formula is as follows: The second feature Input the third fully connected layer, and the nucleus segmentation result is predicted. The formula is as follows: Wherein, M contains three channels, corresponding to the segmentation mask, the vertical gradient, and the horizontal gradient respectively. The intersection of the three is taken as the nucleus segmentation result M.

4. The spatial transcriptome analysis method based on deep learning according to claim 1, wherein Assigning each ribonucleic acid molecule to the cell with the shortest distance thereto according to the nucleus segmentation result and the coordinates of the ribonucleic acid molecule, and forming a single-cell expression matrix, specifically including: Assigning a unique gray value to each nucleus in the nucleus segmentation result, each nucleus constituting an independent region, determining the center of the nucleus according to the principle of connected components, and taking the center of the nucleus as the assumed cell center; Calculating the distance between each ribonucleic acid molecule and each cell center, and assigning each ribonucleic acid molecule to the cell with the shortest distance; cell RNAl = cell j , if distance cell j = min(d cell 1 , d cell 2 , d cell 3 , …, d celln ); where l = 1, 2, …, M, cell RNAl represents the cell to which the RNA l belongs, and d cell j represents the distance between the RNA l and cell j ; distance cellj represents the distance between the RNA l and the cell with the smallest distance thereto; M represents the total amount of ribonucleic acid molecules, and n represents the total amount of cells; Counting the ribonucleic acid molecules corresponding to each cell to obtain the single-cell expression matrix.

5. The spatial transcriptome analysis method based on deep learning according to claim 1, characterized in that Performing cell type annotation on cells according to the single-cell expression matrix to obtain cell types, and identifying anatomical regions according to cell types and the position information of cells, specifically including: Performing cell type annotation according to the single-cell expression matrix using single-cell analysis techniques or Seurat, or performing cell type annotation using the Tangram algorithm according to single-cell data and the single-cell expression matrix to determine cell types; Generating a neighborhood cell type composition vector of the cell type distribution of K cells around each cell according to the cell type and the position information of the cell; Using the kmeans clustering algorithm to cluster the neighborhood cell type composition vectors to obtain anatomical regions with different cell types.

6. The spatial transcriptome analysis method based on deep learning according to claim 4, characterized in that Aligning a plurality of nucleus images according to the position information of the cells to obtain aligned nucleus images, and performing three-dimensional visualization on the basis of the aligned nucleus images using point cloud technology to reconstruct the surface contour of a tissue or an organ, specifically including: Determining the centroid of each nucleus image based on the cell center, and the formula is as follows: Centroid section r = ((r cell 1 + r cell 2 + r cell 3 + … + r celln ) / n, (c cell 1 + c cell 2 + c cell 3 + … + c celln ) / n); Among them, Centroid section r represents the centroid of the nucleus image r, where r cell j represents the row coordinate of the cell center of cell j, and c cell j represents the column coordinate of the cell center of cell j; Calculate the difference in pixel values between the aligned cell nucleus images, and select the angle that minimizes the difference as the final rotation angle Rotate for each cell nucleus image section r , and the formula is as follows: Rotate section r = min(d angle 0 , d angle 1 , d angle 2 , …, d angle k ) {angle = 0:359}; Extracting the external contour points of the cells using the Alpha Shapes scatter contour algorithm and constructing a convex hull, and reprocessing the contour of the cells using an average filter and grid refinement to create a smooth and continuous surface to obtain the surface contour of the tissue or the organ.

7. A spatial transcriptome analysis device based on deep learning, characterized in that, Including: A data acquisition module configured to acquire a plurality of consecutive nucleus images; A model construction module, configured to construct and train an instance segmentation neural network to obtain a nucleus segmentation model. The instance segmentation neural network includes a hierarchical transformation encoder and a fully connected decoder. The hierarchical transformation encoder includes a first encoder block, a second encoder block, a third encoder block, and a fourth encoder block connected in sequence. The first encoder block, the second encoder block, the third encoder block, and the fourth encoder block each include an efficient self-attention layer, a hybrid feed-forward network layer, and an overlapping image patch merging layer. Each image patch in the nucleus image is input into the efficient self-attention layer to extract discriminative features. The discriminative features are input into the hybrid feed-forward network layer to inject local information and obtain a synthetic image patch. The hybrid feed-forward network layer includes a fourth fully connected layer, a 3×3 convolutional layer, a GELU activation function layer, and a fifth fully connected layer connected in sequence. The discriminative features are input into the fourth fully connected layer and sequentially pass through the 3×3 convolutional layer, the GELU activation function layer, and the fifth fully connected layer to obtain local features. The local features are added to the discriminative features to obtain the synthetic image patch. The synthetic image patch is input into the overlapping image patch merging layer for merging to obtain hierarchical features. In the overlapping image patch merging layer, the stride S and padding size P between two adjacent synthetic image patches are respectively defined, and a plurality of synthetic image patches are overlapped and merged according to the stride S and padding size P to obtain hierarchical features; The four hierarchical features corresponding to the first encoder block, the second encoder block, the third encoder block, and the fourth encoder block are input into the fully connected decoder to obtain a nucleus segmentation result and determine the position information of the cells; An RNA allocation module, configured to obtain the coordinates of ribonucleic acid molecules generated by spatial transcriptomics, input the nucleus image into the nucleus segmentation model to obtain a nucleus segmentation result, and allocate each ribonucleic acid molecule to the cell with the shortest distance from it according to the nucleus segmentation result and the coordinates of the ribonucleic acid molecules, and form a single-cell expression matrix; An annotation recognition module, configured to perform cell type annotation on cells according to the single-cell expression matrix to obtain cell types, and identify anatomical regions according to the cell types and the position information of the cells; A three-dimensional reconstruction module, configured to align a plurality of nucleus images according to the position information of the cells to obtain aligned nucleus images, and perform three-dimensional visualization based on the aligned nucleus images using point cloud technology to reconstruct the surface contour of a tissue or organ.

8. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Cell function annotation method, device, equipment and medium

    CN114496099A

  • Spatial transcriptome biological tissue substructure analysis method fused with single cell transcriptome

    CN115359845A