A spatial cell type parsing method based on improved SSD-ResNet

The improved SSD-ResNet model is used to extract and analyze features of spatial transcriptome data, which solves the problem of the traditional ResNet model outputting a single result, achieves accurate analysis of multiple cell types, and improves the analysis accuracy of spatial transcriptome data.

CN117037894BActive Publication Date: 2025-09-19GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311058955.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-21
Publication Date
2025-09-19
Estimated Expiration
2043-08-21

AI Technical Summary

Technical Problem

Existing single-cell RNA sequencing technology and spatial transcription technology cannot achieve the accuracy of single cells. The cell analysis of spatial transcriptome data has the problem of insufficient accuracy. The traditional ResNet model has a single output and cannot meet the analysis requirements of multiple cell types.

Method used

An improved SSD-ResNet model was used to preprocess single-cell data and spatial transcriptome data, construct a ResNet network and add SSD. The SSD-ResNet model was trained to perform cell type analysis, and multiple classes and their probability distributions were output using multi-target detection.

Benefits of technology

It improves the cell parsing accuracy of spatial transcriptome data, can output multiple cell types and their probability distribution, and expands the application field of ResNet in bioinformatics data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117037894B_ABST
    Figure CN117037894B_ABST
Patent Text Reader

Abstract

The present invention relates to bioinformatics spatial transcriptome and single-cell sequencing data analysis, and more specifically, to a spatial cell type parsing method based on an improved SSD-ResNet. The present invention aims to distinguish the cellular composition of each spatial detection position, estimate the proportion of each cell type in each detection area and the gene expression level of each cell. First, the genes retained in the single-cell gene expression data and spatial transcriptome data are constructed into an n*n image. Then, an SSD-ResNet model is constructed and the model is trained using heat map images of gene expression. Finally, the trained model is used to perform cellular parsing on the spatial transcriptome data. The obtained optimal model is used to analyze the spatial transcriptome data to obtain the final cellular parsing results of the spatial transcriptome data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics spatial transcriptome and single-cell sequencing data analysis, and more specifically, to a spatial cell type parsing method based on an improved SSD-ResNet. Background Art

[0002] In recent years, with the development of bioinformatics, single-cell RNA sequencing (scRNA-seq) can link the transcriptome with single cells, but lacks the spatial location information of transcripts in tissues. Spatial transcriptomics (ST) is a natural extension of scRNA-seq. Spatial transcriptome data can include the gene expression of the transcriptome and information about the location distribution of sequenced cells, thereby enhancing our understanding of cell interactions, organ function and pathology. The integration of scRNA-seq and ST can provide a better potential dynamic picture of the development of tissues and organisms, greatly improving people's understanding of biological processes.

[0003] However, there is a problem when integrating scRNA-seq and ST: the spatial resolution of spatial transcriptomics (the number of cells contained in a single detection area) is limited and cannot achieve the accuracy of single cells. Spatial transcriptomics technology can obtain gene expression measurements in the detection area, which are actually composed of several or even dozens of different types of cells. Therefore, each detection area detects a mixture of cells, rather than a single type of single cell. In this case, cellular analysis of the spatial transcriptome is required.

[0004] The purpose of analyzing spatial transcriptome data is to clearly understand the cell composition of each detection area and to estimate the proportion of each cell type in each detection area and the gene expression level of each cell. At present, the deconvolution methods of spatial transcriptome data can be roughly divided into the following three types: probabilistic methods, methods based on non-negative matrix factorization (NMF) and non-negative least squares (NNLS), and other methods. Among them, probabilistic methods explicitly or parameterize the data distribution and use likelihood-based methods for inference, such as Adroit, cell2location, DestVI, RCTD, STdeconvolve and stereoscope, methods based on non-negative matrix factorization and neural network decomposition, such as spatialDWLS and SPOTlight, and other methods use some specially designed method architectures or loss functions to estimate cell type proportions, including DSTG and Tangram.

[0005] With the development of the data age, deep learning has become increasingly prominent because it attempts to directly obtain high-level features from data. Many organizations have also open-sourced their deep learning frameworks, making the application areas of deep learning increasingly broad, and more and more deep learning models have been introduced into bioinformatics data analysis.

[0006] ResNet is a deep convolutional network proposed in 2015. As soon as it was born, it won the championship in image classification, detection, and positioning in ImageNet. ResNet is a milestone model in the history of deep learning. Before the emergence of ResNet, deep learning models were generally between 20 and 30 layers, but after the emergence of ResNet, the number of layers of deep learning models was increased to more than 100 layers, which achieved a qualitative improvement. It is not uncommon to use ResNet to process bioinformatics data, and it can achieve good performance. Currently, ResNet is used to process bioinformatics data mainly in the following two forms: when the data is in the form of images, such as pathological sections, pathologists first mark the tumor area and normal area on the image, and then use the sliding window to traverse and crop the ROI into non-overlapping blocks of K*K pixels. The processed images are then used to train the model to achieve the corresponding purpose, such as cancer type classification, cancer prognosis prediction, etc.

[0007] When the original data is in a non-image format such as The Cancer Genome Atlas (TCGA), we first construct an N*N gene expression image based on the chromosomal location of the gene and the number of expressed genes. Then, we perform subsequent targeted processing and analysis based on the image. The construction idea is as follows: Assuming that the size of the mutation map is N*N, all expressed genes are collected, grouped according to their location on the chromosome, and located on the matrix map. First, each gene is sorted according to its location on the chromosome (chromosomes 1-22, X and Y). For cancer j, the list of mutated genes on chromosome i is r ij Then different genes in the same chromosome are grouped according to their positions. The genome R collected from chromosome i i , length L i Therefore, the number of columns occupied by genes on chromosome i in the gene expression map is k i ,in All genes on the chromosome occupy K columns, (Where 23 is based on the 23 pairs of human chromosomes as an example), according to the above description, we can choose a suitable N value, and then arrange the genes according to their position and order on chromosomes 1-22, X, and Y to obtain an N*N gene expression image. In 2021, the gene mutation data will be converted into a gene mutation map. For each patient sample, a gene mutation map will be constructed to record the gene mutation status and chromosome position information, and the ResNet-50 network will be used for genomic pan-cancer classification. In 2022, Cox-ResNet will also use this idea to construct gene expression images. Gene pixels are arranged vertically according to their position on the chromosome, and then survival analysis is performed. In order to improve the accuracy of cellular analysis of spatial transcriptome data, in the absence of breakthroughs in single-cell RNA sequencing technology and spatial transcription technology, the improved SSD-ResNet model is used to extract and analyze features of spatial transcriptome data, which can obtain better cell analysis results, and enable the original single-output ResNet to output multiple categories and their probability distributions by embedding multi-target detection, making the application field of ResNet broader. Summary of the Invention

[0008] In order to overcome the above-mentioned shortcomings of the prior art in the absence of breakthroughs in single-cell RNA sequencing technology and spatial transcription technology, the present invention adopts an improved SSD-ResNet model to extract and analyze features of spatial transcriptome data, which can obtain better cell analysis results. In addition, the original single-output ResNet can output multiple categories and their probability distributions by embedding multi-target detection, providing a spatial cell type analysis method based on the improved SSD-ResNet.

[0009] In order to solve the above technical problems, the technical solutions of the present invention are as follows:

[0010] A spatial cell type parsing method based on an improved SSD-ResNet comprises the following steps:

[0011] S1: Obtain single-cell data and spatial transcriptome data of the same tissue;

[0012] S2: Preprocess the single-cell data and spatial transcriptome data of the same tissue to obtain a heat map of gene expression;

[0013] S3: Build a ResNet network, add SSD to the ResNet network to obtain an SSD-ResNet model;

[0014] S4: Using the gene expression heat map image obtained in step S2 to train the SSD-ResNet model parameters, so that the loss function converges, a trained SSD-ResNet model is obtained;

[0015] S5: Use the trained SSD-ResNet model to perform cell type analysis.

[0016] Furthermore, in step S2, the single cell data and spatial transcriptome data are preprocessed, specifically:

[0017] S2.1: Construct single-cell gene expression profiles for single-cell data and perform conditional screening on single-cell data;

[0018] S2.2: Obtain genes that are expressed in both single-cell data and spatial transcriptome data after screening in the gene expression profile in step S2.1;

[0019] S2.3: Construct a heat map image based on the genes expressed in both the single-cell data and the spatial transcriptome data after filtering in step S2.2.

[0020] Furthermore, the ResNet network also introduces a residual network structure, which adds a direct forward feedback connection between the input layer and the output layer of the ResNet network.

[0021] Furthermore, in step S2.1, the single-cell data are conditionally screened, specifically: single-cell data with nUMI greater than 1500, counts greater than 500, and mitochondrial gene percentage less than 10% are retained.

[0022] Furthermore, the SSD includes: the input of the SSD passes through six feature prediction layers respectively, then reaches the detector, and finally reaches the classifier.

[0023] Furthermore, the SSD-ResNet model is specifically:

[0024] The SSD-ResNet model includes a basic convolution module, an auxiliary convolution module and a prediction convolution module. The basic convolution module includes multiple ResNet networks, which are connected in sequence. The auxiliary convolution module includes multiple convolution layers. The prediction convolution module includes a classifier and a detector. The auxiliary convolution module is connected to the basic convolution module in sequence. The outputs of the basic convolution module and the auxiliary convolution module are both input into the prediction convolution module. The thermal map image is first normalized after passing through the first ResNet network of the basic convolution module, and then input into the prediction convolution module. The output of the first ResNet network is input into other ResNet network layers and then into the prediction convolution module. The output of the basic convolution module is input into the auxiliary convolution module and then reaches the prediction convolution module. Finally, the output of the prediction convolution module is summarized into the non-maximum suppression layer and the result is output.

[0025] Furthermore, in step S4, the heat map image of gene expression obtained in step S2 is used to train the SSD-ResNet model parameters so that the loss function converges. The model parameters include the number of prior boxes and the scale of the prior box settings.

[0026] Furthermore, the heat map image of gene expression obtained in step S2 is used to train the SSD-ResNet model parameters, specifically:

[0027] S4.1: Find the box with the largest IOU with the actual target in the image, and set this box as a positive sample. The box that does not match the actual target is a negative sample.

[0028] S4.2: Set the box whose actual target IOU is higher than the preset threshold as the matching positive sample;

[0029] S4.3: Arrange and sample negative samples according to background confidence;

[0030] S4.4: Input positive and negative samples into the loss function and train the SSD-ResNet model.

[0031] Furthermore, in step S4.4, the loss function is specifically:

[0032]

[0033] Among them, x is the indicator parameter, α is the weighting coefficient, L is the loss function, c is the category confidence prediction value, l is the location prediction value of the prior frame, g is the location parameter of the actual target, P is the number of positive samples of the prior frame, L conf is the confidence loss, L loc is the position loss.

[0034] Furthermore, the trained SSD-ResNet model is used to perform cell type analysis, specifically:

[0035] The trained SSD-ResNet model is used to perform cell analysis on the spatial transcriptome data homologous to the single-cell gene expression profile data to obtain the spatial cell composition results.

[0036] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0037] The ResNet network solves the network degradation problem of deep networks. Its application in cell analysis helps to improve the accuracy of cell analysis of spatial transcriptome data. A new data preprocessing method is proposed based on data characteristics. Since the two data dimensions are too different, single-cell data may have tens of thousands of samples, while transcriptome data only has hundreds of samples. The present invention also screens and retains genes to reduce data sparsity. ResNet is embedded in SSD to obtain the SSD-ResNet model, which can perform multiple types of outputs and their probability distributions, making the network application field that originally output a traditional single result broader. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a flow chart of a spatial cell type parsing method based on improved SSD-ResNet according to the present invention;

[0039] Figure 2 Flowchart for preprocessing single-cell data and spatial transcriptome data provided by an embodiment of the present invention;

[0040] Figure 3 A diagram of the residual network structure provided by an embodiment of the present invention;

[0041] Figure 4 A block diagram of the SSD structure provided by an embodiment of the present invention;

[0042] Figure 5 8*8*k SSD feature map provided by the embodiment of the present invention;

[0043] Figure 6 4*4*k SSD feature map provided by the embodiment of the present invention;

[0044] Figure 7 This is a block diagram of the SSD-ResNet structure described in the present invention. DETAILED DESCRIPTION

[0045] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;

[0046] In order to better illustrate this embodiment, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product size;

[0047] It is understandable to those skilled in the art that some well-known structures and descriptions thereof may be omitted in the drawings.

[0048] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0049] Example 1

[0050] A spatial cell type parsing method based on improved SSD-ResNet, such as Figure 1 As shown, the following steps are included:

[0051] S1: Obtain single-cell data and spatial transcriptome data of the same tissue;

[0052] S2: Preprocess the single-cell data and spatial transcriptome data of the same tissue to obtain a heat map of gene expression;

[0053] S3: Build a ResNet network, add SSD to the ResNet network to obtain an SSD-ResNet model;

[0054] S4: Using the gene expression heat map image obtained in step S2 to train the SSD-ResNet model parameters, so that the loss function converges, a trained SSD-ResNet model is obtained;

[0055] S5: Use the trained SSD-ResNet model to perform cell type analysis.

[0056] Obtain public single-cell data and spatial transcriptome data, perform conditional screening on the single-cell data, and then preprocess the single-cell data and spatial transcriptome data to construct an n*n heat map image. Then, embed ResNet into SSD to obtain the SSD-ResNet model. The SSD-ResNet model can perform multiple types of outputs and their probability distributions, making the network application field that originally outputs traditional single results broader. Then improve the model to enable the output of mixed multi-classification results, and obtain the cellular composition of each spot and its probability distribution. Then adjust the model parameters to obtain a trained model. Finally, use the trained model to perform cellular analysis of the spatial transcriptome data. The deep learning network model can better extract abstract features and perform analysis. The ResNet network solves the network degradation problem of deep networks. Its application in cellular analysis helps to improve the accuracy of cellular analysis of spatial transcriptome data.

[0057] Example 2

[0058] This embodiment, based on the first embodiment, further discloses the following contents:

[0059] In step S2, the single cell data and spatial transcriptome data are preprocessed, such as Figure 2 As shown, specifically:

[0060] S2.1: Construct single-cell gene expression profiles for single-cell data and perform conditional screening on single-cell data;

[0061] Obtain publicly available single-cell reference data with cell type annotations and spatial transcriptome data with spatial location information. Because the single-cell data in the present invention is the reference data for subsequent analysis of the spatial transcriptome data, it is necessary to ensure that both are from the same tissue and then preprocess them. In step S2.1, conditional screening is performed on the single-cell data, specifically: retain single-cell data with nUMI greater than 1500, counts greater than 500, and mitochondrial gene percentage less than 10%.

[0062] S2.2: Obtain genes that are expressed in both single-cell data and spatial transcriptome data after screening in the gene expression profile in step S2.1;

[0063] S2.3: Construct a heat map image based on the genes expressed in both the single-cell data and the spatial transcriptome data after filtering in step S2.2.

[0064] In the specific implementation process, if the spatial transcriptome data is directly converted into a gene expression map, and then the single-cell data is constructed into several gene expression maps according to the size of the spatial transcriptome, two problems will arise: First, there will be too few images for training the model. Deep learning requires a large amount of data for learning. When the spatial transcriptome data is constructed into a gene expression profile, the single-cell data must also be constructed according to its size. There may be only dozens of expression profiles, which is too few image data. Second, the pixels of each image are very high, and the space occupied by each sample is very narrow, which is a few hundredths of the width of the entire image. This is equivalent to each sample being a very small target, which is difficult to obtain and will affect the accuracy of subsequent precision experiments. Based on past experience, the effect is better when the n value is not large. Therefore, the following method is proposed:

[0065] Method 1: A single sample constitutes a gene expression map. If the number of genes retained is n gene , then the image size n satisfies (n-1) 2 <n gene ≤n 2 , so that each heat map can be a gene expression map of a sample, and the single-cell data is consistent with the spatial transcriptome data.

[0066] Method 2: Several samples constitute a gene expression map. After comparing the data sets, the number of retained genes is clear. Based on this number, a gene expression map composed of several samples can be constructed to effectively reduce data sparsity. If the number of retained genes is 10,000, a gene expression map with n of 200, 50 rows per sample, and the expression of genes from 4 samples can be constructed. Alternatively, a gene expression map with n of 400, 25 rows per sample, and the expression of genes from 20 samples can be constructed, and so on. Based on the volume of the data, an appropriate n*n gene expression spectrum can be constructed. This can expand the identifiability of each sample, allow the data to be expressed as much as possible, reduce data sparsity, and reduce training time. The above two construction ideas are proposed based on the amount of data. Generally, after screening and processing, single-cell data of more than 10,000, 20,000, or even 30,000 samples will still be retained. Considering the actual training needs, such as training time and computing resources, method 1 can be used when the data is moderate, and method 2 can be considered when the data is too much.

[0067] Example 3

[0068] This embodiment, based on Embodiments 1 and 2, further discloses the following contents:

[0069] The ResNet network also introduces a residual network structure, such as Figure 3 As shown in Figure 2, the residual network structure adds a direct forward feedback connection between the input layer and the output layer of the ResNet network.

[0070] y=F(x,W i )+x

[0071] x l+1 =f(y l )

[0072] Among them, y is the output of this layer, F represents the residual mapping, directly adding x indicates that the input here is the identity mapping, and f is RELU.

[0073] Recursion can get arbitrarily deep units:

[0074]

[0075] Obtain the gradient representation of the reverse process:

[0076]

[0077] The "1" is the result of the identity mapping of the direct connection mechanism, indicating that the direct connection can propagate the gradient losslessly. The residual gradient added later is the result of passing through the weight layer. Due to the existence of "1" in the residual structure, the gradient transmitted in reverse will not be zero. ResNet solves the degradation problem of deep networks through residual learning and can train deeper networks. The more layers of the deep learning network, the richer the features that can be extracted from different layers and the better the effect. However, when the deeper network can begin to converge, a degradation problem is exposed. As the depth of the network increases, the accuracy will tend to saturate and then degenerate rapidly. ResNet introduces a residual network structure to solve the problem of network degradation well, that is, introducing a forward feedback direct connection between the input and output. Direct connections are those that skip one or more layers, and their output is added to the output of the weight layer. Direct connections neither add additional parameters nor increase computational complexity.

[0078] Example 4

[0079] This embodiment, based on Embodiments 1, 2, and 3, further discloses the following contents:

[0080] The SSD includes: Figure 4 As shown in Figure 1, the input of SSD passes through six feature prediction layers, then reaches the detector, and finally reaches the classifier.

[0081] The SSD-ResNet model is specifically:

[0082] SSD-ResNet model, such as Figure 7As shown, it includes a basic convolution module (extracting low-scale features), an auxiliary convolution module (extracting high-scale features) and a prediction convolution module (outputting the location information and classification information of the feature map). The basic convolution module includes multiple ResNet networks, which are connected in sequence. The auxiliary convolution module includes multiple convolution layers. The prediction convolution module includes a classifier and a detector. The auxiliary convolution module is connected to the basic convolution module in sequence. The outputs of the basic convolution module and the auxiliary convolution module are input into the prediction convolution module. The thermal map image is first normalized after passing through the first ResNet network of the basic convolution module so that the features have a consistent scale and range, and then input into the prediction convolution module. The output of the first ResNet network is input into other ResNet network layers and then into the prediction convolution module. The output of the basic convolution module is input into the auxiliary convolution module and then reaches the prediction convolution module. Finally, the output of the prediction convolution module is summarized into the non-maximum suppression layer, the repeated boxes of the same type are deleted, and the box with the highest score is selected and Make the output and obtain the final positioning result. It should be noted that corresponding adjustments and fine-tuning should be made according to the specific problem requirements and data sets, such as adjusting the scale and aspect ratio of the default box. This can ensure that the SSD model can achieve the best detection performance when using ResNet as a feature extractor. After the above conventional network is built, it will be found that the network outputs only a certain class that the network believes to have the highest confidence, rather than a mixture of multiple classes. SSD usually includes a classification branch and a regression branch. The classification branch is used to predict the category of the target, and the regression branch is used to locate the target box. In order to output multiple classification results, the classification branch needs to be modified. The classification prediction layer usually uses a convolutional layer with an output channel number of num_anchors*(num_classes+1), where num_anchors represents the number of bounding boxes at each spatial position, num_classes represents the number of object categories (excluding background), and +1 is to add an additional category to represent the background. The specific modification steps are as follows:

[0083] Modify the number of output channels of the classification prediction layer: set the number of output channels to num_anchors*(num_classes*n), where n represents the number of categories predicted for each target, use the reshape operation: reshape the output of the classification prediction layer, divide the original channel dimension into num_classes dimension and num_anchors dimension, and perform a softmax operation on the classification prediction result: perform a softmax operation on the classification result of each target to obtain the probability of each category. In this way, an SSD network based on ResNet can be obtained, including components such as feature extraction, multi-scale convolution, normalization, positioning and classification prediction, and forward propagation and loss calculation are implemented. SSD-ResNet is a deep learning model that can perform mixed multi-classification and its confidence output. ResNet output is generally directly a single prediction result, but the current spatial transcription technology does not achieve the accuracy of single cells. Each spot contains multiple cells, and the spatial resolution also changes dynamically depending on the spatial transcription technology used. Therefore, this single type of output result does not conform to the actual situation and cannot achieve the expected purpose. Therefore, the ResNet output must also be adjusted. A common problem encountered in target detection is that its training model may have zero to dozens of categories, and the model may output multiple prediction results. The classifier will take the image as input and generate a probability distribution of the predicted class. Each spot can obtain the possible cell type and its probability distribution.

[0084] In the specific implementation process, the convolutional network generally has larger feature maps at the front, and then reduces the feature maps by setting the step size or pooling. Regardless of the size, these feature maps are used for detection. For example, the first feature layer can be used to detect smaller targets. As the abstract feature extraction capability increases, the subsequent prediction layer can detect larger targets. Figure 5 and 6 The two pictures on the left and right are feature maps of different sizes. For multiple feature maps at the top of the network, a set of default bounding boxes are associated with each feature map unit. The default bounding boxes tile the feature maps in a convolutional manner so that the position of each bounding box relative to its corresponding cell is fixed, such as Figure 5 and 6 The three sets of bounding boxes shown in each figure and the positions of their corresponding cells are the benchmarks of the predicted bounding boxes. Assuming that the feature map size is m*h and each cell has k boxes, Figure 5 There are 8*8*k boxes, Figure 6 There are 4*4*k boxes, and each box needs to be classified and regressed. The number of convolution kernels used for classification is ck (c represents the number of categories), and the number of convolution kernels for regression is 4k. Then each box requires (c+4)k prediction values, and the feature map requires (c+4)kmh prediction values.

[0085] Example 5

[0086] This embodiment, based on Embodiments 1, 2, 3, and 4, further discloses the following contents:

[0087] After the single-cell data is preprocessed, the data set can be divided into training data and test data, and then the training data is used to train the SSD-ResNet classification model. As the training progresses, the various parameters are appropriately adjusted to obtain the optimal model. In step S4, the SSD-ResNet model parameters are trained using the gene expression heat map image obtained in step S2 to make the loss function converge. The model parameters include the number of prior boxes and the scale of the prior box settings. During the training process, we need to determine which bounding boxes correspond to the ground truth (actual target) and train the network accordingly. The SSD-ResNet model parameters trained using the gene expression heat map image obtained in step S2 are specifically:

[0088] S4.1: Find the box with the largest IOU with the actual target in the image, and set this box as a positive sample. The box that does not match the actual target is a negative sample.

[0089] S4.2: Set the box whose actual target IOU is higher than the preset threshold as the matching positive sample;

[0090] S4.3: Arrange and sample negative samples according to background confidence;

[0091] S4.4: Input positive and negative samples into the loss function and train the SSD-ResNet model.

[0092] To avoid excessive imbalance between positive and negative samples, we can set an IOU above a threshold to be considered a match. To better balance positive and negative samples, we will sort and sample negative samples according to the background confidence, ensuring a positive-to-negative sample ratio of about 1:3. We then train according to the loss function, which is a weighted sum of the location loss (loc) and the confidence loss (conf):

[0093]

[0094] Where x is an indicator parameter, such as The indicator for matching the i-th default box with the actual target of the j-th category a, c is the category confidence prediction value, l is the position prediction value of the prior box, g is the position parameter of the actual target, P is the number of positive samples of the prior box, if P = 0, set the loss to 0,

[0095]

[0096]

[0097]

[0098]

[0099]

[0100] The position loss is a smooth L1 loss between the predicted box (l) and the actual target (g) parameters, regressing to the center (cx, cy) of the default bounding box (d) and its width (w) and height (h) offsets, which is used to measure the difference between the position and size of the predicted box and the real box.

[0101]

[0102]

[0103] The confidence loss is a softmax loss on the multi-class confidence (c), which is used to measure the difference between the category probability of the predicted box and the true box category.

[0104] In step S4.4, the loss function is specifically:

[0105]

[0106] Among them, x is the indicator parameter, α is the weighting coefficient, L is the loss function, c is the category confidence prediction value, l is the location prediction value of the prior frame, g is the location parameter of the actual target, P is the number of positive samples of the prior frame, L conf is the confidence loss, L loc is the position loss.

[0107] Use the trained SSD-ResNet model to perform cell type analysis, specifically:

[0108] The trained SSD-ResNet model is used to perform cell analysis on the spatial transcriptome data homologous to the single-cell gene expression profile data to obtain the spatial cell composition results.

[0109] In the specific implementation process, since the cells in each spot of the actual spatial transcriptome data are unknown, the virtual spatial transcriptome dataset constructed by the experimenter himself must be used here first. Since it is a dataset constructed by himself, the cell composition of the virtual spatial transcriptome dataset can be known at this time, and the effect of the model can be obtained by comparing the cell composition of the spot with the output result. Finally, the model is applied to the actual spatial transcriptome data. Although it is impossible to know whether the result is completely correct, the spatial transcriptome data itself contains the spatial position information of each spot. The experimenter can obtain the spatial distribution of cells based on the cell type output by each spot. Comparing this distribution with the actual tissue spatial structure can obtain the feasibility of the model. The trained model is used to perform cell analysis on the spatial transcriptome data homologous to the single-cell gene expression spectrum data to obtain the spatial cell composition result. This result should reflect the cell type that each spot may contain in space. Depending on the situation, the output of a certain cell, two cells or more cell types depends on the specific threshold setting.

[0110] The same or similar reference numerals correspond to the same or similar components;

[0111] The terms used in the drawings to describe positional relationships are for illustrative purposes only and should not be construed as limiting this patent;

[0112] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A spatial cell type parsing method based on improved SSD-ResNet, characterized in that: The following steps are involved: S1: Obtain single-cell data and spatial transcriptome data of the same tissue; S2: Preprocess the single-cell data and spatial transcriptome data of the same tissue to obtain a heat map of gene expression, specifically: S2.1: Construct single-cell gene expression profiles for single-cell data and perform conditional screening on single-cell data; S2.2: Obtain genes that are expressed in both single-cell data and spatial transcriptome data after screening in the gene expression profile in step S2.1; S2.3: Genes expressed in both the single-cell data and the spatial transcriptome data after filtering in step S2.2 are used to construct a heat map image. S3: Build a ResNet network, add SSD to the ResNet network, and obtain an SSD-ResNet model. The SSD-ResNet model is specifically as follows: The SSD-ResNet model includes a basic convolution module, an auxiliary convolution module and a prediction convolution module. The basic convolution module includes multiple ResNet networks, which are connected in sequence. The auxiliary convolution module includes multiple convolution layers. The prediction convolution module includes a classifier and a detector. The auxiliary convolution module is connected to the basic convolution module in sequence. The outputs of the basic convolution module and the auxiliary convolution module are both input into the prediction convolution module. The thermal map image is first normalized after passing through the first ResNet network of the basic convolution module, and then input into the prediction convolution module. The output of the first ResNet network is input into other ResNet network layers and then into the prediction convolution module. The output of the basic convolution module is input into the auxiliary convolution module and then reaches the prediction convolution module. Finally, the output of the prediction convolution module is summarized into the non-maximum suppression layer and the result is output. S4: Using the gene expression heat map image obtained in step S2 to train the SSD-ResNet model parameters, so that the loss function converges, a trained SSD-ResNet model is obtained; S5: Use the trained SSD-ResNet model to perform cell type analysis.

2. A spatial cell type parsing method based on improved SSD-ResNet according to claim 1, characterized in that: The ResNet network also introduces a residual network structure, which adds a direct forward feedback connection between the input layer and the output layer of the ResNet network.

3. The spatial cell type parsing method based on improved SSD-ResNet according to claim 1, characterized in that: In step S2.1, single-cell data are conditionally screened, specifically: single-cell data with nUMI greater than 1500, counts greater than 500, and mitochondrial gene percentage less than 10% are retained.

4. The spatial cell type parsing method based on improved SSD-ResNet according to claim 1, characterized in that: The SSD includes: the input of the SSD passes through six feature prediction layers, then reaches the detector, and finally reaches the classifier.

5. The spatial cell type parsing method based on improved SSD-ResNet according to claim 1, characterized in that: In step S4, the heat map image of gene expression obtained in step S2 is used to train the SSD-ResNet model parameters so that the loss function converges. The model parameters include the number of prior boxes and the scale of the prior box settings.

6. The spatial cell type parsing method based on improved SSD-ResNet according to claim 5, characterized in that: The SSD-ResNet model parameters are trained using the heat map image of gene expression obtained in step S2, specifically: S4.1: Find the box with the largest IOU with the actual target in the image, and set this box as a positive sample. The box that does not match the actual target is a negative sample. S4.2: Set the box whose actual target IOU is higher than the preset threshold as the matching positive sample; S4.3: Arrange and sample negative samples according to background confidence; S4.4: Input positive and negative samples into the loss function and train the SSD-ResNet model.

7. The spatial cell type parsing method based on improved SSD-ResNet according to claim 6, characterized in that: In step S4.4, the loss function is specifically: Among them, x is the indicator parameter, α is the weighting coefficient, L is the loss function, c is the category confidence prediction value, l is the location prediction value of the prior frame, g is the location parameter of the actual target, P is the number of positive samples of the prior frame, L conf is the confidence loss, L loc is the position loss.

8. The spatial cell type parsing method based on improved SSD-ResNet according to claim 1, characterized in that: The trained SSD-ResNet model is used to perform cell type analysis. Specifically, the trained SSD-ResNet model is used to perform cell analysis on the spatial transcriptome data homologous to the single-cell gene expression profile data to obtain the spatial cell composition results.

Citation Information

Patent Citations

  • Single cell sequencing gene expression data interpolation method and system based on deep learning

    CN115394358A

  • Cancer patient specimen property distinguishing method based on optical image and application

    CN115984231A