Cell region determination method and device, electronic device, and storage medium
By obtaining gene expression information from spatial transcription data and calculating target distribution parameters and adjacent site probabilities, the problem of inaccurate cell region determination in existing technologies is solved, achieving higher accuracy and universality.
Patent Information
- Application Number
- CN202311079104.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-23
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-08-23
AI Technical Summary
In existing technologies, cell region determination methods based on staining images have low average RNA capture rates, while methods based on partial RNA signal delineation are not suitable for spatial omics data with sparse RNA signal amounts, and there is a lack of methods to accurately determine cell regions.
Gene expression information is obtained from the spatial transcription data of the target slice. The target distribution parameters of gene expression at each spatial sequencing site are calculated to ensure that the gene expression level follows a preset distribution. The probability of cell coverage area is calculated based on the target distribution parameters of adjacent sites. Cell regions are determined by combining the location data.
It improves the accuracy and universality of cellular region identification, reduces dependence on specific genes, and enhances the consistency between RNA signaling and cellular region identification results.
Smart Images

Figure CN119560023B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of cell segmentation technology, specifically to a method and apparatus for determining cell regions, electronic equipment, and storage medium. Background Technology
[0002] Currently, identifying cellular regions is a crucial step in space-based cell research. Related techniques for identifying cellular regions include methods based on staining images obtained using techniques such as Heterochromatography (H&E) and methods that delineate cells using partial RNA signals. However, methods based on staining images rely solely on image features and do not consider RNA signals, resulting in a low average RNA capture amount in the cells after segmentation, which is detrimental to downstream analysis of space omics data. Methods that delineate cells based on partial RNA signals require a high level of a specific RNA signal, meaning they depend on the capture amount of a particular gene. Therefore, these methods are unsuitable for sequencing-based space omics data where RNA signals are sparse. Thus, a method that can accurately identify cellular regions and achieve reasonable cell segmentation is currently lacking. Summary of the Invention
[0003] This disclosure provides a method and apparatus for determining cell regions, an electronic device, and a storage medium, which can improve the accuracy of determining cell regions.
[0004] To achieve the above objectives, a first aspect of this application provides a method for determining cell regions, the method comprising:
[0005] Gene expression level information is obtained from the spatial transcription data of the target slice; wherein, the gene expression level information includes the location data of multiple spatial sequencing sites and the gene expression level of each spatial sequencing site;
[0006] Calculate the target distribution parameters corresponding to the gene expression level of each of the spatial sequencing sites following a preset distribution;
[0007] The first probability that each spatial sequencing site belongs to a cell-covered region is calculated based on the target distribution parameters of the adjacent sites of each spatial sequencing site; wherein, the adjacent sites of the spatial sequencing site are the spatial sequencing sites adjacent to the spatial sequencing site;
[0008] The cell regions of the target slice are determined based on the first probability and the location data.
[0009] In some embodiments, a first log-likelihood function of the preset distribution is constructed based on the gene expression level at each of the spatial sequencing sites;
[0010] The latent variables are determined based on the first log-likelihood function, and the target steps are executed iteratively. The target steps include: calculating the conditional probability of the latent variables based on the randomly initialized original distribution parameters; maximizing the first log-likelihood function based on the conditional probability to obtain preliminary distribution parameters; if the number of executions of the target steps reaches a preset number, the preliminary distribution parameters at the preset number of executions will be used as the target distribution parameters, or if the preliminary distribution parameters converge, the preliminary distribution parameters will be used as the target distribution parameters.
[0011] In some embodiments, constructing the first log-likelihood function of the preset distribution based on the gene expression level at each of the spatial sequencing sites includes:
[0012] Probabilistic expression data of the preset distribution are constructed based on the gene expression levels at each of the spatial sequencing sites;
[0013] The state variables of the spatial sequencing sites are constructed based on the probability expression data; wherein, the state variables are used to indicate whether the spatial sequencing sites belong to cell-covered regions or non-cell-covered regions;
[0014] The distribution variables of the preset distribution are constructed based on the probability expression data; wherein, the distribution variables are used to indicate whether the type of the preset distribution is a cell-covered region type or a non-cell-covered region type;
[0015] The first log-likelihood function is constructed based on the state variable and the distribution variable.
[0016] In some embodiments, the conditional probability includes cell region probability, first variable probability, and second variable probability, and the original distribution parameters include first distribution parameters and second distribution parameters;
[0017] The step of calculating the conditional probability of the latent variable based on the randomly initialized original distribution parameters includes:
[0018] The probability that the spatial sequencing site belongs to the cell coverage area is calculated based on the preset conditions and the original distribution parameters to obtain the cell region probability.
[0019] The probability of the first variable is obtained by performing a logarithmic calculation based on the first distribution parameter.
[0020] The probability of the second variable is obtained by performing logarithmic differentiation based on the second distribution parameters.
[0021] In some embodiments, the step of maximizing the first log-likelihood function according to the conditional probability to obtain preliminary distribution parameters includes:
[0022] The probabilities of the cell regions are summed to obtain the probability sum;
[0023] Calculate the sum of the products of the cell region probability and the second variable probability, and then calculate the ratio of the sum of the products to the cumulative value of the cell region probabilities to obtain the first probability ratio.
[0024] The second probability ratio is calculated based on the cell region probability, the first variable probability, and the second variable probability.
[0025] The expected function of the first log-likelihood function is calculated, and the expected function is maximized based on the sum of probabilities, the first probability ratio, and the second probability ratio to obtain the preliminary distribution parameters.
[0026] In some embodiments, determining the target distribution parameters corresponding to the gene expression level of each of the spatial sequencing sites following a preset distribution includes:
[0027] A second probability density function of the preset distribution is constructed based on the gene expression level at each of the spatial sequencing sites;
[0028] Obtain observation data of the spatial sequencing site belonging to the cell coverage area in a preset number of times, and construct a second likelihood function based on the observation data and the second probability density function;
[0029] Logarithmic transformation of the second likelihood function yields the log-likelihood function;
[0030] The target distribution parameters are obtained by differentiating the log-likelihood function.
[0031] In some embodiments, determining the first probability that the spatial sequencing site belongs to a cell-covered region based on the target distribution parameters of the adjacent sites includes:
[0032] Calculate the adjacency message data between the adjacent sites and the spatial sequencing sites based on the target distribution parameters;
[0033] The probability of determining that the spatial sequencing site belongs to the cell-covered region is determined based on the adjacency message data.
[0034] In some embodiments, calculating the adjacency message data between the adjacent sites and the spatial sequencing sites based on the target distribution parameters includes:
[0035] Calculate the conditional probability that the spatial sequencing site belongs to the cell coverage region based on the target distribution parameters of the spatial sequencing site;
[0036] The first potential difference of the adjacency message is calculated based on the first state data of the adjacent sites, the target distribution parameters of the adjacent sites, and a preset potential function; wherein, the first state data indicates whether the adjacent sites belong to cell-covered or non-cell-covered regions.
[0037] The adjacency message data is calculated based on the conditional probability and the first potential difference.
[0038] In some embodiments, determining the first probability that the spatial sequencing site belongs to a cell-covered region based on the adjacency message data includes:
[0039] The total message data is obtained by summing the multiple adjacent message data.
[0040] The second potential difference of the spatial sequencing site is calculated based on the second state data of the spatial sequencing site and the target distribution parameters of the spatial sequencing site; wherein, the second state data indicates whether the spatial sequencing site belongs to a cell-covered region or a non-cell-covered region;
[0041] The probability that the spatial sequencing site belongs to the cell-covered region is calculated based on the total message data and the second potential difference.
[0042] In some embodiments, determining the cell region of the target slice based on the first probability and the location data includes:
[0043] A cell mask matrix is constructed based on the first probability and the location data, and a spatial counting matrix is constructed based on the gene expression level information;
[0044] The cell regions of the target slice are determined based on the cell mask matrix and the spatial counting matrix.
[0045] In some embodiments, constructing a cell mask matrix based on the first probability and the location data includes:
[0046] The second probability that the spatial sequencing site belongs to a non-cell-covered region is determined based on the target distribution parameters of the adjacent sites;
[0047] The first probability is compared with the second probability, and the cell mask matrix is constructed based on the comparison result, the preset mask value, and the position data.
[0048] In some embodiments, the step of determining the cell region of the target slice based on the cell mask matrix and the spatial counting matrix includes:
[0049] Convolution kernels are constructed based on preset cell diameters, and convolution calculations are performed based on the convolution kernels and the spatial counting matrix to obtain a convolution matrix;
[0050] A cell tag matrix is obtained by performing a dot product between the convolution matrix and the cell mask matrix; wherein the row and column dimensions of the cell tag matrix are the same as those of the spatial counting matrix, the cell tag matrix includes the location data of the spatial sequencing sites and the cell tags corresponding to the spatial sequencing sites, and the cell tags include cell region tags;
[0051] The cell regions of the target slice are determined based on the location data corresponding to the cell region labels.
[0052] To achieve the above objectives, a second aspect of this application provides a cell region determination device, the device comprising:
[0053] An information acquisition unit is used to acquire gene expression level information from the spatial transcription data of the target slice; wherein, the gene expression level information includes the location data of multiple spatial sequencing sites and the gene expression level of each spatial sequencing site;
[0054] The distribution parameter calculation unit is used to calculate the target distribution parameter corresponding to the gene expression level of each spatial sequencing site following a preset distribution.
[0055] A probability calculation unit is used to calculate a first probability that each spatial sequencing site belongs to a cell-covered region based on the target distribution parameters of the adjacent sites of each spatial sequencing site; wherein, the adjacent sites of the spatial sequencing site are spatial sequencing sites adjacent to the spatial sequencing site.
[0056] A cell region determination unit is used to determine the cell region of the target slice based on the first probability and the location data.
[0057] To achieve the above objectives, a third aspect of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0058] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0059] To achieve the above objectives, a fifth aspect of this application provides a computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the method described in the first aspect.
[0060] The cell region determination method, apparatus, electronic device, and storage medium proposed in this application calculate a preset distribution parameter based on the gene expression level of each spatial sequencing site. Based on the target distribution parameter of adjacent sites, a first probability is calculated that the spatial sequencing site belongs to a cell-covered region. Then, the cell region of the target slice is determined based on the first probability and the location data of the spatial sequencing site. Therefore, this application example determines the cell region of the target slice entirely based on gene expression levels. Compared to methods based on staining images in related technologies, this application embodiment directly processes the gene expression information obtained from spatial transcription data, reducing the possibility of low RNA signal levels in the delineated cell regions and enhancing the consistency between RNA signal and cell region determination results. Compared to methods that determine cell regions based on partial RNA, this application embodiment reduces dependence on specific genes and improves the universality of the cell region determination method. Attached Figure Description
[0061] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.
[0062] Figure 1 This is a flowchart of the cell region determination method provided in the embodiments of this application;
[0063] Figure 2 yes Figure 1 The flowchart of step S102 in the document;
[0064] Figure 3 yes Figure 2 The flowchart of step S201 in the text;
[0065] Figure 4 yes Figure 2 The flowchart for calculating the conditional probability of the latent variable in step S202;
[0066] Figure 5 yes Figure 2 The flowchart for calculating the preliminary distribution parameters in step S202;
[0067] Figure 6 yes Figure 1 A flowchart of another embodiment of step S102;
[0068] Figure 7 yes Figure 1 The flowchart of step S103 in the process;
[0069] Figure 8 yes Figure 7 The flowchart of step S701 in the process;
[0070] Figure 9 yes Figure 7 The flowchart of step S702 in the process;
[0071] Figure 10 yes Figure 1 The flowchart of step S104 in the process;
[0072] Figure 11 yes Figure 10 The flowchart of step S1001 in the text;
[0073] Figure 12 yes Figure 10 The flowchart of step S1002 in the document;
[0074] Figure 13 This is a schematic diagram of a cell region determined by artificial labeling based on RNA signals;
[0075] Figure 14 This is a schematic diagram of a cell region determined using the method provided in the embodiments of this application;
[0076] Figure 15 It is a schematic diagram of the cell region determined using image-referenced methods (such as the Stardist method);
[0077] Figure 16 This is a schematic diagram of a cell region determined by a method that delineates cell regions based on partial RNA signals (such as the Baysor method);
[0078] Figure 17 This is a schematic diagram of the cell region determination device provided in the embodiments of this application;
[0079] Figure 18 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0080] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.
[0081] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations:
[0082] Spatial transcriptomics (ST) sequencing is a technique for spatially analyzing RNA, used to analyze all RNA in a single tissue slice. 10x Genomics spatial transcriptomics uses four capture regions per slide for library construction, each measuring 6.5 x 6.5 mm. Each capture region contains 5000 barcoded spots, each 55 μm in diameter, with a center-to-center distance of 100 μm between spots. Each spot includes multiple capture probes that bind to RNA, each with unique spatial barcodes to mark the spatial location of the captured RNA. Sequencing allows mapping the transcribed sequence of each RNA back to its original location in the tissue slice, providing data support for downstream tasks. Therefore, ST sequencing is a technique capable of simultaneously acquiring RNA expression levels and RNA spatial location information in a single experiment.
[0083] Spatial Transcriptomics with Error-Corrected Barcoding and Sequencing (Stereo-seq) encompasses the design of patterned DNA nanosphere (DNB) arrays, in-situ sequencing to determine the spatial coordinates of uniquely barcoded oligonucleotides, point ligation of UMI-polyT oligonucleotides, in-situ RNA capture from tissues, cDNA amplification, library construction and sequencing, and data analysis. DNB sequencing, based on patterned arrays used in-situ sequencing, serves as the foundation for transcription techniques with high-resolution and large-field-of-view spatial resolution. Compared to other sequencing methods such as Stereo-seq, HDST (a high-resolution spatial transcriptomics technology that uses numerous 2.05 μm pores carved into a slide and silica magnetic beads distributed within these pores to capture mRNA from cells at corresponding locations for reverse transcription, library construction, and transcriptome sequencing), Slide-seqV2 (a high-sensitivity near-cellular resolution spatial transcriptomics analysis that achieves better RNA capture efficiency through magnetic bead synthesis and array indexing), Visium (a spatial transcriptomics sequencing technology that integrates gene expression with immunohistochemical images of tissue sections to locate gene expression levels in different cells within a tissue to their original spatial locations), and DBiT-seq (a microfluidic barcode-based spatial multi-omics sequencing technology), Stereo-seq offers higher resolution per 100 μm pores. 2With more light spots and smaller spot size and center-to-center distance, Stereo-seq has higher resolution.
[0084] Currently, for subcellular resolution spatial transcriptome sequencing technology, identifying cellular regions in tissue sections is a crucial step in spatial cell research. Related technologies employ two main methods for identifying cellular regions: the first method uses staining images obtained through techniques such as H&E to determine cellular regions; the second method delineates cellular regions based on partial RNA analysis.
[0085] The methods based on stained images include the following two types: First, using machine learning models such as feature extraction from computer vision and random forests (e.g., the Ilastik algorithm, a machine learning-based image segmentation and analysis method) to identify stained images and determine cell regions. Second, using deep learning models based on convolutional neural networks (e.g., the DeepCell model and the Stardist model, a model for object detection and segmentation) to identify stained images and determine cell regions.
[0086] Methods for delineating cell regions based on partial RNA (such as the Baysor algorithm (which is an algorithm based on the MRF segmentation idea) and the ClusterMap algorithm (which is a clustering algorithm)) use image-based in situ transcription sequencing methods such as MERFISH to stain specific genes of a few genes, and then delineate cells based on the stained images, that is, determine the cell regions.
[0087] Among the various methods mentioned above, machine learning-based methods can automatically detect cellular regions in fluorescently stained images, but they require labeled datasets for model training. In this case, the trained machine learning model may have poor generalization ability to other datasets, meaning that the accuracy of cell region identification is low when applied to other datasets.
[0088] Deep learning-based methods require manual adjustments to the results to achieve satisfactory cell region identification, and the identified cell regions may differ from the actual cell distribution. This is because subsequent experimental acquisitions involving ST processing of images can lead to inconsistencies between the stained images and RNA signals.
[0089] Methods based on partial RNA delineation of cell regions require specific genes to have high capture rates to achieve accurate cell delineation results. Therefore, this method is unsuitable for processing sequencing-based spatial transcriptome data with sparse RNA signal density. Furthermore, this method is time-consuming and memory-intensive.
[0090] Based on this, embodiments of this application propose a method and apparatus for determining cell regions, an electronic device, and a storage medium, which can improve the accuracy of cell region determination.
[0091] The method for determining cell regions provided in the embodiments of this application will be described below.
[0092] Reference Figure 1 In some embodiments, the cell region determination method provided in this application includes, but is not limited to, steps S101 to S104.
[0093] Step S101: Obtain gene expression level information from the spatial transcription data of the target slice; wherein, the gene expression level information includes the location data of multiple spatial sequencing sites and the gene expression level of each spatial sequencing site;
[0094] Step S102: Calculate the target distribution parameters corresponding to the gene expression level of each spatial sequencing site following a preset distribution.
[0095] Step S103: Calculate the first probability that each spatial sequencing site belongs to the cell coverage area based on the target distribution parameters of the adjacent sites of each spatial sequencing site; wherein, the adjacent sites of the spatial sequencing site are the spatial sequencing sites adjacent to the spatial sequencing site.
[0096] Step S104: Determine the cell region of the target slice based on the first probability and location data.
[0097] Steps S101 to S104 of this embodiment involve calculating target distribution parameters for a preset distribution based on the gene expression level of each spatial sequencing site, calculating the first probability that a spatial sequencing site belongs to a cell-covered region based on the target distribution parameters of adjacent sites, and then determining the cell region of the target slice based on the first probability and the location data of the spatial sequencing site. Therefore, this embodiment determines the cell region of the target slice entirely based on gene expression levels. Compared to methods based on staining images in related technologies, this embodiment directly processes the gene expression information obtained from spatial transcription data, reducing the possibility of low RNA signal levels in the delineated cell regions and enhancing the consistency between RNA signal and cell region determination results. Compared to methods that determine cell regions based on partial RNA, this embodiment reduces dependence on specific genes and improves the universality of the cell region determination method.
[0098] In step S101 of some embodiments, the target slice refers to a cryo-embedded tissue slice to be processed by ST to determine cellular regions. The target slice is a biological slice obtained in accordance with relevant legal regulations; biological slices include plant slices and animal slices. ST processing of the target slice yields spatial transcription data, and gene expression information is obtained from the spatial transcriptome data. Gene expression information is used to describe the location data of spatial sequencing sites and the gene expression levels corresponding to those sites. Here, a spatial sequencing site refers to a spot in ST technology; location information indicates the position of the spot in the corresponding capture region; and gene expression level refers to the aggregated value of RNA expression levels captured by the spot from the corresponding location on the target slice. For example, after ST processing, a list can be obtained, describing the sub-expression levels of various genes captured at each spatial sequencing site. The sub-expression levels of various genes are aggregated (i.e., summed) to obtain the gene expression level of the corresponding spatial sequencing site. It is understood that gene expression information can be represented in matrix form, and the corresponding location information can be represented in (x,y) coordinate form; the gene expression level is the value of the (x,y) coordinate in that matrix.
[0099] In step S102 of some embodiments, the preset distribution refers to a pre-set probability distribution, which can be a negative binomial distribution, a Poisson distribution, a Poisson distribution with a zero-expansion term, etc. For ease of explanation, the following embodiments use a negative binomial distribution as an example. The negative binomial distribution is used to describe the relationship between success and failure in an independent event. For example, in the embodiments of this application, whether each spatial sequencing site belongs to a cell-covered region can be regarded as an independent event. According to the negative binomial distribution, the relationship between each spatial sequencing site belonging to a cell-covered region (i.e., success) and belonging to a non-cell-covered region (i.e., failure) can be determined, and the probability of each spatial sequencing site belonging to a cell-covered region can be estimated. It should be understood that adaptive modifications made to the embodiments of this application by setting the preset distribution to other probability distributions should also fall within the protection scope of this application. The target distribution parameter refers to the distribution parameter corresponding to the preset distribution when the gene expression level of each spatial sequencing site follows the preset distribution. It can be understood that the gene expression level of each spatial sequencing site can follow an independent preset distribution, so the target distribution parameter corresponding to following the preset distribution can be calculated based on the gene expression level.
[0100] Reference Figure 2 In some embodiments, step S102 includes, but is not limited to, steps S201 to S202.
[0101] Step S201: Construct the first log-likelihood function of the preset distribution based on the gene expression level of each spatial sequencing site;
[0102] Step S202: Determine the latent variables based on the first log-likelihood function, and repeatedly execute the target step. The target step includes: calculating the conditional probability of the latent variables based on the randomly initialized original distribution parameters; maximizing the first log-likelihood function based on the conditional probability to obtain the preliminary distribution parameters; if the number of executions of the target step reaches a preset number, the preliminary distribution parameters at the preset number of executions will be used as the target distribution parameters, or if the preliminary distribution parameters converge, the preliminary distribution parameters will be used as the target distribution parameters.
[0103] In step S201 of some embodiments, the first log-likelihood function is a function of parameters of a preset distribution. The first log-likelihood function for each spatial sequencing site can be constructed based on the gene expression level of that spatial sequencing site.
[0104] Reference Figure 3 In some embodiments, step S201 includes, but is not limited to, steps S301 to S304.
[0105] Step S301: Construct probability expression data with a preset distribution based on the gene expression level at each spatial sequencing site;
[0106] Step S302: Construct state variables for spatial sequencing sites based on probability expression data; wherein, state variables are used to indicate whether the spatial sequencing site belongs to a cell-covered region or a non-cell-covered region;
[0107] Step S303: Construct a distribution variable for a preset distribution based on the probability expression data; wherein, the distribution variable is used to indicate whether the type of the preset distribution is a cell-covered region type or a non-cell-covered region type;
[0108] Step S304: Construct the first log-likelihood function based on the state variables and distribution variables.
[0109] In step S301 of some embodiments, when the preset distribution is a negative binomial distribution, for each spatial sequencing site, an expression (i.e., probability expression data) can be constructed based on gene expression levels as shown in equation (1).
[0110] P(X|p0,p1,r,r′,θ,θ′)=p1·NB(X|r,θ)+p0·NB(X|r′,θ′)...Equation (1)
[0111] Where X represents the gene expression level at the spatial sequencing site, p1 represents the probability that the gene expression level belongs to the foreground signal expression level, and p0 represents the probability that the gene expression level belongs to the background signal gene expression level. Foreground signal refers to the RNA signal corresponding to the cellular region, while background signal refers to the RNA signal caused by technical noise parameters. Technical noise can include transcript localization errors, gene expression level estimation errors, etc. Transcript localization errors can be caused by factors such as uncertainties in the measurement method and changes in signal intensity. Gene expression level estimation errors can be caused by factors such as noise in the RNA signal and limitations in data processing and analysis methods. r and θ represent the parameters of the preset distribution corresponding to the foreground signal. r′ and θ′ represent the parameters of the preset distribution corresponding to the background signal.
[0112] In steps S302 to S304 of some embodiments, to facilitate calculation and improve the calculation speed of the EM algorithm, dummy variables are constructed based on the expression shown in equation (1) to separate the variables in the expression. Specifically, state variable S, distribution variable M, variable Y, and variable Z are constructed, where variables Y and Z are proposed for ease of calculation and have no practical meaning. State variable S is used to represent the state of the spatial sequencing site, that is, whether the spatial sequencing site belongs to a cell-covered region or a non-cell-covered region. Distribution variable M represents the type of preset distribution, that is, which set of preset distributions the spatial sequencing site follows. The cell-covered region type represents the preset distribution of the set of parameters (r, θ), and the non-cell-covered region type represents the preset distribution of the set of parameters (r′, θ′). Based on state variable S, distribution variable M, variable Y, and variable Z, the first log-likelihood function shown in equation (2) is constructed.
[0113]
[0114] Here, T represents the set of state variables S. When a spatial sequencing site includes multiple states, the set T can be used to represent it. φ k Let represent the k-th set of distribution parameters, which includes two sets of distribution parameters: (r,θ) and (r′,θ′).
[0115] The value of variable Y can be calculated according to the following formula (3), and the value of variable Z can be calculated according to the following formula (4).
[0116]
[0117]
[0118] The advantage of steps S301 to S303 is that, since the log-likelihood function of the classic negative binomial distribution contains a quadratic gamma function, directly using the EM algorithm to calculate the parameters of such a negative binomial distribution will result in the inability to provide an accurate numerical solution, and the EM algorithm calculation consumes a lot of computation time. This embodiment of the application solves the above problems by introducing dummy variables, thereby improving the computation speed of the subsequent EM algorithm.
[0119] In step S202 of some embodiments, the target distribution parameters can be calculated according to the EM algorithm. The EM algorithm includes an E-step (i.e., the expectation step) and an M-step (i.e., the maximization step). The target distribution parameters are obtained by iteratively executing the target step, alternating between the E-step and M-step, until convergence. It can be understood that the target step can be viewed as performing one E-step and one M-step. In the target step, latent variables refer to variables that appear in the first log-likelihood function but cannot be directly observed. In the E-step, the conditional probabilities of the latent variables are calculated based on the parameter values of a randomly initialized preset distribution (i.e., the original distribution parameters). In the M-step, the first log-likelihood function is maximized based on the conditional probabilities calculated in the E-step to update the parameters of the preset distribution, thus obtaining new parameter values for the preset distribution (i.e., the initial distribution parameters). It can be understood that the distribution parameters used in the next E-step are the calculation results of the previous M-step (i.e., the initial distribution parameters). The goal of the M-step is to find the parameter values that maximize the first log-likelihood function. The above target step is repeated until convergence. Convergence refers to the number of repeated executions reaching a preset number, or the initial distribution parameters converging (i.e., the initial distribution parameters stabilizing and no longer changing significantly). The initial distribution parameters corresponding to convergence are used as the target distribution parameters. It is understood that the specific values of the preset number of executions and the randomly initialized original distribution parameters can be adaptively set according to actual needs, and this application embodiment does not specifically limit this.
[0120] Reference Figure 4 In some embodiments, the conditional probability includes cell region probability, first variable probability, and second variable probability, and the original distribution parameters include first distribution parameters and second distribution parameters. Step S202, "calculating the conditional probability of the latent variable based on the randomly initialized original distribution parameters," includes, but is not limited to, steps S401 to S403.
[0121] Step S401: Calculate the probability that the spatial sequencing site belongs to the cell coverage area based on the preset conditions and the original distribution parameters to obtain the cell region probability;
[0122] Step S402: Perform logarithmic calculation based on the first distribution parameter to obtain the probability of the first variable;
[0123] Step S403: Perform logarithmic differential calculation based on the second distribution parameters to obtain the probability of the second variable.
[0124] It should be noted that steps S401 to S403 can be regarded as the E step in the EM algorithm. The specific calculation process of the E step will be explained below.
[0125] In step S401 of some embodiments, the cell region probability refers to the probability that a spatial sequencing site belongs to a cell-covered region. The cell region probability is calculated based on preset conditions, original distribution parameters, and the following equation (5).
[0126] The preset condition refers to a certain value (i.e., x). t ) and a set of references (i.e. The constraints of ) are as follows. h represents each iteration, for example, the original distribution parameters randomly initialized during the first E-step are denoted as θ. (0) r (0) After updating the original distribution parameters based on the initial distribution parameters obtained in step M, the distribution parameters during the second step E are represented as θ. (1) r (1) . Denotes the first distribution parameter. This represents the second distribution parameter.
[0127] In steps S402 to S403 of some embodiments, the first variable probability and the second variable probability have no practical significance; they are used to facilitate the calculation of iterative data introduced for variables Y and Z. The first variable probability is calculated based on the first distribution parameter and the following equation (6). The probability of the second variable is calculated based on the second distribution parameter and the following equation (7).
[0128]
[0129]
[0130] in:
[0131]
[0132]
[0133] Equation (8) represents the differential of the logarithm of the γ function, and Equation (9) represents the conversion relationship between λ and θ and r.
[0134] Reference Figure 5In some embodiments, step S202, "maximizing the first log-likelihood function according to the conditional probability to obtain preliminary distribution parameters", includes, but is not limited to, steps S501 to S504.
[0135] Step S501: Sum the cell region probabilities to obtain the probability sum;
[0136] Step S502: Calculate the sum of the product of the cell region probability and the second variable probability, and calculate the ratio of the sum of the product to the cumulative value of the cell region probability to obtain the first probability ratio.
[0137] Step S503: Calculate the second probability ratio based on the cell region probability, the first variable probability, and the second variable probability;
[0138] Step S504: Calculate the expected function of the first log-likelihood function, and maximize the expected function based on the sum of probabilities, the first probability ratio, and the second probability ratio to obtain the preliminary distribution parameters.
[0139] It should be noted that steps S501 to S504 can be regarded as the M step in the EM algorithm. The specific calculation process of the M step will be explained below.
[0140] In steps S501 to S503 of some embodiments, the cell region probability is determined according to the following formula (10). Perform summation to obtain the probability sum. The cell region probability is calculated according to the following formula (11). With the second variable probability The sum of products, and the sum of products with the cell region probability. The first probability ratio is obtained by calculating the ratio of the accumulated values. Based on cell region probability First variable: probability The second variable is probability. The second probability ratio is calculated using the following formula (12).
[0141]
[0142]
[0143]
[0144] In step S504 of some embodiments, the expectation of the first log-likelihood function is obtained by referring to the following equation (13), thus obtaining the expectation function (i.e., the Q function). The sum of probabilities, the first probability ratio, and the second probability ratio are substituted into the expectation function, and the expectation function is maximized according to the gradient direction of the expectation function to obtain the parameters (i.e., the initial distribution parameters) corresponding to the maximization of the expectation function.
[0145] Q(φ,φ h = E[log(S,M,Y,Z;φ)|x,φ (h) ]=Q(p,p (h) )+Q(φ,φ (h) Equation (13)
[0146] Where Q(p,p) (h) Let Q(φ,φ) represent the Q function with respect to the distribution variable M. (h) Let Q(p,p) represent the Q function with respect to the parameter variable of the predefined distribution. (h) The specific expression is shown in equation (14) below, Q(φ,φ (h) The specific expression is shown in equation (15).
[0147]
[0148]
[0149]
[0150] Reference Figure 6 In some other embodiments, step S102 may include, but is not limited to, steps S601 to S604.
[0151] Step S601: Construct a second probability density function of a preset distribution based on the gene expression level at each spatial sequencing site;
[0152] Step S602: Obtain observation data of spatial sequencing sites belonging to cell coverage areas in a preset number of times, and construct a second likelihood function based on the observation data and the second probability density function;
[0153] Step S603: Perform logarithmic processing on the second likelihood function to obtain the log-likelihood function;
[0154] Step S604: Differentiate the log-likelihood function to obtain the target distribution parameters.
[0155] It should be noted that in steps S601 to S604, the parameters of the preset distribution (i.e., the target distribution parameters) are solved based on the maximum likelihood probability. It is understood that, in addition to the EM algorithm and maximum likelihood probability, other methods such as maximum a posteriori probability can also be used to solve for the target distribution parameters. The specific calculation process of the maximum likelihood probability is explained below.
[0156] In step S601 of some embodiments, a corresponding second probability density function is constructed based on the expression of the negative binomial distribution and the gene expression level at each spatial sequencing site.
[0157] In step S602 of some embodiments, the observation data refers to the results of spatial sequencing sites belonging to cell-covered regions in a preset number of experiments. A result set can be constructed based on the observation data. For example, assuming the preset number of experiments is 10, the result set can be represented as {1,0,1,1,0,1,0,0,1,1}, where 1 represents a result where the spatial sequencing site belongs to a cell-covered region, and 0 represents a result where the spatial sequencing site belongs to a non-cell-covered region. A second likelihood function is constructed based on the result set and the second probability density function.
[0158] In step S603 of some embodiments, in order to simplify the calculation, the logarithm of the second likelihood function is taken to obtain the log-likelihood function.
[0159] In step S604 of some embodiments, the log-likelihood function is maximized to solve for the target distribution parameters of the preset distribution. The method for maximizing the log-likelihood function can be to take the derivative of the log-likelihood function. The target distribution parameters are the parameter values that maximize the log-likelihood function.
[0160] In step S103 of some embodiments, an adjacent site refers to a spot near a spatial sequencing site. "Adjacent" means that the distance between the adjacent site and the spatial sequencing site is within a preset distance range. This distance can be calculated based on location data, and the specific value of the preset distance range can be adaptively set according to actual needs. For example, it can be adaptively set based on the precision of determining the cell region. This application embodiment does not specifically limit this. The first probability that the corresponding spatial sequencing site belongs to a cell-covered region is calculated based on the target distribution parameters of the adjacent sites. It is understood that belonging to a cell-covered region means the cell region of the target slice corresponding to the spatial sequencing site.
[0161] Reference Figure 7 In some embodiments, step S103 includes, but is not limited to, steps S701 to S702.
[0162] Step S701: Calculate the adjacency message data between adjacent sites and spatial sequencing sites based on the target distribution parameters;
[0163] Step S702: Determine the first probability that the spatial sequencing site belongs to the cell coverage area based on the adjacency message data.
[0164] It should be noted that when the target distribution parameters (including r, r′, θ, θ′) are calculated according to step S102, the conditional probability of each spatial sequencing site belonging to the cell coverage region can actually be calculated according to the following formula (16), thereby determining some high-confidence sites (i.e., spatial sequencing sites with high conditional probability). However, in reality, the possibility that most of a cell region corresponds to high-confidence sites is small. If the cell region is determined directly based on these high-confidence sites, the obtained cell region will be incomplete. Therefore, a method (such as the method shown in steps S701 to S702) is needed to take all sites in the high-confidence site concentration area as sites of the cell coverage region in order to obtain a complete cell region.
[0165]
[0166] A confidence matrix can be constructed based on the conditional probability calculated according to equation (16). Each spatial sequencing site can be represented as a point in the confidence matrix. The location data of the spatial sequencing site can be represented as the coordinate data of the corresponding point. The gene expression level of the spatial sequencing site can be represented as the value of the corresponding point in the confidence matrix.
[0167] It is understandable that the method shown in steps S701 to S702 is actually a belief propagation algorithm (BP algorithm). The BP algorithm is based on a Markov random field (MRF) and is used to describe the correlation of all nodes (i.e., spatial sequencing sites) in the field. In the BP algorithm, the joint probability distribution of each spatial sequencing site and its neighboring sites can be assumed to be as shown in the following equation (17).
[0168]
[0169] Where x represents the spatial sequencing site, and y represents the adjacent site of that spatial sequencing site. (i,j) y represents the gene expression level at spatial sequencing sites. (i,j) Indicates the state of the spatial sequencing site, y (k,l) This represents the state of adjacent sites, and N represents the set of adjacent sites.
[0170] Each spatial sequencing site has a state value (e.g., whether it belongs to a cell-covered region) and an observed value (e.g., the gene expression level of the spatial sequencing site). The likelihood function between the state value and the observed value reflects the dependency between them. The potential energy between adjacent spatial sequencing sites can reflect the correlation between them. The BP algorithm uses the mutual message passing between adjacent spatial sequencing sites to update the labeling state of the entire MRF. The BP algorithm is an iterative algorithm. After multiple iterations, if the confidence of all spatial sequencing sites (i.e., the first probability of belonging to a cell-covered region) no longer changes, the state of each spatial sequencing site at this time is considered the optimal state. Steps S701 to S702 are explained in detail below.
[0171] In step S701 of some embodiments, a spatial sequencing site is randomly selected, and its adjacent sites are determined according to the above method. Adjacency message data is then calculated based on the target distribution parameters of the adjacent sites. The adjacency message data refers to the message passed from the adjacent sites to the spatial sequencing site, which includes information such as the distribution and state that the spatial sequencing site should be in.
[0172] Reference Figure 8 In some embodiments, step S701 includes, but is not limited to, steps S801 to S803.
[0173] Step S801: Calculate the conditional probability that the spatial sequencing site belongs to the cell coverage area based on the target distribution parameters of the spatial sequencing site.
[0174] Step S802: Calculate the first potential difference of the adjacency message based on the first state data of the adjacent sites, the target distribution parameters of the adjacent sites, and the preset potential function; wherein, the first state data indicates whether the adjacent sites belong to cell-covered or non-cell-covered regions.
[0175] Step S803: Calculate the adjacent message data based on the conditional probability and the first potential difference.
[0176] In step S801 of some embodiments, the method for calculating the conditional probability that a spatial sequencing site belongs to a cell-covered region based on the target distribution parameters refers to the above formula (16), X (i,j) This indicates the gene expression level at the spatial sequencing site.
[0177] In step S802 of some embodiments, the first potential difference represents the difference between adjacent messages belonging to high confidence and those belonging to low confidence. The first potential difference can be calculated according to the following formula (18).
[0178] λ (k,l) =log(ψ′(x) (k,l) ,1) / ψ′(x (k,l),0))......Equation (18)
[0179] Where, ψ′(x (k,l) ,1) indicates that the adjacent message belongs to high confidence, ψ′(x (k,l) ,0) indicates that the adjacent message belongs to the low confidence category. ψ′(x (k,l) ,1) and ψ′(x (k,l) ,0) can be calculated according to the potential energy equation (i.e., potential function) shown in equation (19).
[0180]
[0181] When ψ′(x) is calculated according to equation (19) (k,l) ,1) and ψ′(x (k,l) When , 0), y in equation (19) (i,j) Replace the first state data of the adjacent sites with r, θ, r′, and θ′, which are replaced with the target distribution parameters of the adjacent sites. The first state data indicates whether the adjacent site belongs to a cell-covered region or a non-cell-covered region. It can be understood that in equation (19), y... (i,j) =1 indicates that it belongs to the cell coverage area, y (i,j) =0 indicates that it belongs to a non-cell-covered area.
[0182] In step S803 of some embodiments, the adjacency message data l is calculated based on the conditional probability, the first potential difference, and the following formula (20). (k,l)→(i,j) .
[0183]
[0184] Where p represents the conditional probability, and (o,p) represents the other adjacent sites of the spatial sequencing site (i,j) besides the adjacent site (k,l).
[0185] It is understandable that the adjacency message data shown in equation (20) is actually the message data m shown in the following equation. (k,l)→(i,j) Log-likelihood form:
[0186]
[0187] Furthermore, the marginal probability distribution can be expressed as follows:
[0188]
[0189] In step S702 of some embodiments, a first probability is determined based on adjacency message data that the spatial sequencing site belongs to the cell coverage area.
[0190] It is understandable that steps S701 to S702 only represent one iteration. It is necessary to select another spatial sequencing site and repeat steps S701 to S702 until the first probability of all spatial sequencing sites no longer changes. The first probability at which it no longer changes is used as the confidence level that each spatial sequencing site ultimately belongs to the cell-covered region.
[0191] Reference Figure 9 In some embodiments, step S702 includes, but is not limited to, steps S901 to S903.
[0192] Step S901: Sum the multiple adjacent message data to obtain the total message data;
[0193] Step S902: Calculate the second potential difference of the spatial sequencing site based on the second state data of the spatial sequencing site and the target distribution parameters of the spatial sequencing site; wherein, the second state data indicates whether the spatial sequencing site belongs to a cell-covered region or a non-cell-covered region.
[0194] Step S903: Calculate the first probability that the spatial sequencing site belongs to the cell-covered region based on the total message data and the second potential difference.
[0195] In steps S901 to S903 of some embodiments, the adjacency message data of multiple adjacent sites corresponding to the spatial sequencing site are summed to obtain the total message data. The second potential difference λ of the spatial sequencing sites is calculated according to the methods shown in steps S801 to S803. (i,j) The first probability of a spatial sequencing site is calculated based on the total message data, the second potential difference, and the following formula (21). The first probability represents the probability that the spatial sequencing site belongs to a cell-covered region.
[0196]
[0197] The advantage of steps S801 to S803 and steps S901 to S903 is that, from message data m (k,l)→(i,j) The calculation formula shows that the message data m (k,l)→(i,j)In reality, this is the result of aggregating the adjacency message data of all adjacent sites of a spatial sequencing site. The adjacency message data refers to the information conveyed by the adjacent sites to the spatial sequencing site, indicating its distribution and state. Therefore, this embodiment estimates the state of each spatial sequencing site based on the state of its adjacent sites. Thus, for spatial sequencing sites directly determined as non-high-confidence sites based on conditional probability, the final state of the spatial sequencing site can be re-estimated based on its adjacent sites. In this way, when multiple high-confidence spatial sequencing sites exist among the adjacent sites of a spatial sequencing site, the spatial sequencing site can also be identified as a high-confidence spatial sequencing site. That is, spatial sequencing sites at the edge of a high-confidence region can also be included in the high-confidence region, avoiding omissions. At this point, a complete cellular region can be obtained based on multiple spatial sequencing sites belonging to the cell coverage area.
[0198] In step S104 of some embodiments, multiple spatial sequencing sites can be divided according to a first probability to obtain high-confidence sites belonging to the cell coverage area. Based on the location data of the high-confidence sites, the enclosing region of the high-confidence sites can be determined. This enclosing region is the region in the capture area that corresponds to the cell region of the target slice, thus the cell region of the target slice can be determined.
[0199] Reference Figure 10 In some embodiments, step S104 includes, but is not limited to, steps S1001 to S1002.
[0200] Step S1001: Construct a cell mask matrix based on the first probability and location data, and construct a spatial counting matrix based on gene expression information;
[0201] Step S1002: Determine the cell regions of the target slice based on the cell mask matrix and the spatial counting matrix.
[0202] In step S1001 of some embodiments, a cell mask matrix M is constructed based on the first probability and location data of each spatial sequencing site. Specifically, it is determined whether the spatial sequencing site belongs to a cell-covered region based on the first probability. If it belongs to a cell-covered region, the location data of the spatial sequencing site in the cell mask matrix M can be assigned a value of 1. If it belongs to a non-cell-covered region, the location data of the spatial sequencing site in the cell mask matrix M can be assigned a value of 0. A spatial counting matrix X is constructed based on the location data of each spatial sequencing site and the gene expression level of the spatial sequencing site. In the spatial counting matrix X, the gene expression level is used as the value assigned to the corresponding location data.
[0203] Reference Figure 11 In some embodiments, step S1001 includes, but is not limited to, steps S1101 to S1102.
[0204] Step S1101: Determine the second probability that the spatial sequencing site belongs to a non-cell-covered region based on the target distribution parameters of the adjacent sites;
[0205] Step S1102: Compare the first probability with the second probability, and construct a cell mask matrix based on the comparison result, preset mask value, and position data.
[0206] In step S1101 of some embodiments, a second probability that a spatial sequencing site belongs to a non-cell-covered region is calculated based on the target distribution parameters of adjacent sites and the following formula (22). It is understood that the specific method for calculating the second probability can refer to the method for calculating the first probability, and will not be described again in this embodiment.
[0207]
[0208] In step S1102 of some embodiments, the first probability P(y) is... (i,j) =1) and the second probability P(y (i,j) The values of the spatial sequencing sites are compared with those of the cell mask matrix M, and the position data of each spatial sequencing site is assigned a value of 1 based on the comparison results. Specifically, if the first probability is greater than the second probability, it indicates that the corresponding spatial sequencing site belongs to a cell-covered region, and the position data of the spatial sequencing site is assigned a value of 1. If the first probability is less than the second probability, it indicates that the corresponding spatial sequencing site belongs to a non-cell-covered region, and the position data of the spatial sequencing site is assigned a value of 0.
[0209] In step S1002 of some embodiments, cell segmentation is performed based on the cell mask matrix and the spatial counting matrix to ultimately determine multiple cell regions of the target slice.
[0210] Reference Figure 12 In some embodiments, step S1002 includes, but is not limited to, steps S1201 to S1203.
[0211] Step S1201: Construct a convolution kernel based on a preset cell diameter, and perform convolution calculation based on the convolution kernel and the spatial counting matrix to obtain a convolution matrix;
[0212] Step S1202: Perform dot product calculation on the convolution matrix and the cell mask matrix to obtain the cell tag matrix; wherein, the row and column dimensions of the cell tag matrix are the same as those of the spatial counting matrix, the cell tag matrix includes the location data of the spatial sequencing sites and the cell tags corresponding to the spatial sequencing sites, and the cell tags include cell region tags.
[0213] Step S1203: Determine the cell regions of the target slice based on the location data corresponding to the cell region labels.
[0214] In step S1201 of some embodiments, a convolution kernel is constructed based on a preset cell diameter. The cell diameter can be determined based on prior knowledge, such as determining the cell diameter corresponding to each cell in the target slice based on prior knowledge, and then calculating the average value of multiple cell diameters. A convolution kernel G is constructed based on the average value of multiple cell diameters. The convolution kernel G can be an odd number of convolution kernels. For example, if the cell diameter s is 15, then the convolution kernel G = s * s. It is understood that when performing image representation, the unit of cell diameter is pixels. The convolution kernel G and the spatial counting matrix X are convolved as shown in the following equation (23) to obtain the convolution matrix V.
[0215] V = -Conv(X,G)......Equation (23)
[0216] In step S1202 of some embodiments, the convolution matrix V and the cell mask matrix M are multiplied by a dot product to obtain the cell tag matrix L. The row and column dimensions of the cell tag matrix L are the same as those of the spot in the ST processing; that is, when the gene expression information is in matrix form, the row and column dimensions of the cell tag matrix L are the same as those of the gene expression information. The cell tag matrix L is used to describe whether the location data corresponding to each spatial sequencing site belongs to a cell-covered region. When it belongs to a cell-covered region, the cell tag corresponding to the spatial sequencing site is a cell region tag. When it belongs to a non-cell-covered region, the cell tag corresponding to the spatial sequencing site is a non-cell region tag. It can be understood that the cell region tag can be in the form of a number, where a number of 0 represents a non-cell region tag, and a number greater than 0 represents a cell region tag. In the cell region tag, different number values represent different cell regions.
[0217] In step S1203 of some embodiments, the cell regions of the target slice are segmented according to the position data of cell region labels in the cell label matrix L.
[0218] The advantage of steps S1201 to S1203 is that it can not only determine the cell regions of the target slice, but also distinguish each cell region according to the number corresponding to the cell region label.
[0219] In one specific embodiment, refer to Figures 13 to 16 Based on the cell region determination method provided in the embodiments of this application, cell regions are determined in mouse brain spatial transcriptome data generated by Stereo-seq sequencing technology. Figures 13 to 16 To determine the effect image for a portion of the cell region captured in a real experiment. Among them, Figure 13 These are cellular regions identified using artificial labeling methods based on RNA signals. Figure 14 The cell regions are determined using the methods provided in the embodiments of this application. Figure 15It is a cell region determined using image-referenced methods (such as the Stardist method). Figure 16 This refers to cell regions identified using methods that delineate cell regions based on partial RNA signals (such as the Baysor method). Through... Figures 13 to 16 The comparison shows that... Figure 16 The method shown has the drawback of dividing a single cell region into multiple cell regions, namely... Figure 16 The method shown does not yield a complete cellular region. In contrast, the method described in this application's embodiments yields results close to those of artificial labeling methods. The method described in this application's embodiments delineates a larger cell region, eliminating discrete, scattered cells, and allowing adjacent regions to be integrated into a single cellular region. Furthermore, the cells delineated by the method described in this application's embodiments are closer to the theoretical size of real cells. Therefore, compared to methods based on images or partial RNA to determine cellular regions, the method provided in this application's embodiments is more suitable for analyzing high-throughput spatial transcriptome sequencing data. The cell region determination results obtained by the method provided in this application's embodiments are highly reliable, and the delineated cellular regions exhibit strong RNA signals, thus facilitating downstream data analysis.
[0220] The cell region determination device provided in the embodiments of this application will be described below.
[0221] Reference Figure 17 In some embodiments, this application also provides a cell region determination device, which includes:
[0222] The information acquisition unit 1701 is used to acquire gene expression level information from the spatial transcription data of the target slice; wherein, the gene expression level information includes the location data of multiple spatial sequencing sites and the gene expression level of each spatial sequencing site;
[0223] The distribution parameter calculation unit 1702 is used to calculate the target distribution parameter corresponding to the gene expression level of each spatial sequencing site following a preset distribution.
[0224] The probability calculation unit 1703 is used to calculate the first probability that each spatial sequencing site belongs to the cell coverage area based on the target distribution parameters of the adjacent sites of each spatial sequencing site.
[0225] Cell region determination unit 1704 is used to determine the cell region of the target slice based on a first probability and location data.
[0226] It is evident that the content of the above-described cell region determination method embodiments is applicable to the embodiments of this cell region determination device. The specific functions implemented by this cell region determination device embodiment are the same as those of the above-described cell region determination method embodiments, and the beneficial effects achieved are also the same as those achieved by the above-described cell region determination method embodiments.
[0227] Reference Figure 18 , Figure 18 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0228] The processor 1801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0229] The memory 1802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1802 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1802 and is called and executed by the processor 1801 using the cell region determination method of the embodiments of this application.
[0230] The input / output interface 1803 is used to implement information input and output;
[0231] The communication interface 1804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0232] Bus 1805 transmits information between various components of the device (e.g., processor 1801, memory 1802, input / output interface 1803, and communication interface 1804);
[0233] The processor 1801, memory 1802, input / output interface 1803 and communication interface 1804 are connected to each other within the device via bus 1805.
[0234] This application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the cell region determination method described above.
[0235] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.
[0236] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0237] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.
[0238] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0239] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0240] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0241] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0242] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.
[0243] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.
Claims
1. A method for determining cell regions, characterized in that, The method includes: Gene expression level information is obtained from the spatial transcription data of the target slice; wherein, the gene expression level information includes the location data of multiple spatial sequencing sites and the gene expression level of each spatial sequencing site; Calculate the target distribution parameters corresponding to the gene expression level of each of the spatial sequencing sites following a preset distribution; The first probability that each spatial sequencing site belongs to a cell-covered region is calculated based on the target distribution parameters of the adjacent sites of each spatial sequencing site; wherein, the adjacent sites of the spatial sequencing site are the spatial sequencing sites adjacent to the spatial sequencing site; The cell regions of the target slice are determined based on the first probability and the location data.
2. The method according to claim 1, characterized in that, The calculation of the target distribution parameters corresponding to the gene expression level of each spatial sequencing site following a preset distribution includes: A first log-likelihood function of the preset distribution is constructed based on the gene expression level at each of the spatial sequencing sites; The latent variables are determined based on the first log-likelihood function, and the target steps are executed iteratively. The target steps include: calculating the conditional probability of the latent variables based on the randomly initialized original distribution parameters; maximizing the first log-likelihood function based on the conditional probability to obtain preliminary distribution parameters; if the number of executions of the target steps reaches a preset number, the preliminary distribution parameters at the preset number of executions will be used as the target distribution parameters, or if the preliminary distribution parameters converge, the preliminary distribution parameters will be used as the target distribution parameters.
3. The method according to claim 2, characterized in that, The step of constructing the first log-likelihood function of the preset distribution based on the gene expression level of each of the spatial sequencing sites includes: Probabilistic expression data of the preset distribution are constructed based on the gene expression levels at each of the spatial sequencing sites; The state variables of the spatial sequencing sites are constructed based on the probability expression data; wherein, the state variables are used to indicate whether the spatial sequencing sites belong to cell-covered regions or non-cell-covered regions; The distribution variables of the preset distribution are constructed based on the probability expression data; wherein, the distribution variables are used to indicate whether the type of the preset distribution is a cell-covered region type or a non-cell-covered region type; The first log-likelihood function is constructed based on the state variable and the distribution variable.
4. The method according to claim 2, characterized in that, The conditional probability includes cell region probability, first variable probability, and second variable probability; the original distribution parameters include first distribution parameters and second distribution parameters. The step of calculating the conditional probability of the latent variable based on the randomly initialized original distribution parameters includes: The probability that the spatial sequencing site belongs to the cell coverage area is calculated based on the preset conditions and the original distribution parameters to obtain the cell region probability. The probability of the first variable is obtained by performing a logarithmic calculation based on the first distribution parameter. The probability of the second variable is obtained by performing logarithmic differentiation based on the second distribution parameters.
5. The method according to claim 4, characterized in that, The step of maximizing the first log-likelihood function based on the conditional probability to obtain preliminary distribution parameters includes: The probabilities of the cell regions are summed to obtain the probability sum; Calculate the sum of the products of the cell region probability and the second variable probability, and then calculate the ratio of the sum of the products to the cumulative value of the cell region probabilities to obtain the first probability ratio. The second probability ratio is calculated based on the cell region probability, the first variable probability, and the second variable probability. The expected function of the first log-likelihood function is calculated, and the expected function is maximized based on the sum of probabilities, the first probability ratio, and the second probability ratio to obtain the preliminary distribution parameters.
6. The method according to claim 1, characterized in that, The calculation of the target distribution parameters corresponding to the gene expression level of each spatial sequencing site following a preset distribution includes: A second probability density function of the preset distribution is constructed based on the gene expression level at each of the spatial sequencing sites; Obtain observation data of the spatial sequencing site belonging to the cell coverage area in a preset number of times, and construct a second likelihood function based on the observation data and the second probability density function; Logarithmic transformation of the second likelihood function yields the log-likelihood function; The target distribution parameters are obtained by differentiating the log-likelihood function.
7. The method according to claim 1, characterized in that, The step of determining the first probability that the spatial sequencing site belongs to the cell-covered region based on the target distribution parameters of the adjacent sites includes: Calculate the adjacency message data between the adjacent sites and the spatial sequencing sites based on the target distribution parameters; The probability of determining that the spatial sequencing site belongs to the cell-covered region is determined based on the adjacency message data.
8. The method according to claim 7, characterized in that, The step of calculating the adjacency message data between the adjacent sites and the spatial sequencing sites based on the target distribution parameters includes: Calculate the conditional probability that the spatial sequencing site belongs to the cell coverage region based on the target distribution parameters of the spatial sequencing site; The first potential difference of the adjacency message is calculated based on the first state data of the adjacent sites, the target distribution parameters of the adjacent sites, and a preset potential function; wherein, the first state data indicates whether the adjacent sites belong to cell-covered or non-cell-covered regions. The adjacency message data is calculated based on the conditional probability and the first potential difference.
9. The method according to claim 7, characterized in that, The step of determining the first probability that the spatial sequencing site belongs to the cell-covered region based on the adjacency message data includes: The total message data is obtained by summing the multiple adjacent message data. The second potential difference of the spatial sequencing site is calculated based on the second state data of the spatial sequencing site and the target distribution parameters of the spatial sequencing site; wherein, the second state data indicates whether the spatial sequencing site belongs to a cell-covered region or a non-cell-covered region; The probability that the spatial sequencing site belongs to the cell-covered region is calculated based on the total message data and the second potential difference.
10. The method according to any one of claims 1 to 9, characterized in that, Determining the cell region of the target slice based on the first probability and the location data includes: A cell mask matrix is constructed based on the first probability and the location data, and a spatial counting matrix is constructed based on the gene expression level information; The cell regions of the target slice are determined based on the cell mask matrix and the spatial counting matrix.
11. The method according to claim 10, characterized in that, The step of constructing a cell mask matrix based on the first probability and the location data includes: The second probability that the spatial sequencing site belongs to a non-cell-covered region is determined based on the target distribution parameters of the adjacent sites; The first probability is compared with the second probability, and the cell mask matrix is constructed based on the comparison result, the preset mask value, and the position data.
12. The method according to claim 10, characterized in that, Determining the cell region of the target slice based on the cell mask matrix and the spatial counting matrix includes: Convolution kernels are constructed based on preset cell diameters, and convolution calculations are performed based on the convolution kernels and the spatial counting matrix to obtain a convolution matrix; A cell tag matrix is obtained by performing a dot product between the convolution matrix and the cell mask matrix; wherein the row and column dimensions of the cell tag matrix are the same as those of the spatial counting matrix, the cell tag matrix includes the location data of the spatial sequencing sites and the cell tags corresponding to the spatial sequencing sites, and the cell tags include cell region tags; The cell regions of the target slice are determined based on the location data corresponding to the cell region labels.
13. A cell region determination device, characterized in that, The device includes: An information acquisition unit is used to acquire gene expression level information from the spatial transcription data of the target slice; wherein, the gene expression level information includes the location data of multiple spatial sequencing sites and the gene expression level of each spatial sequencing site; The distribution parameter calculation unit is used to calculate the target distribution parameter corresponding to the gene expression level of each spatial sequencing site following a preset distribution. A probability calculation unit is used to calculate a first probability that each spatial sequencing site belongs to a cell-covered region based on the target distribution parameters of the adjacent sites of each spatial sequencing site; wherein, the adjacent sites of the spatial sequencing site are spatial sequencing sites adjacent to the spatial sequencing site. A cell region determination unit is used to determine the cell region of the target slice based on the first probability and the location data.
14. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the cell region determination method according to any one of claims 1 to 12.
15. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the cell region determination method according to any one of claims 1 to 12.
16. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the cell region determination method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Cell display method and device, computer equipment and computer readable storage medium
CN110334604A
Accurate And Robust Information-Deconvolution From Bulk Tissue Transcriptomes
US20210142867A1