Cell region determination method and device, storage medium and computer equipment

By acquiring staining images and spatial transcriptome sequencing data, determining the cell center and constructing a cell center map, and using the inflection point of gene expression to fit the cell area, the problem of low accuracy in the image after cell segmentation was solved, and higher accuracy in cell area determination was achieved.

CN120656165APending Publication Date: 2025-09-16SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410261006.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing technology has cell adhesion, overlap and holes in the image after cell segmentation, resulting in low accuracy in cell area determination.

Method used

By obtaining staining images and spatial transcriptome sequencing data of biological samples, the cell center is determined and a cell center map is constructed. The cell area is fitted using the inflection point of gene expression, reducing dependence on staining images and making full use of the gene expression changes in spatial transcriptome sequencing data.

Benefits of technology

The accuracy of cell region determination is improved, the dependence on high-precision cell segmentation is reduced, and the accuracy of cell boundary delineation is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656165A_ABST
    Figure CN120656165A_ABST
Patent Text Reader

Abstract

The invention provides a cell region determination method and device, a storage medium and computer equipment, and belongs to the field of biological information. The method comprises the following steps: acquiring a dyeing image of a biological sample and space transcriptome sequencing data, and determining cell centers of a plurality of cells based on the dyeing image; for any target cell, a corresponding cell center map is constructed, the cell center map comprises multiple rays dividing the dyeing image into multiple space areas, and the multiple rays take the target cell center of the target cell as a common starting point; determining a gene expression quantity inflection point in each spatial region according to the spatial transcriptome sequencing data to obtain a plurality of gene expression quantity inflection points of the target cell, the gene expression quantity inflection points being points with the highest signal-to-noise ratio in the corresponding spatial regions; and fitting according to the plurality of gene expression quantity inflection points to obtain a first cell region of the target cell. The method can improve the accuracy of determining the cell region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of bioinformatics, and in particular to a method, apparatus, storage medium, and computer equipment for determining a cell region. Background Art

[0002] Cell segmentation is one of the most fundamental and important areas in medical image processing. It is the fundamental prerequisite for identifying and counting cell images. Cell segmentation involves dividing a cell image into several non-overlapping regions based on characteristics such as grayscale, color, and geometric shape, ensuring that these characteristics appear consistent or similar within the same region, but distinctly different across different regions. Cell segmentation can be used to study cell morphology and structure, understanding the characteristics of cells and tissues, and is of great significance for the diagnosis and treatment of various diseases.

[0003] Among the related technologies, some use watershed segmentation algorithms for cell segmentation, while others use wavelet transform-based edge detection integrated algorithms and deep learning algorithms for cell segmentation. Cell regions are determined based on the cell segmentation results, and the determined cell regions are used to classify gene expression molecules into belonging cells.

[0004] However, due to the complexity of cell images, uneven illumination of microscopic images, and grayscale changes of the target objects themselves, there are still some problems in the images after cell segmentation, mainly cell adhesion, cell overlap, and hole phenomena. These problems lead to low accuracy of the determined cell areas. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to provide a method, apparatus, storage medium and computer equipment for determining a cell region, which can improve the accuracy of determining a cell region.

[0006] To achieve the above objectives, a first aspect of an embodiment of the present application provides a method for determining a cell region, the method comprising:

[0007] Acquire a staining image and spatial transcriptome sequencing data of a biological sample, and determine the cell centers of a plurality of cells based on the staining image;

[0008] For any target cell, construct a corresponding cell center map, wherein the cell center map includes a plurality of rays that divide the stained image into a plurality of spatial regions, and the plurality of rays have a target cell center of the target cell as a common starting point;

[0009] Determining a gene expression inflection point in each spatial region according to the spatial transcriptome sequencing data to obtain multiple gene expression inflection points of the target cells, wherein the gene expression inflection point is a point with the highest signal-to-noise ratio in the corresponding spatial region;

[0010] The first cell region of the target cell is obtained according to the inflection points of the multiple gene expression levels.

[0011] A second aspect of the present application provides a cell region determination device, comprising:

[0012] an acquisition unit, configured to acquire a staining image and spatial transcriptome sequencing data of a biological sample, and determine cell centers of a plurality of cells based on the staining image;

[0013] a cell center map construction unit, configured to construct a corresponding cell center map for any target cell, wherein the cell center map includes a plurality of rays that divide the stained image into a plurality of spatial regions, and the plurality of rays have a target cell center of the target cell as a common starting point;

[0014] a gene expression inflection point determination unit, configured to determine a gene expression inflection point in each spatial region based on the spatial transcriptome sequencing data, to obtain a plurality of gene expression inflection points of the target cell, wherein the gene expression inflection point is a point with the highest signal-to-noise ratio in the corresponding spatial region;

[0015] A cell region fitting unit is used to obtain a first cell region of the target cell according to the inflection points of the multiple gene expression levels.

[0016] In one embodiment, the gene expression level inflection point determination unit includes:

[0017] A first calculation subunit is configured to calculate, based on the spatial transcriptome sequencing data, gene expression values ​​in a plurality of fan-shaped regions corresponding to different radii in each spatial region;

[0018] A second calculation subunit is used to calculate the signal-to-noise ratio of each sector-shaped area based on the gene expression value and the radius of each sector-shaped area;

[0019] The first determination subunit is used to determine the gene expression inflection point of the corresponding spatial region according to the target sector region with the highest signal-to-noise ratio.

[0020] In one embodiment, the first determining subunit includes:

[0021] The first determining module is used to determine the center point of the arc corresponding to the target sector area with the highest signal-to-noise ratio;

[0022] The second determination module is used to determine the gene expression inflection point of the corresponding spatial area according to the center point of the arc.

[0023] In one embodiment, the second computing subunit includes:

[0024] The first calculation module is used to calculate the ratio of the gene expression value in each sector area to the corresponding radius;

[0025] The third determining module is configured to determine the signal-to-noise ratio of each sector area according to the ratio.

[0026] In one embodiment, the first computing subunit includes:

[0027] A second calculation module is used to calculate the gene expression value of the sampling point in each sector area and the Pearson correlation coefficient between each sampling point and the adjacent sampling points in each sector area according to the spatial transcriptome sequencing data;

[0028] A third calculation module is used to calculate the average Pearson correlation coefficient of each sampling point in the fan-shaped area according to the Pearson correlation coefficient of each sampling point and the adjacent sampling points;

[0029] The summarizing subunit is used to summarize the gene expression value of each sampling point whose gene expression value is higher than the first preset threshold and whose average Pearson coefficient is higher than the second preset threshold, to obtain the gene expression value in each of the sector areas.

[0030] In one embodiment, the cell-centric map construction unit includes:

[0031] A ray drawing subunit is used to draw multiple rays from any target cell, with the target cell center of the target cell as a common starting point, based on the common starting point, with the angles between any two adjacent rays being equal;

[0032] The spatial region division subunit is used to divide the stained image into multiple spatial regions according to the multiple rays to obtain a cell center map of the target cell.

[0033] In one embodiment, the ray extraction subunit includes:

[0034] a fourth determining module, configured to determine the number of spatial regions included in the cell center map of each target cell;

[0035] a fourth calculation module, configured to calculate a ray angle based on the quantity, wherein the ray angle is an angle between adjacent rays in the cell center map;

[0036] The ray generation module is configured to use a preset ray starting from the target cell center of the target cell as an initial ray, and to generate multiple rays starting from the target cell center one by one based on the ray angle.

[0037] In one embodiment, the cell region fitting unit includes:

[0038] A minimum convex polygon construction subunit, used to construct a minimum convex polygon containing the multiple gene expression inflection points;

[0039] The second determining subunit is configured to determine a first cell region of the target cell according to the minimum convex polygon.

[0040] In one embodiment, the acquiring unit includes:

[0041] a third determining subunit, configured to determine a second cell region of the plurality of cells based on the staining image;

[0042] The fourth determining subunit is configured to determine a plurality of cell centers according to the geometric centers of the plurality of second cell regions.

[0043] In one embodiment, the acquiring unit further includes:

[0044] a cell nucleus center determining subunit, configured to determine the cell nucleus centers of a plurality of cells based on the stained image;

[0045] The fifth determining subunit is used to determine multiple cell centers based on the multiple cell nucleus centers.

[0046] In one embodiment, after obtaining the first cell region of the target cell according to the inflection point fitting of the multiple gene expression levels, the cell region determining device is further configured to:

[0047] Determining gene expression data in the target cells based on the first cell region and the spatial transcriptome sequencing data, and assigning the gene expression data to cells;

[0048] The cell type of the target cell is determined based on the gene expression data of the target cell.

[0049] A third aspect of an embodiment of the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the cell region determination method described in the first aspect when executing the computer program.

[0050] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the cell region determination method described in the first aspect.

[0051] To achieve the above-mentioned purpose, the fifth aspect of the embodiment of the present application proposes a computer program product, which includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the cell area determination method described in the first aspect.

[0052] The cell region determination method, device, storage medium and computer equipment proposed in the embodiments of the present application obtain a staining image and spatial transcriptome sequencing data of a biological sample, and determine the cell centers of multiple cells based on the staining image; for any target cell, construct a corresponding cell center map, the cell center map includes multiple rays that divide the staining image into multiple spatial regions, and the multiple rays have the target cell center of the target cell as a common starting point; determine the gene expression inflection point in each spatial region based on the spatial transcriptome sequencing data, and obtain multiple gene expression inflection points of the target cell, where the gene expression inflection point is the point with the highest signal-to-noise ratio in the corresponding spatial region; and obtain the first cell region of the target cell based on fitting of the multiple gene expression inflection points.

[0053] Therefore, in the embodiment of the present disclosure, only the cell center and approximate position of the cell are determined by the staining image, and the cell area is not determined directly based on the staining image. After determining the cell center based on the staining image and constructing the cell center map, the gene expression inflection point in each spatial region in the cell center map is determined based on the spatial transcriptome sequencing data. The gene expression inflection point is the point with the highest signal-to-noise ratio in the corresponding spatial region. Since the gene expression level inside the cell is usually higher than the gene expression level outside the cell, there is a change inflection point in the distribution of gene expression levels inside and outside the cell, and the gene expression inflection point can better distinguish the cell boundaries. Determining the cell area by the gene expression inflection point can reduce the degree of dependence on the division of cell boundaries in the staining image, does not require high-precision cell segmentation, makes full use of the spatial transcriptome sequencing data and the changes in gene expression levels inside and outside the cell to determine the cell area, and improves the accuracy of the cell area determination by accurately dividing the gene expression molecules in the cell. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.

[0055] Figure 1 is a flow chart of the cell region determination method provided in an embodiment of the present application;

[0056] Figure 2A is a schematic diagram of the gray-white pixel map of the second cell area;

[0057] Figure 2B is a schematic diagram of the cell center determined based on the geometric center of the second cell region;

[0058] Figure 3 It is a schematic diagram of the cell center map containing 16 spatial regions;

[0059] Figure 4 yes Figure 1Flowchart of step 103;

[0060] Figure 5 It is a schematic diagram of a sector area with a radius of d in the space area numbered t;

[0061] Figure 6 This is a schematic diagram of the gene expression inflection point provided in the examples of this application;

[0062] Figure 7 This is a schematic diagram of the first cell region fitted using the minimum convex polygon;

[0063] Figure 8 This is a schematic diagram of the experimental results of applying the method of the present application to 135D1 mouse brain data;

[0064] Figure 9 This is a schematic diagram of cell annotation results corrected by the cell region determination method of the present application;

[0065] Figure 10A This is a schematic diagram of the results of clustering mouse brain cell DGGRC2 according to one embodiment of the present application;

[0066] Figure 10B This is a schematic diagram of the results of clustering DGGRC2 in mouse brain cells based on the Fast correction method;

[0067] Figure 11A This is a schematic diagram of the results of MBDOP2 clustering in mouse brain cells according to one embodiment of the present application;

[0068] Figure 11B This is a schematic diagram of the results of MBDOP2 clustering in mouse brain cells based on the Fast correction method;

[0069] Figure 12A This is a schematic diagram of the results of TEGLU3 clustering in mouse brain cells according to one embodiment of the present application;

[0070] Figure 12B This is a schematic diagram of the results of TEGLU3 clustering in mouse brain cells based on the Fast correction method;

[0071] Figure 13A This is a schematic diagram of the results of TEGLU8 clustering in mouse brain cells according to one embodiment of the present application;

[0072] Figure 13B This is a schematic diagram of the results of clustering TEGLU8 in mouse brain cells based on the Fast correction method;

[0073] Figure 14A This is a schematic diagram of the results of TEGLU6 clustering in mouse brain cells according to one embodiment of the present application;

[0074] Figure 14BThis is a schematic diagram of the results of clustering TEGLU6 in mouse brain cells based on the Fast correction method;

[0075] Figure 15 This is a schematic diagram of the size distribution of 135D1 mouse brain cells provided in the examples of this application;

[0076] Figure 16 A structural diagram of a cell region determination device provided in an embodiment of the present application;

[0077] Figure 17 This is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0078] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0079] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations:

[0080] Several main parameters for clustering effect evaluation:

[0081] (1) Silhouette coefficient: Silhouette coefficient is an indicator to measure the clustering effect. It comprehensively considers the cohesion and separation of the clustering results. The mean of S(i) of all samples is called the silhouette coefficient of the clustering result. The value of the silhouette coefficient is between [-1, 1]. The larger the value, the closer the similar samples are and the farther the different samples are, and the better the clustering effect. S(i) close to 1 means that the clustering of sample i is reasonable. S(i) close to -1 means that sample i should be classified into another cluster. S(i) close to 0 means that sample i is on the boundary of two clusters.

[0082] (2) Davies-Bouldin score: The Davies-Bouldin score measures the average similarity between each cluster and its most similar cluster, where similarity is defined as the ratio of the intra-cluster distance (the distance from cluster center to cluster center) to the inter-cluster distance (the distance between cluster centers). The DB score is used to evaluate the quality of clustering by measuring the similarity between clusters and the difference within clusters. The smaller the value, the more similar the clusters are, and the smaller the difference within the clusters, the better the classification effect.

[0083] (3) Calinski-Harabasz score: A clustering effect evaluation metric. Its essence is the ratio of the inter-cluster distance to the intra-cluster distance. The overall calculation process is similar to the variance calculation, and it is also called the variance ratio criterion. This metric mainly evaluates the clustering effect by calculating the variance ratio between categories in the clustering results. The larger the value, the greater the inter-cluster distance is. In other words, the farther the clusters are from each other, the better the clustering result is, and the better the classification effect is.

[0084] (4) Dunn validity index: Dunn index, referred to as DVI, is calculated by dividing the shortest distance between any two cluster elements (between classes) by the maximum distance in any cluster (within classes). The larger the value, the greater the inter-class distance and the smaller the intra-class distance, and the better the classification effect.

[0085] In related technologies, gene expression molecules are assigned to cells based on registered cell outline images and spatial transcriptome sequencing data. Cell outline images are processed using image processing algorithms such as the watershed algorithm to generate cell segmentation results. Because the amount of RNA in cells obtained through cell segmentation is insufficient, cell labeling is required to more clearly visualize the identified cell regions and gene expression molecules within them. Cell regions are determined directly based on the cell segmentation results.

[0086] The results of gene expression molecule attribution to cells are corrected using a spatiotemporal data correction algorithm. Specifically, a Gaussian mixture model is established based on the coordinates of the gene expression molecules within the cell and the number of unique molecular identifiers, and the model parameters are fitted. The model parameters are used to calculate the probability that the gene expression molecules outside the cell belong to the cell. A threshold is set to determine whether the gene expression molecules belong to the cell. Gene expression molecules that meet the threshold conditions are classified as cells, while those that do not meet the threshold conditions are classified as extracellular.

[0087] However, this method is highly dependent on the cell area determined by the cell contour image and has high requirements for the accuracy of the cell segmentation results. However, the final cell area is inaccurate, resulting in large errors in the subsequent results of assigning gene molecules to cells.

[0088] Based on this, the present application provides a method, device, storage medium, and computer equipment for determining a cell region. This method uses a cell outline image to determine the approximate location of the cell center and the cell, determines the inflection point of gene expression inside and outside the cell based on spatial transcriptome sequencing data, and fits the inflection point of gene expression changes by a minimum convex polygon to determine the cell region. Below, the cell region determination method provided in the present application is described.

[0089] Reference Figure 1 In some embodiments, the cell region determination method provided in the embodiments of the present application includes but is not limited to steps 101 to 104.

[0090] Step 101: Acquire a staining image and spatial transcriptome sequencing data of a biological sample, and determine the cell centers of multiple cells based on the staining image;

[0091] Step 102: construct a corresponding cell center map for any target cell;

[0092] Step 103, determining the gene expression inflection point in each spatial region based on the spatial transcriptome sequencing data, and obtaining multiple gene expression inflection points of the target cells;

[0093] Step 104 : obtaining a first cell region of target cells based on the inflection points of multiple gene expression levels by fitting.

[0094] In the steps 101 to 104 shown in the embodiment of the present application, before determining the cell area, it is necessary to obtain a staining image of the biological sample. The biological sample may include biological tissues, cells, or cell suspensions. The biological sample embedded in (Optimal Cutting Temperature compound, OCT) is made into slices and placed on a spatiotemporal transcriptome sequencing chip. The biological sample is stained with ssDNA dye, and the stained biological sample is photographed with a microscope to obtain a staining image of the biological sample. ssDNA dye can be adsorbed with nucleic acids and emit fluorescence. Since the DNA and RNA in the cell nucleus are highly enriched, the spatial position of the cell nucleus can be identified by ssDNA staining. After the slice is attached to the chip, a series of operations are performed to obtain a sequencing library that can be tested on the machine. The obtained sequencing library is sequenced using a sequencer, and the obtained sequencing data is reduced to the spatial expression information of the biological sample by a bioinformatics analysis method to obtain spatial transcriptome sequencing data. Spatial transcriptome sequencing data can explore the gene expression pattern of each cell in the tissue at spatial resolution, combining gene expression data with its spatial positioning in tissue structure. In this embodiment, the cell region determination method is preferably applied to Stereo-seq spatial transcriptome data to divide cell regions, but its application scenario is not specifically limited and can be adjusted accordingly according to actual needs or possible needs.

[0095] After obtaining the dyeing image of biological sample, further, need to obtain the spatial transcriptome sequencing data of biological sample, dyeing image and spatial transcriptome sequencing data are carried out image registration to generate the gene expression information of cell.Specifically, spatial transcriptome sequencing data can be through gene expression large-scale data repository (Gene Expression Omnibus data base, GEO), multidisciplinary genome sequencing data repository (Sequence Read Archive, SRA) or some online open platforms that provide detailed biological tissue molecular analysis atlas.Dyed image is carried out cell segmentation by watershed algorithm to obtain the corresponding relationship between pixel coordinates and cell position, obtain the pixel gray-white figure corresponding to cell outline.By carrying out the cell image figure after cell segmentation, the approximate position and cell center of cell can be determined.

[0096] In some embodiments, determining the cell centers of a plurality of cells based on the staining image includes but is not limited to the following steps:

[0097] determining a second cell region of the plurality of cells based on the stained image;

[0098] A plurality of cell centers are determined based on the geometric centers of the plurality of second cell regions.

[0099] Wherein, after the stained image is subjected to cell segmentation, a gray-white pixel image corresponding to the cell outline is obtained, and the cell outline pixel image characterizes the approximate position and cell morphology of the cell. The second cell region is the approximate outline of the cell determined according to the gray-white pixel image of the cell outline. The pixel coordinates corresponding to a plurality of points on the cell outline are obtained, and the coordinate position of the geometric center of the second cell region is calculated according to the average value of the pixel coordinates of each point, and the geometric center is used as the cell center of the cell. By determining the cell center of the cell, the spatial position of the cell is located, and the accuracy of the subsequent spatial region division and the determination of the gene expression inflection point in each spatial region can be improved.

[0100] Please refer to Figure 2A , is a schematic diagram of the gray-white pixel map of the second cell area, refer to Figure 2B , is a schematic diagram of the cell center determined based on the geometric center of the second cell region. Figure 2A In the image, the approximate outline of the cell can be seen, resembling a convex polygon. The geometric center of the cell is determined based on the pixel coordinates of multiple points on the cell outline. These points can be used to select the vertices of the convex polygon. The geometric center of the second cell region is calculated based on the vertex coordinates and is used as the center of the cell, also known as the cell center.

[0101] In some embodiments, determining the cell centers of a plurality of cells based on the staining image further includes but is not limited to the following steps:

[0102] Determine the nuclear centers of multiple cells based on staining images;

[0103] Multiple cell centers were determined based on multiple nuclear centers.

[0104] In the aforementioned steps, the biological sample is stained with ssDNA dye, and a stained image is obtained by taking a photo under a microscope. The ssDNA dye can well identify the position of the cell nucleus. The Cellbin algorithm in the Stereo-seq spatiotemporal group data analysis scheme can achieve analysis at the single cell level. By fitting the stained image with the spatial transcriptome sequencing data, the position of the cell nucleus can be identified by image. Since DNA and RNA are enriched in the cell nucleus, the cell nucleus will show strong fluorescence after being stained, which is convenient for determining the position of the cell nucleus. After determining the position of the cell nucleus, the position of the center of the cell nucleus is further determined by an image recognition algorithm. No restrictions are placed on the image recognition algorithm here. Specifically, the pixel coordinate data of the cell nucleus is extracted to determine the outline of the cell nucleus. The center coordinates of the cell nucleus are obtained by calculating the average value of the coordinate values ​​of each pixel point on the cell nucleus outline, and the determined cell nucleus center is used as the cell center point. Whether the geometric center of the cell is used as the cell center or the cell nuclear center is used as the cell center point, the cell can be accurately located to obtain the accurate initial position of the cell.

[0105] In some embodiments, step 102 includes but is not limited to the following steps:

[0106] For any target cell, the target cell center of the target cell is used as a common starting point, and multiple rays are drawn based on the common starting point, and the angles between any two adjacent rays are equal;

[0107] The stained image is divided into multiple spatial regions according to multiple rays to obtain the cell center map of the target cell.

[0108] In the aforementioned steps, the cell center is determined, and the origin of a rectangular coordinate system is used as the coordinate origin. Multiple rays are derived from this origin. The angle between any two adjacent rays is equal, and each pair of adjacent rays forms a spatial region. Each spatial region is the angular region enclosed by the two rays, with the common origin as the vertex. The multiple rays divide the stained image into multiple spatial regions, resulting in a cell center map containing the multiple rays dividing the stained image into multiple spatial regions.

[0109] In some embodiments, for any target cell, the target cell center of the target cell is used as a common starting point, and multiple rays are drawn from the gene common starting point, including but not limited to the following steps:

[0110] Determine the number of spatial regions contained in the cytocentric map of each target cell;

[0111] Calculate ray angle based on quantity;

[0112] A preset ray starting from the target cell center of the target cell is used as an initial ray, and multiple rays starting from the target cell center are generated one by one based on the ray angle.

[0113] In an embodiment of the present application, the number of spatial regions contained in the cell center map of each target cell can be artificially defined, for example, it can be divided into an even number of spatial regions such as 12, 16, 18, etc. On the one hand, an even number of spatial regions can facilitate the calculation of the angle between two adjacent rays. On the other hand, since the plane rectangular coordinate system contains four quadrants, an even number of spatial regions can make each quadrant contain the same number of complete spatial regions, which is convenient for the subsequent statistics of the number of unique molecular identifiers in each spatial region.

[0114] In the present embodiment, the ray angle is the angle between adjacent rays in the cell center diagram. The angle is calculated by dividing the number of degrees of a circle by the number of corresponding spatial regions. For example, if the number of divided spatial regions is determined to be 16, the angle between two adjacent rays is the ratio of 360 degrees to 16, that is, the angle between the two adjacent rays is 22.5 degrees.

[0115] In the embodiment of the present application, the initial ray can be a semi-axis (positive semi-axis or negative semi-axis) of the horizontal or vertical axis of the plane rectangular coordinate system. According to the calculated angle, multiple rays with the cell center as the common starting point are generated one by one until the generated rays return to the initial ray after passing a circle, thus completing the division of the spatial region.

[0116] Please refer to Figure 3 , Figure 3 The figure shows a schematic diagram of a cell-centric map consisting of 16 spatial regions. Starting with the positive half of the vertical axis as the initial ray, the spatial regions are divided clockwise, resulting in a cell-centric map consisting of 16 spatial regions. Numbering each spatial region further enables the identification of gene expression inflection points in different spatial regions, facilitating gene expression data statistics and calculation of the signal-to-noise ratio in each spatial region.

[0117] Please refer to Figure 4 In some embodiments, step 103 includes but is not limited to the following steps 401 to 403.

[0118] Step 401, calculating the gene expression values ​​in a plurality of fan-shaped regions corresponding to different radii in each spatial region based on the spatial transcriptome sequencing data;

[0119] Step 402, calculating the signal-to-noise ratio of each sector-shaped area based on the gene expression value and the radius of each sector-shaped area;

[0120] Step 403 : determining the gene expression inflection point of the corresponding spatial region according to the target sector region with the highest signal-to-noise ratio.

[0121] In the embodiment of the present application, multiple spatial regions are divided, and each spatial region corresponds to a direction. Taking any two adjacent rays as an example, the distance from any point on the ray to the center point of the cell is d, and the unit of d is pixel. The range of d can be set according to actual needs. Here, d is greater than or equal to 1 and less than or equal to 20 as an example, and d takes an integer value. Taking the center of the cell as the starting point, a line segment of length d is intercepted on two adjacent rays. The spatial region uses adjacent line segments of the same length as the radius, the angle between the two adjacent rays as the center angle, and the fan-shaped area enclosed by the arc connecting the ends of the radius. Please refer to Figure 5 , Figure 5 This is a schematic diagram of a sector-shaped area with a radius of d in the space area numbered t.

[0122] Further, the gene expression values ​​in the fan-shaped area are calculated according to the spatial transcriptome sequencing data. When the RNA diffusion in the cell is not serious, there is a change inflection point in the number of RNA distribution from the cell to the outside of the cell, and the number of gene expression molecules in the cell is usually higher than the number of gene expression molecules outside the cell, and the cell type can also be determined according to the number and type of gene expression molecules in the cell. At the inflection point, the gene expression amount changes significantly, and therefore, the cell boundary can be determined by the position of the change point, and the cell area and cell type determined by combining the spatial position of the cell and the gene expression molecules of the corresponding spatial position, make full use of the spatial transcriptome data and cell segmentation results, but do not rely entirely on the cell morphology obtained by the cell segmentation result, so that the cell area determined is also more accurate.

[0123] Among them, the gene expression value of a cell can be characterized by the number of unique molecular identifiers. A unique molecular identifier is a small random sequence, typically 10-12 nucleotides in length, which is used to distinguish homologous sequences of different gene expression molecules. The Stereo-seq spatiotemporal transcriptome sequencing chip has ultra-high spatial resolution, and multiple DNA nanoballs (DNBs) can be used to capture gene expression molecules in a cell. Each DNA nanoball on the sequencing chip is connected to hundreds or even thousands of capture oligonucleotide chains, and the UMI sequence on each capture oligonucleotide chain is different. The capture oligonucleotide chain captures the mRNA, which is then reverse transcribed to generate cDNA. After a series of operations such as linker removal and amplification, it can become a sequencing library for the machine. Therefore, the number of unique molecular identifiers in a spatial region can represent the gene expression value in that region.

[0124] In the embodiment of the present application, step 401 includes but is not limited to the following steps:

[0125] Based on the spatial transcriptome sequencing data, the gene expression values ​​of the sampling points in each sector area and the Pearson correlation coefficient between each sampling point and the adjacent sampling points in each sector area were calculated;

[0126] The average Pearson correlation coefficient of each sampling point in the fan-shaped area is calculated based on the Pearson correlation coefficient of each sampling point and the adjacent sampling points;

[0127] The gene expression values ​​of each sampling point whose gene expression value is higher than the first preset threshold and whose average Pearson correlation coefficient is higher than the second preset threshold are summarized to obtain the gene expression value in each sector area.

[0128] In the embodiment of the present application, each sector contains multiple DNB sampling points, each of which has hundreds or even thousands of capture oligonucleotide chains connected to UMI sequences. By counting the number of UMI sequences at each DNB sampling point, the gene expression value at each sampling point is statistically calculated.

[0129] Because the concentration of the corresponding RNA in the cell segmentation results obtained by cell segmentation is far from sufficient for downstream analysis. The low concentration of RNA results in low readings of the sequencing results, which cannot detect all the gene expression information in the cell. The resulting sequencing results are not representative and comparable, affecting the accuracy of gene expression analysis. In addition, for spatial transcriptome sequencing data, there is a contamination problem with UMI technology. For example, the counts of highly expressed genes are low, which leads to incorrect grouping and classification of cells in differential expression analysis or cluster analysis. Ideally, spatial transcriptome sequencing technology captures mRNA in cells through microarray chips, and the UMI sequence of a given point identifies the expression of the gene at that point. However, the gene expression molecules actually captured at each sampling point do not represent the expression of the gene at that point. The mRNA flowing out from adjacent sampling points will cause contamination of the UMI counts, and the obtained gene expression information comes from multiple cells.

[0130] Therefore, it is necessary to preset a first threshold and a second threshold. Specifically, the first threshold and the second threshold can be set based on human experience. The first threshold is the value that the number of UMIs at each sampling point must reach, and is used to filter out sampling points with a small number of UMIs. Setting the first threshold can prevent interference from data from some sampling points that cannot be used for downstream analysis. The Pearson correlation coefficient represents the similarity of gene expression at different sampling points. The second threshold is used to filter out the number of UMIs counted at sampling points whose average Pearson coefficient does not reach the corresponding threshold. This ensures that the gene expression molecules in the resulting cells have stronger correlation and gene expression similarity, thereby reducing the interference of gene expression molecules collected by sampling points in other cells on gene expression molecules collected by sampling points in the target cell.

[0131] Specifically, the number of UMIs at the sampling points in each sector is counted. The average Pearson correlation coefficient of each sampling point is calculated by calculating the Pearson correlation coefficient between each sampling point and adjacent sampling points. The number of UMIs at each sampling point is compared with a first threshold, and the average Pearson correlation coefficient of each sampling point is compared with a second threshold. Sampling points that meet both the first and second thresholds are selected. The number of UMIs at the sampling points that meet the conditions in each sector is counted to obtain the gene expression value in the sector. By setting a threshold to filter sampling points, UMI count contamination can be reduced, improving the accuracy of gene expression values ​​in each sector, thereby improving the accuracy of determining gene expression inflection points and ultimately improving the accuracy of cell region determination.

[0132] In the embodiment of the present application, step 402 includes but is not limited to the following steps:

[0133] Calculate the ratio of gene expression value in each sector area to the corresponding radius;

[0134] The signal-to-noise ratio of each sector area is determined based on the ratio.

[0135] Within any spatial region of the cell center map, the number of UMIs in the sector-shaped regions corresponding to different radius lengths is calculated according to the preset radius length intervals. For example, if the number of spatial regions is 16, the central angle of the sector-shaped region in any spatial region is 22.5 degrees. The radius d is selected as 5 pixels. The spatial resolution of the spatiotemporal transcriptome sequencing chip is fixed, and the value corresponding to the standard length unit of the radius can be calculated based on the correspondence between the chip's spatial resolution and pixels.

[0136] Furthermore, based on the gene expression value in each sector determined in the above steps and the radius value corresponding to the sector, the ratio of the two is calculated. This ratio is the signal-to-noise ratio value in the sector, thereby obtaining multiple signal-to-noise ratios in a certain spatial area. The signal-to-noise ratio calculation formula corresponding to a sector is as follows:

[0137]

[0138] Where ∑UMI is the total number of UMIs in the sector area, d t is the radius length corresponding to the sector area, and t represents the number corresponding to the spatial area in a certain direction of the division.

[0139] In the embodiment of the present application, step 403 includes but is not limited to the following steps:

[0140] Determine the center point of the arc corresponding to the target sector area with the highest signal-to-noise ratio;

[0141] The inflection point of gene expression in the corresponding spatial region is determined according to the center point of the arc.

[0142] Specifically, based on the multiple signal-to-noise ratios in each spatial region calculated in the previous step, the multiple signal-to-noise ratios are compared to determine the target sector region with the highest signal-to-noise ratio. The arc midpoint of the target sector region is used as the gene expression inflection point in the corresponding spatial region, and the gene expression inflection points of all spatial regions in the cell center map are determined in sequence. Figure 6 , in some embodiments, Figure 6 Schematic diagram of gene expression inflection points. The gray dots in the figure are the arc midpoints of each target sector area, i.e., the gene expression inflection points in the corresponding spatial area.

[0143] In the embodiment of the present application, step 104 includes but is not limited to the following steps:

[0144] Construct the minimum convex polygon containing multiple gene expression inflection points;

[0145] The first cell region of the target cell is determined according to the minimum convex polygon.

[0146] Please refer to Figure 7 , in some embodiments, Figure 7 This is a schematic diagram of the first cell region fitted using the minimum convex polygon. The first cell region is the cell region of the final cell. Specifically, the cell boundary shape is relatively complex and is usually not a standard shape, such as a rectangle or circle. Fitting the cell boundary using a standard shape is prone to errors. Using a convex polygon to fit the cell boundary can better adapt to changes in cell morphology. Among the multiple convex polygons that fit the inflection point of gene expression, the convex polygon with the smallest area is used as the cell boundary. The fitting result is closer to the actual morphology of the cell, and the fitted cell size is close to the actual size of the cell, which improves the accuracy of cell region determination.

[0147] In some embodiments, after obtaining the first cell region of the target cell based on the inflection point fitting of multiple gene expression levels, the cell region determination method provided by the present application further includes:

[0148] Determining gene expression data in target cells based on the first cell region and spatial transcriptome sequencing data, and assigning the gene expression data to cells;

[0149] The cell type of the target cells is determined based on the gene expression data of the target cells.

[0150] Specifically, the location of the cell is obtained according to the determined cell area, and the gene expression spectrum in the corresponding cell area is obtained based on the location of the cell and the spatial transcription sequencing data. The gene expression spectrum usually refers to the expression data of all genes in one or more cells, as well as the corresponding sample information. The sample information includes cell number, sample source, strain, etc. The matrix composed of this information is the gene expression matrix. In the spatial transcriptome sequencing data, the gene expression data is the gene expression matrix corresponding to each cell. Each row of the gene expression matrix represents a gene, each column represents a cell, and each element in the matrix represents the expression level of the corresponding gene in the corresponding cell.

[0151] After determining the gene expression data in each cell, it is necessary to compare it with a reference standard to obtain a more accurate cell type. The reference standard needs to be obtained through analysis and comparison based on a large amount of single-cell RNA sequencing data. For example, a clustering algorithm can be used to perform cluster analysis on this type of sequencing data, clustering cells into different subpopulations. The cell type is determined based on the discovery patterns and characteristics of gene expression in different subpopulations. After obtaining the reference standard, the measured data can be compared with the standard to determine the cell type to which the target cell belongs.

[0152] The following is a detailed description of the gene variation detection method provided by the present application using a specific embodiment.

[0153] Please refer to Figure 8 , in some embodiments, Figure 8 This is a schematic diagram of the experimental results of applying the cell region determination method of this application to 135D1 mouse brain data. The figure shows the determination results of cell boundaries and gene expression molecules on the spatiotemporal transcriptome sequencing chip. The white lines represent cell boundaries determined based on the inflection points of gene expression changes, the gray areas represent gene expression molecules, and the cross lines in the figure represent the track lines on the chip. As can be seen in the figure, the molecular weight of gene expression inside the cell is significantly higher than the molecular weight of gene expression outside the cell, and the cell region fitted by the convex polygon conforms to the actual cell morphology.

[0154] Please refer to Figure 9 , in some embodiments, Figure 9 Figure 1 is a schematic diagram of the cell annotation results after correction by the cell region determination method of the application. Clustering and annotation results of different types of cells are shown in the figure, and it can be seen that the difference between each tissue in the mouse brain cell is obvious, and different types of cell positions are reasonable. According to the cell region determination method correction of the application, the mouse brain data provided by the online neuron analysis and visualization platform can be compared, and the platform provides the cell type and its corresponding position in the mouse brain, as well as information such as Marker genes. Marker genes refer to genes with unique expression patterns for cells or tissues, which can be used to distinguish different types of cells or tissues, or for detecting the difference in gene expression or protein expression. According to the type of manual inspection cell annotation and the distribution position of different cells, it can be known that the cell annotation type obtained by the cell region determination method of the application is accurate, and the position of different cell types is accurate, further illustrating that the method provided by the application improves the accuracy of determining the gene expression molecules in the cell, and then improves the accuracy determined by the cell region.

[0155] Please refer to Figure 10A 、 Figure 10B , in some embodiments, Figure 10A This is a schematic diagram of the results of clustering mouse brain cell DGGRC2 according to one embodiment of the present application. Figure 10B This is a schematic diagram of the results of clustering mouse brain cell DGGRC2 based on the Fast correction method.

[0156] Please refer to Figure 11A 、 Figure 11B , in some embodiments, Figure 11A This is a schematic diagram of the results of MBDOP2 clustering in mouse brain cells according to one embodiment of the present application. Figure 11BThis is a schematic diagram of the results of MBDOP2 clustering in mouse brain cells based on the Fast correction method.

[0157] Please refer to Figure 12A 、 Figure 12B , in some embodiments, Figure 12A This is a schematic diagram of the results of TEGLU3 clustering in mouse brain cells according to one embodiment of the present application. Figure 12B This is a schematic diagram of the results of TEGLU3 clustering in mouse brain cells based on the Fast correction method.

[0158] Please refer to Figure 13A 、 Figure 13B , in some embodiments, Figure 13A This is a schematic diagram of the results of clustering TEGLU8 in mouse brain cells according to one embodiment of the present application. Figure 13B This is a schematic diagram of the results of clustering TEGLU8 in mouse brain cells based on the Fast correction method.

[0159] Please refer to Figure 14A 、 Figure 14B , in some embodiments, Figure 14A This is a schematic diagram of the results of clustering TEGLU6 in mouse brain cells according to one embodiment of the present application. Figure 14B This is a schematic diagram of the results of TEGLU6 clustering in mouse brain cells based on the Fast correction method.

[0160] Specifically, the above clustering results are obtained by the scanpy clustering function, and DGGRC2, MBDOP2, TEGLU3, TEGLU8, and TEGLU6 represent a type of mouse brain cell respectively. Each type of cell usually has a specific position or forms a specific tissue shape, and the cell position of the same type is relatively concentrated, and there is rarely a dispersed situation. The rapid correction method is a method for single-cell sequencing data analysis, and the cell outline obtained based on cell segmentation is expanded outward to obtain the gene expression molecular area of ​​the cell. It can automatically correct inaccurate information in the sequencing data, such as sample batch effect, sequencing plate difference, etc., so that the result of cell cluster analysis is more accurate. Clustering is carried out for cells of the same type, and the clustering results are corrected according to the rapid correction method. The threshold value for cell identification and segmentation is set to 10 pixels, and the clustering results obtained using the cell area method provided in the present application embodiment. By comparing the two clustering results, it can be learned that for the clustering result of DGGRC2 cells, cells are more concentrated in the clustering result obtained in the present application embodiment, and the tissue shape formed is clearer. For the clustering results of TEGLU6 cells, the clustering results obtained in the examples of the present application are more concentrated and numerous than those obtained by the rapid correction method, and the resulting tissue shape is clearer, and the cell morphology can be seen more clearly, which is closer to the actual distribution of cells. It is further explained that the cell area determined by the gene expression inflection point provided by the present application is more accurate, can better divide RNA gene expression molecules into cells, and the resulting cell number and distribution are also more reasonable.

[0161] Please refer to Figure 15 , in some embodiments, Figure 15 Schematic diagram of cell size distribution in the brain of 135D1 mice.

[0162] Specifically, in Figure 15 In the example, the units corresponding to the horizontal and vertical coordinates are both pixels. The cell size and cell distribution obtained using the embodiment of the present application are similar to the cell size and cell distribution obtained using the rapid correction method, indicating that the cell area size determined by the present embodiment is reasonable and consistent with the actual size of the cells. In addition, after calculating the determined cell area, it can be found that the average cell size obtained by the method provided in the embodiment of the present application is 12% smaller than the average cell size of the cells obtained by the rapid correction method. Therefore, the cell area obtained by the cell area determination method provided in the present application is consistent with the actual size of the cells and also consistent with the actual distribution of the cells.

[0163] As shown in Table 1 below, the comparison results of the original RNA segmentation results and the B02304A2 clustering effect evaluation parameters of the inflection point of single gene expression are shown; the better the evaluation parameters represented by the bold values ​​in the table, the better the clustering effect. B02304A2 is a specific single-cell RNA sequencing data set or the sample number used in a study. Clustering is to apply the sample data to a clustering algorithm, such as the scanpy clustering algorithm, to classify cells with similar gene expression patterns into one category, and different groups represent different cell subtypes or types. The figure shows the clustering evaluation parameters obtained by dividing gene expression molecules into cells based on the cell segmentation results in the related art, as well as the clustering evaluation parameters obtained by dividing gene expression molecules into cells based on the cell area determined by the gene expression inflection point provided in this application. For the convenience of calculation here, a single gene expression inflection point is used to calculate the clustering evaluation parameters. According to the comparison results, the silhouette coefficient obtained by clustering the cell regions determined in the embodiment of the present application is greater than the silhouette coefficient obtained by the original RNA segmentation result, and is closer to 0, indicating that the compactness and separation of the clustering are better, the similarity of the data points in the same cluster is higher, and the clustering effect is better; in addition, the Karlinsky-Harabas score and Dunn index calculated from the inflection point of the expression of a single gene are also better than the Karlinsky-Harabas score and Dunn index obtained from the original RNA segmentation result, which further illustrates that the clustering effect obtained by clustering the cell regions determined according to the inflection point of the gene expression is better, the cell type division is more accurate, and the distinction between tissues composed of different cells is more obvious.

[0164] Clustering effect evaluation parameters Raw RNA segmentation results Inflection point of single gene expression Silhouette coefficient -0.229102674 -0.189760737 Davis Bolding scores 7.095641655 7.39122761 Kalinsky-Harabas score 0.940174852 1.017554255 Dunn Index 1.524571E-05 1.524900E-05

[0165] Table 1 Comparison results of B02304A2 clustering effect evaluation parameters

[0166] The cell region determination device provided in the embodiments of the present application is described below.

[0167] Reference Figure 16 In some embodiments, the present application further provides a cell region determination device 1600, which includes:

[0168] An acquisition unit 1601 is configured to acquire a staining image and spatial transcriptome sequencing data of a biological sample, and determine cell centers of multiple cells based on the staining image;

[0169] The cell center map construction unit 1602 is used to construct a corresponding cell center map for any target cell. The cell center map includes a plurality of rays that divide the stained image into a plurality of spatial regions. The plurality of rays have the target cell center of the target cell as a common starting point.

[0170] A gene expression inflection point determination unit 1603 is configured to determine a gene expression inflection point in each spatial region based on the spatial transcriptome sequencing data, thereby obtaining multiple gene expression inflection points of target cells. The gene expression inflection point is the point with the highest signal-to-noise ratio in the corresponding spatial region.

[0171] The cell region fitting unit 1604 is configured to obtain a first cell region of the target cell according to the inflection points of the multiple gene expression levels.

[0172] In one embodiment, the gene expression level inflection point determination unit includes:

[0173] The first calculation subunit is used to calculate the gene expression values ​​in a plurality of fan-shaped regions corresponding to different radii in each spatial region based on the spatial transcriptome sequencing data;

[0174] The second calculation subunit is used to calculate the signal-to-noise ratio of each sector-shaped area based on the gene expression value and the radius of each sector-shaped area;

[0175] The first determination subunit is used to determine the gene expression inflection point of the corresponding spatial region according to the target sector region with the highest signal-to-noise ratio.

[0176] In one embodiment, the first determining subunit includes:

[0177] The first determining module is used to determine the center point of the arc corresponding to the target sector area with the highest signal-to-noise ratio;

[0178] The second determination module is used to determine the inflection point of the gene expression level in the corresponding spatial area according to the center point of the arc.

[0179] In one embodiment, the second computing subunit includes:

[0180] The first calculation module is used to calculate the ratio of the gene expression value in each sector area to the corresponding radius;

[0181] The third determining module is configured to determine the signal-to-noise ratio of each sector area according to the ratio.

[0182] In one embodiment, the first computing subunit includes:

[0183] The second calculation module is used to calculate the gene expression value of the sampling point in each sector area and the Pearson correlation coefficient between each sampling point and the adjacent sampling points in each sector area based on the spatial transcriptome sequencing data;

[0184] The third calculation module is used to calculate the average Pearson correlation coefficient of each sampling point in the fan-shaped area according to the Pearson correlation coefficient of each sampling point and the adjacent sampling points;

[0185] The summarizing subunit is used to summarize the gene expression value of each sampling point whose gene expression value is higher than the first preset threshold and whose average Pearson coefficient is higher than the second preset threshold, to obtain the gene expression value in each sector area.

[0186] In one embodiment, the cell-centric map construction unit includes:

[0187] The ray drawing subunit is used to draw multiple rays from any target cell, with the target cell center of the target cell as a common starting point, and the angles between any two adjacent rays are equal;

[0188] The spatial region division subunit is used to divide the stained image into multiple spatial regions according to multiple rays to obtain the cell center map of the target cell.

[0189] In one embodiment, the ray extraction subunit includes:

[0190] a fourth determining module, configured to determine the number of spatial regions included in the cell center map of each target cell;

[0191] A fourth calculation module is used to calculate the ray angle based on the quantity, where the ray angle is the angle between adjacent rays in the cell center map;

[0192] The ray generation module is used to use a preset ray based on the target cell center of the target cell as the initial ray, and to generate multiple rays based on the ray angle one by one with the target cell center as the starting point.

[0193] In one embodiment, the cell region fitting unit includes:

[0194] The minimum convex polygon construction subunit is used to construct the minimum convex polygon containing multiple gene expression inflection points;

[0195] The second determining subunit is configured to determine a first cell region of the target cell according to the minimum convex polygon.

[0196] In one embodiment, the acquiring unit includes:

[0197] a third determining subunit, configured to determine a second cell region of the plurality of cells based on the staining image;

[0198] The fourth determining subunit is configured to determine a plurality of cell centers according to the geometric centers of the plurality of second cell regions.

[0199] In one embodiment, the acquiring unit further includes:

[0200] a cell nucleus center determination subunit, used to determine the cell nucleus centers of multiple cells based on staining images;

[0201] The fifth determining subunit is used to determine multiple cell centers based on multiple cell nucleus centers.

[0202] In one embodiment, after obtaining the first cell region of the target cell based on the inflection point fitting of the multiple gene expression levels, the cell region determining device is further configured to:

[0203] Determining gene expression data in target cells based on the first cell region and spatial transcriptome sequencing data, and assigning the gene expression data to cells;

[0204] The cell type of the target cells is determined based on the gene expression data of the target cells.

[0205] It can be seen that the contents of the above-mentioned cell region determination method embodiment are all applicable to the embodiment of the present cell region determination device. The functions specifically implemented by the present cell region determination device embodiment are the same as those in the above-mentioned cell region determination method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned cell region determination method embodiment.

[0206] Reference Figure 17 , Figure 17 The hardware structure of a computer device according to another embodiment is shown. The computer device includes:

[0207] The processor 1701 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0208] The memory 1702 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1702 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1702 and is called by the processor 1701 to execute the cell region determination method of the embodiments of this application.

[0209] Input / output interface 1703, used to implement information input and output;

[0210] Communication interface 1704, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0211] Bus 1705 , which transmits information between various components of the device (e.g., processor 1701 , memory 1702 , input / output interface 1703 , and communication interface 1704 );

[0212] The processor 1701 , the memory 1702 , the input / output interface 1703 and the communication interface 1704 are connected to each other in communication within the device via a bus 1705 .

[0213] The present application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device implements the above-mentioned cell region determination method.

[0214] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "comprises" and "comprising," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.

[0215] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0216] It should be understood that in the description of the embodiments of the present application, multiple (or multiple items) means more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.

[0217] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0218] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0219] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0220] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0221] It should also be understood that the various implementation methods provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0222] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A method for determining a cell region, characterized in that: include: Acquire a staining image and spatial transcriptome sequencing data of a biological sample, and determine the cell centers of a plurality of cells based on the staining image; For any target cell, construct a corresponding cell center map, wherein the cell center map includes a plurality of rays that divide the stained image into a plurality of spatial regions, and the plurality of rays have a target cell center of the target cell as a common starting point; Determining a gene expression inflection point in each spatial region according to the spatial transcriptome sequencing data to obtain multiple gene expression inflection points of the target cells, wherein the gene expression inflection point is a point with the highest signal-to-noise ratio in the corresponding spatial region; The first cell region of the target cell is obtained according to the inflection points of the multiple gene expression levels.

2. The cell region determination method according to claim 1, wherein: The determining the gene expression inflection point in each spatial region according to the spatial transcriptome sequencing data to obtain multiple gene expression inflection points of the target cell includes: Calculating gene expression values ​​in a plurality of fan-shaped regions corresponding to different radii in each spatial region according to the spatial transcriptome sequencing data; Calculating the signal-to-noise ratio of each sector based on the gene expression value and the radius of each sector; The gene expression inflection point in the corresponding spatial region is determined based on the target sector region with the highest signal-to-noise ratio.

3. The method for determining a cell region according to claim 2, wherein determining the gene expression inflection point of the corresponding spatial region based on the target sector region with the highest signal-to-noise ratio comprises: Determine the center point of the arc corresponding to the target sector area with the highest signal-to-noise ratio; The gene expression inflection point of the corresponding spatial region is determined according to the center point of the arc.

4. The cell region determination method according to claim 2, characterized in that: The calculating the signal-to-noise ratio of each sector-shaped area based on the gene expression value and the radius of each sector-shaped area includes: Calculate the ratio of gene expression value in each sector area to the corresponding radius; The signal-to-noise ratio of each sector area is determined according to the ratio.

5. The cell region determination method according to claim 2, characterized in that: The step of calculating the gene expression values ​​in a plurality of fan-shaped regions corresponding to different radii in each spatial region based on the spatial transcriptome sequencing data includes: Calculate the gene expression value of the sampling point in each sector area and the Pearson correlation coefficient between each sampling point and the adjacent sampling points in each sector area according to the spatial transcriptome sequencing data; Calculate the average Pearson correlation coefficient of each sampling point in the fan-shaped area based on the Pearson correlation coefficient of each sampling point and the adjacent sampling points; The gene expression values ​​of each sampling point whose gene expression value is higher than the first preset threshold and whose average Pearson coefficient is higher than the second preset threshold are summarized to obtain the gene expression value in each of the sector-shaped areas.

6. The cell region determination method according to claim 1, characterized in that: The method of constructing a corresponding cell center map for any target cell includes: For any target cell, taking the target cell center of the target cell as a common starting point, deriving multiple rays based on the common starting point, and arranging the angles between any two adjacent rays being equal; The stained image is divided into multiple spatial regions according to the multiple rays to obtain a cell center map of the target cell.

7. The cell region determination method according to claim 6, characterized in that: For any target cell, the target cell center of the target cell is used as a common starting point, and multiple rays are drawn based on the common starting point, including: Determine the number of spatial regions contained in the cytocentric map of each target cell; Calculating a ray angle based on the quantity, the ray angle being the angle between adjacent rays in the cell center map; A preset ray starting from the target cell center of the target cell is used as an initial ray, and multiple rays starting from the target cell center are generated one by one based on the ray angles.

8. The cell region determination method according to claim 1, characterized in that: The obtaining the first cell region of the target cell according to the inflection points of the multiple gene expression levels by fitting includes: Constructing a minimum convex polygon containing the multiple gene expression inflection points; A first cell region of the target cell is determined according to the minimum convex polygon.

9. The cell region determination method according to claim 1, characterized in that: Determining the cell centers of the plurality of cells based on the staining image comprises: determining a second cell region of the plurality of cells based on the stained image; A plurality of cell centers are determined based on the geometric centers of the plurality of second cell regions.

10. The cell region determination method according to claim 1, characterized in that: The determining of the cell centers of the plurality of cells based on the stained image further comprises: determining the nucleus centers of a plurality of cells based on the stained image; A plurality of cell centers are determined based on the plurality of cell nuclear centers.

11. The cell region determination method according to any one of claims 1 to 10, characterized in that: After obtaining the first cell region of the target cell according to the inflection points of the multiple gene expression levels, the method further includes: Determining gene expression data in the target cells based on the first cell region and the spatial transcriptome sequencing data, and assigning the gene expression data to cells; The cell type of the target cell is determined based on the gene expression data of the target cell.

12. A cell region determination device, characterized in that: include: an acquisition unit, configured to acquire a staining image and spatial transcriptome sequencing data of a biological sample, and determine cell centers of a plurality of cells based on the staining image; a cell center map construction unit, configured to construct a corresponding cell center map for any target cell, wherein the cell center map includes a plurality of rays that divide the stained image into a plurality of spatial regions, and the plurality of rays have a target cell center of the target cell as a common starting point; a gene expression inflection point determination unit, configured to determine a gene expression inflection point in each spatial region based on the spatial transcriptome sequencing data, to obtain a plurality of gene expression inflection points of the target cell, wherein the gene expression inflection point is a point with the highest signal-to-noise ratio in the corresponding spatial region; A cell region fitting unit is used to obtain a first cell region of the target cell according to the inflection points of the multiple gene expression levels.

13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the cell region determination method according to any one of claims 1 to 11 is implemented.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the cell region determination method according to any one of claims 1 to 11 is implemented. 15 . A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the cell region determination method according to claim 1 .