Cell region identification method and device, electronic equipment and storage medium

By combining spatial transcription data and stained images, different algorithms are used to generate cell images and match images, the problem of adhesion phenomenon in cell-intensive area recognition is solved and the accuracy of recognition is improved.

CN120088776APending Publication Date: 2025-06-03SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311643320.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is prone to cell boundary adhesion when identifying cell-intensive areas in spatial group sections, which affects subsequent analysis.

Method used

By combining spatial transcription data and stained images, first and second cell images are generated using different recognition algorithms and image matching is performed to identify target cell regions.

Benefits of technology

It improves the accuracy of cell region recognition and reduces cell region adhesions, especially in dense areas, the recognition effect is better than traditional watershed algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088776A_ABST
    Figure CN120088776A_ABST
Patent Text Reader

Abstract

The invention provides a cell region recognition method and device, electronic equipment and a storage medium. The cell region recognition method comprises the steps of performing cell region recognition on a target section based on spatial transcription data of the target section and a preset first recognition algorithm to obtain a first cell image; performing cell region identification on the target section based on the dyed image of the target section and a preset second identification algorithm to obtain a second cell image; and carrying out image matching on the first cell image and the second cell image to obtain a target cell region of the target section. According to the embodiment of the invention, adhesion of the identified cell region can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of cell region recognition, and in particular, to a method and device for cell region recognition, an electronic device, and a storage medium. Background Art

[0002] Currently, cell region recognition is an important step in spatial cell research.

[0003] In the related art, the watershed algorithm can be used to recognize and segment cell regions. However, when recognizing regions where cells are closely adjacent in a spatial transcriptomics section based on this algorithm, the recognition result is prone to cell boundary adhesion, which will affect subsequent analysis of cell regions. Therefore, how to provide a cell region recognition method that can reduce the adhesion situation has become an urgent technical problem to be solved. Summary of the Invention

[0004] The main objective of the embodiments of the present application is to propose a method and device for cell region recognition, an electronic device, and a storage medium, aiming to reduce the adhesion of recognized cell regions.

[0005] To achieve the above objective, a first aspect of the embodiments of the present application proposes a method for cell region recognition, and the method includes:

[0006] Performing cell region recognition on the target section based on the spatial transcriptomic data of the target section and a preset first recognition algorithm to obtain a first cell image;

[0007] Performing cell region recognition on the target section based on the staining image of the target section and a preset second recognition algorithm to obtain a second cell image;

[0008] Performing image matching on the first cell image and the second cell image to obtain the target cell region of the target section.

[0009] In some embodiments, the first cell image includes a first baseline corresponding to the sequencing chip, and the second cell image includes a second baseline corresponding to the sequencing chip;

[0010] The performing image matching on the first cell image and the second cell image to obtain the target cell region of the target section includes:

[0011] Aligning the first baseline with the second baseline to obtain the overlapping probability between the first cell region in the first cell image and the second cell region in the second cell image;

[0012] Screening the first cell region based on the overlapping probability and a preset probability threshold to obtain the target cell region.

[0013] In some embodiments, the spatial transcriptomic data includes gene expression levels and gene location data;

[0014] Performing cell region recognition on the target slice based on the spatial transcriptomic data of the target slice and a preset first recognition algorithm to obtain a first cell image, including:

[0015] Constructing an expression grayscale image according to the gene expression levels and the gene location data;

[0016] Performing cell region recognition on the target slice according to the expression grayscale image and the first recognition algorithm to obtain the first cell image.

[0017] In some embodiments, the method further includes:

[0018] Calculating a coloring score based on the cell types of the target slice and the coloring order of a preset color palette;

[0019] Screening a target color palette from the preset color palette according to the coloring score, and performing display processing on the target cell region corresponding to the cell type according to the coloring order of the target color palette.

[0020] In some embodiments, calculating the coloring score based on the cell types of the target slice and the coloring order of a preset color palette includes:

[0021] Performing slicing processing on the target slice according to a preset slice window to obtain sub-slices;

[0022] Determining the neighbor cells of the cells corresponding to each cell type in the sub-slices, and obtaining the number of adjacent cells according to the cell types corresponding to the neighbor cells;

[0023] Performing data splicing on the number of adjacent cells according to the slice window to obtain adjacent cell data;

[0024] Calculating the coloring score according to the adjacent cell data and the coloring order.

[0025] In some embodiments, the preset color palette includes preset colors, and calculating the coloring score according to the adjacent cell data and the coloring order includes:

[0026] Calculating the distance between the preset color and the origin in a preset color space to obtain a first color difference of the preset color;

[0027] Calculating the coloring score according to the first color difference, the coloring order, and the adjacent cell data.

[0028] In some embodiments, screening the target color palette from the preset color palette according to the coloring score includes:

[0029] Comparing a plurality of the coloring scores to obtain a first comparison result;

[0030] Taking the preset color palette corresponding to the coloring score with the largest value in the first comparison result as the target color palette.

[0031] In some embodiments, the preset color palette includes preset colors. Calculating the coloring score based on the cell type of the target slice and the coloring order of the preset color palette includes:

[0032] Calculating the distance between any two of the cell types to obtain a type distance;

[0033] Calculating the distance between any two of the preset colors based on the coloring order to obtain a second color difference;

[0034] Calculating the coloring score based on the second color difference and the type distance.

[0035] In some embodiments, screening the target color palette from the preset color palette according to the coloring score includes:

[0036] Comparing a plurality of the coloring scores to obtain a second comparison result;

[0037] Taking the preset color palette corresponding to the coloring score with the smallest value in the second comparison result as the target color palette.

[0038] In some embodiments, before calculating the coloring score based on the cell type of the target slice and the coloring order of the preset color palette, the method further includes: determining the cell type, including:

[0039] Obtaining the expression data corresponding to each cell in the target slice based on the spatial transcript data;

[0040] Screening the cells in the target slice based on the expression data, and performing normalization processing on the expression data corresponding to the screened cells to obtain standard data;

[0041] Performing dimensionality reduction processing on the standard data to obtain key data;

[0042] Performing clustering processing on the cells corresponding to the key data to obtain the cell type.

[0043] To achieve the above object, a second aspect of the embodiments of the present application provides a cell region recognition device, and the device includes:

[0044] The first recognition module is configured to recognize the cell region of the target slice based on the spatial transcriptomic data of the target slice and a preset first recognition algorithm, so as to obtain a first cell image;

[0045] The second recognition module is configured to recognize the cell region of the target slice based on the stained image of the target slice and a preset second recognition algorithm, so as to obtain a second cell image;

[0046] The image matching module is configured to perform image matching on the first cell image and the second cell image to obtain the target cell region of the target slice.

[0047] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.

[0048] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0049] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer program product. The computer program product includes a computer program, and when the computer program is read and executed by a processor of a computer device, the computer device executes the method described in the first aspect.

[0050] The cell region recognition method, device, electronic device, and storage medium provided by the embodiments of the present application obtain a first cell image through a first recognition algorithm and spatial transcriptomic data, and obtain a second cell image through a second recognition algorithm and a stained image. It can be seen from this that the first cell image and the second cell image are obtained by processing different data using different algorithms. Therefore, when performing image matching according to the first cell image and the second cell image to obtain the target cell region, the accuracy of the target cell region can be improved, and the situation of adhesion of the target cell region can be reduced. Description of the Drawings

[0051] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. They are used to explain the technical solutions of the present disclosure together with the embodiments of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.

[0052] Figure 1 is a flowchart of the cell region recognition method provided by the embodiments of the present application;

[0053] Figure 2 is Figure 1 a flowchart of step S101 in

[0054] Figure 3 is Figure 1 the flowchart of step S103 in

[0055] Figure 4 is the flowchart of another embodiment of the cell region recognition method provided by the embodiments of the present application;

[0056] Figure 5 is the flowchart of the cell type determination method provided by the embodiments of the present application;

[0057] Figure 6 is Figure 4 the flowchart of step S401 in

[0058] Figure 7 is Figure 6 the flowchart of step S604 in

[0059] Figure 8 is Figure 4 the flowchart of step S402 in

[0060] Figure 9 is Figure 4 the flowchart of another embodiment of step S401 in

[0061] Figure 10 is Figure 4 the flowchart of another embodiment of step S402 in

[0062] Figure 11 is a schematic diagram of displaying and processing the cell region according to the default parameters;

[0063] Figure 12 is a schematic diagram of displaying and processing according to the coloring order corresponding to the target color palette of the embodiments of the present application;

[0064] Figure 13 is a schematic diagram of the cell region recognition device provided by the embodiments of the present application;

[0065] Figure 14 is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0066] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0067] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0068] Spatial Transcriptomics (ST) sequencing technology: It is a technology for analyzing RNA at the spatial level and is used to analyze all RNAs in a single tissue section. Each slide for library construction in 10x Genomics Spatial Transcriptomics has four capture regions, and the size of each capture region is 6.5 x 6.5 mm. Each capture region contains 5000 barcoded spots, the diameter of each spot is 55 μm, and the center-to-center distance between spots is 100 μm. Each spot includes multiple capture probes that can bind to RNA, and each probe is attached with a unique spatial barcode to label the spatial position of the captured RNA. By performing sequencing on the machine, the sequence of each RNA transcript can be mapped back to its original position in the tissue section, thereby providing data support for downstream tasks. It can be seen that the ST sequencing technology is a technology that can simultaneously obtain RNA expression levels and RNA spatial position information in a single experiment.

[0069] Spatial Transcriptomics with Error-Corrected Barcoding and Sequencing (Stereo-seq): It is a method that combines spatial transcriptomics and single-cell RNA sequencing technologies. It can analyze gene expression levels in tissues at spatial resolution and improve the accuracy and sensitivity of gene expression levels through error-corrected barcoding and sequencing methods. Stereo-seq includes the design of DNA nanoball (DNB) patterned array chips, in situ sequencing to determine the spatial coordinates of unique barcode oligonucleotides, dot ligation of UMI-polyT oligonucleotides, in situ capture of RNA from tissues, cDNA amplification, library construction sequencing, and data analysis. Among them, DNB sequencing is based on an in situ sequencing patterned array, which can serve as the basis for a spatial resolution transcription technology with high resolution and large field of view. Compared with sequencing methods such as HDST (HDST is a high-resolution spatial transcriptomics technology that carves a large number of 2.05 μm holes on the surface of a glass slide, distributes silica magnetic beads in the holes, captures the mRNA of corresponding cells at the corresponding positions through the magnetic beads for reverse transcription, and conducts library construction and transcriptome sequencing), Slide-seqV2 (Slide-seqV2 is a method for achieving high-sensitivity near-cell-resolution spatial transcriptome analysis, and can obtain better RNA capture efficiency through steps such as magnetic bead synthesis and array indexing), Visium (Visium is a spatial transcriptome sequencing technology that integrates gene expression with immunohistochemical images of tissue sections to localize the gene expression levels of different cells in tissues to the original spatial positions of the tissues), and DBiT-seq (DBiT-seq is a spatial multi-omics sequencing technology based on microfluidic barcoding), Stereo-seq has more spots per 100 μm2, and the spot size and center-to-center distance are smaller, making Stereo-seq have higher resolution.

[0070] Currently, the commonly used cell region recognition and segmentation algorithm is the watershed algorithm. The watershed algorithm can recognize relatively dispersed cells, but has poor recognition effect on cells in dense regions and is prone to adhesion phenomena. That is to say, the recognition effect of the watershed algorithm depends on the quality of the spatial group slices. For slices with relatively poor quality, the number of recognized cells is small and the recognition time is long.

[0071] In addition, in single-cell spatial omics data analysis, different cell regions can be represented by different numbers. The closer the numbers are, the closer the pixel values of the corresponding regions in the image are, making it difficult to distinguish between different cell regions. In the cell clustering results, the number of cells corresponding to different cell types is small and sparsely distributed. Different numbers can be assigned to the cell regions corresponding to different cell types, and the numbers can be defined as the magnitudes of the pixel values of the corresponding cell regions. In a two-dimensional space, when coloring the corresponding cell regions according to the magnitudes of the pixel values and the cell clustering results, it is easy for the colors of the cell regions corresponding to adjacent cell types to be similar, making it difficult to distinguish the distributions of different cell types in the section according to the colors.

[0072] Based on this, the embodiments of the present application propose a cell region recognition method, device, electronic device, and storage medium, which can reduce the adhesion of recognized cell regions.

[0073] The cell region recognition method provided by the embodiments of the present application will be described below.

[0074] Referring to Figure 1 , in some embodiments, the cell region recognition method provided by the embodiments of the present application includes, but is not limited to, steps S101 to S103.

[0075] Step S101: Perform cell region recognition on the target section based on the spatial transcript data of the target section and a preset first recognition algorithm to obtain a first cell image.

[0076] Step S102: Perform cell region recognition on the target section based on the stained image of the target section and a preset second recognition algorithm to obtain a second cell image.

[0077] Step S103: Perform image matching on the first cell image and the second cell image to obtain the target cell region of the target section.

[0078] Steps S101 to S103 illustrated in the embodiments of the present application obtain a first cell image through the first recognition algorithm and the spatial transcript data, and obtain a second cell image through the second recognition algorithm and the stained image. It can be seen from this that the first cell image and the second cell image are obtained by processing different data using different algorithms. Therefore, when performing image matching according to the first cell image and the second cell image to obtain the target cell region, the accuracy of the target cell region can be improved, and the adhesion of the target cell region can be reduced.

[0079] In step S101 of some embodiments, the target slice refers to the tissue slice to be subjected to cell region recognition. The target slice is a biological slice obtained in compliance with relevant legal regulations, and the biological slice includes plant slices and animal slices. Spatial transcriptomic data is obtained by performing ST processing on the target slice, or spatial transcriptomic data is obtained according to Stereo-seq. The first recognition algorithm refers to an algorithm that is preset and can perform cell region recognition on the target slice based on the spatial transcriptomic data. For example, the first recognition algorithm can be the IcellSeg algorithm. According to the first recognition algorithm, a first cell image of the target slice can be obtained, and the first cell image can be used to represent the cell regions corresponding to each cell in the target slice.

[0080] Referring to Figure 2 , in some embodiments, the spatial transcriptomic data includes gene expression levels and gene location data. Step S101 includes, but is not limited to, steps S201 to S202.

[0081] Step S201, constructing an expression grayscale image based on the gene expression levels and gene location data;

[0082] Step S202, performing cell region recognition on the target slice according to the expression grayscale image and the first recognition algorithm to obtain a first cell image.

[0083] In step S201 of some embodiments, taking the spatial transcriptomic data obtained by ST processing as an example, the gene expression level can refer to the RNA expression level captured from the corresponding position of each spot in the target slice. The gene location data can refer to the spatial location information of the captured RNA expression level. The gene expression levels and gene location data can be visually processed, such as displaying the gene expression levels in the form of a grayscale image on the corresponding spatial distribution of the target slice according to the gene location data, to obtain an expression grayscale image. In the expression grayscale image, different gene expression levels can be represented based on different shades of gray, for example, light color represents high expression, and dark color represents low expression, etc.

[0084] In step S202 of some embodiments, the first recognition algorithm can be the IcellSeg algorithm. The IcellSeg algorithm is a morphological-based image segmentation method, such as determining the cell region through methods such as binarization, edge detection, and morphological operations. The expression grayscale image is used as the processing data of the first recognition algorithm to perform cell region recognition on the expression grayscale image to obtain a first cell image corresponding to the target slice. It can be understood that before performing cell region recognition on the expression grayscale image based on the first recognition algorithm, preprocessing such as balancing, filtering, and threshold processing can also be performed on the expression grayscale image to remove noise in the expression grayscale image, and the embodiments of the present application do not make specific limitations on this.

[0085] The advantages of steps S201 to S202 are that the expression grayscale image can reflect the gene expression levels at different spatial positions of the target slice, that is, each pixel in the expression grayscale image corresponds to the gene expression level at the corresponding position of the target slice. Moreover, the expression grayscale image can intuitively display the spatial expression differences of genes. Therefore, when the cell region is recognized based on the first recognition algorithm for the expression grayscale image, the accuracy of cell region recognition can be improved, thereby reducing the situation of cell region adhesion.

[0086] In step S102 of some embodiments, the stained image of the target slice may refer to the image obtained after staining the target slice. Among them, the staining method for the target slice includes conventional staining, immunohistochemical staining, immunofluorescent staining, etc. Among them, conventional staining refers to staining different cells and tissue structures in tissue sections to facilitate the observation and analysis of cell tissue morphology and tissue characteristics. Conventional staining includes H&E staining, etc. Immunohistochemical staining refers to detecting the expression and localization of specific proteins in tissue sections by using specific antibodies conjugated with markers. Immunofluorescent staining refers to using fluorescently labeled antibodies to detect the expression and legal localization of specific proteins in tissue sections. The application embodiments do not specifically limit the staining method for the target slice. However, for the sake of illustration, in the following embodiments, the stained image is taken as an ssDNA stained image for illustration. It can be understood that ssDNA staining belongs to immunohistochemical staining because ssDNA staining can detect and visualize the presence and distribution of single-stranded DNA by using specific antibodies conjugated with markers. The second recognition algorithm refers to a pre-set algorithm that can recognize the cell region of the target slice based on the stained image. For example, the second recognition algorithm can be the Stardist algorithm. According to the second recognition algorithm, the second cell image of the target slice can be obtained, and the second cell image can be used to represent the cell regions corresponding to each cell in the target slice.

[0087] It can be understood that the Stardist algorithm is a deep learning algorithm for cell segmentation. The core idea of the Stardist algorithm is to use a convolutional neural network to learn and predict cell boundaries, and at the same time combine morphological analysis methods to improve the accuracy of cell segmentation.

[0088] In step S103 of some embodiments, image matching of the first cell image and the second cell image may refer to aligning the first cell image and the second cell image to match multiple cell regions in the first cell image with multiple cell regions in the second cell image. When a cell region in the first cell image matches a cell region in the second cell image, it can be indicated that the two cell regions correspond to the same cell region of the target slice. According to the degree of matching, the target cell region can be screened from the cell region corresponding to the first cell image or the second cell image. The target cell region can be used to indicate the region in the target slice that the cell actually corresponds to. Based on the target cell region, downstream biological function analysis and the like can be implemented.

[0089] Reference Figure 3 In some embodiments, the first cell image includes a first baseline corresponding to the sequencing chip, and the second cell image includes a second baseline corresponding to the sequencing chip. Step S103 includes but is not limited to steps S301 to S302.

[0090] Step S301, aligning the first baseline with the second baseline to obtain the overlap probability of the first cell region in the first cell image and the second cell region in the second cell image;

[0091] Step S302: Screen the first cell region based on the overlap probability and a preset probability threshold to obtain a target cell region.

[0092] In step S301 of some embodiments, the sequencing chip refers to a chip used for ST processing or Stereo-seq processing. In the sequencing chip, there are regions in which there are no probes and the shapes are regular, and these regions can present a cross shape. Since the expression grayscale image and the staining image obtained according to the spatial transcription data are from the same sequencing chip, the first cell image and the second cell image both include a cross-shaped baseline. Therefore, the first baseline can be used as the alignment coordinate system of the first cell image, and the second baseline can be used as the alignment coordinate system of the second cell image. When the first baseline is aligned with the second baseline, the alignment of the first cell image and the second cell image can be achieved. According to the aligned first cell image and the second cell image, a plurality of corresponding matching cell region pairs can be obtained, and these cell region pairs can be used to indicate the same cell region in the target slice. It can be understood that the cell region pair can be expressed as the first cell region-the second cell region. Among them, the first cell region can be a cell region indicated from the first cell image, and the second cell region can be a cell region indicated from the second cell image. Determine the probability of regional overlap between the first cell region and the second cell region in each cell region pair.

[0093] In step S302 of some embodiments, the preset probability threshold may refer to a pre-set probability threshold. The specific value of the preset probability threshold can be adaptively set according to the actual situation, and the embodiments of the present application do not make specific limitations thereto. For example, it can be set to 95%. Compare the overlap probability with the preset probability threshold, retain the cell region pairs with an overlap probability greater than or equal to the preset probability threshold, and filter out the cell region pairs with an overlap probability less than the preset probability threshold. It can be understood that since the first cell region and the second cell region in the cell region pair are obtained according to different recognition algorithms, the accuracy of the retained cell region pair corresponding to the true cell region of the target section is relatively high.

[0094] It can be understood that the first cell region in the retained cell region pair can be used as the target cell region, or the second cell region in the cell region pair can be used as the target cell region, or the overlapping region of the first cell region and the second cell region can be used as the target cell region. The embodiments of the present application do not make specific limitations thereto.

[0095] It can be understood that in some embodiments, after determining the target cell region of the target section by the above method, the target cell region can also be visualized based on the cell type to intuitively display the spatial distribution of different cell types in the target section. The visualization method will be described in detail below.

[0096] Refer to Figure 4 , in some embodiments, the cell region recognition method provided by the embodiments of the present application further includes but is not limited to steps S401 to S402.

[0097] Step S401, calculating a coloring score based on the cell type of the target section and the coloring order of the preset color palette;

[0098] Step S402, screening the target color palette from the preset color palette according to the coloring score, and performing display processing on the target cell region corresponding to the cell type according to the coloring order of the target color palette.

[0099] In step S401 of some embodiments, different target cell regions may correspond to different or the same cell type. The cell type corresponding to the target cell region can be determined by methods such as cell annotation. The preset color palette is a pre-set color palette including different preset colors. Different preset colors in the preset color palette can be arranged in a certain order, and the target cell regions corresponding to different cell types can be colored according to the arrangement order. That is to say, the arrangement order of the preset colors can be used as the coloring order for the target cell region. For example, the target section includes C 1 , C 2 ,..., C n a total of n cell types, corresponding to including A1 , A 2 ,..., A n There are a total of n preset colors. The coloring order can be A 1 , A 3 , A 4 ..., or it can be A 2 , A 5 , A 1 ..., or it can be other orders. When the coloring order is A 1 , A 3 , A 4 ..., it means that the cell region corresponding to cell type C can be colored based on the preset color A 1 for the cell region corresponding to cell type C 1 , and the cell region corresponding to cell type C can be colored based on the preset color A 3 for the cell region corresponding to cell type C 2 , and the cell region corresponding to cell type C can be colored based on the preset color A 4 for the cell region corresponding to cell type C 3 . It can be seen that different coloring orders can obtain different coloring effects. The coloring score can be calculated based on the cell type and the coloring order, and the coloring score is used to measure the coloring effect. In the embodiments of the present application, it can be considered that the coloring effect that can distinguish the target cell regions corresponding to different cell types is a better coloring effect.

[0100] Referring to Figure 5 , in some embodiments, the method for determining the cell type includes but is not limited to steps S501 to S504.

[0101] Step S501: Obtain the expression data corresponding to each cell in the target section based on the spatial transcript data;

[0102] Step S502: Screen the cells in the target section based on the expression data, and perform normalization processing on the expression data corresponding to the screened cells to obtain standard data;

[0103] Step S503: Perform dimensionality reduction processing on the standard data to obtain key data;

[0104] Step S504: Perform clustering processing on the cells corresponding to the key data to obtain the cell type.

[0105] In step S501 of some embodiments, the expression data is used to represent the expression level of RNA corresponding to each cell in the target section. The expression data can be obtained from the spatial transcript data. It can be understood that the expression data can be represented in the form of an expression matrix. When represented in the form of an expression matrix, the rows are used to represent the cells in the target section, the columns are used to represent the RNAs, and the values corresponding to the rows and columns are used to represent the RNA expression levels.

[0106] In step S502 of some embodiments, cells in the target slice are screened based on the expression data to filter out cells with less expression data. The expression data corresponding to the remaining cells after filtering is normalized to obtain the standard data corresponding to each cell. Among them, the normalization process refers to converting the expression data according to certain rules to reduce the technical differences or batch effects between different slices, so that the expression data between different slices is comparable and interpretable. The normalization process includes TPM normalization (TPM normalization is to divide the expression data by the total number of transcripts and then multiply by one million), RPKM normalization (RPKM normalization means dividing the expression data by the length of the corresponding gene and the total number of transcripts, and then dividing by one million), FPKM normalization (FPKM normalization means dividing the expression data by the length of the corresponding gene and the total number of fragments, and then multiplying by one million), etc. The embodiments of the present application do not make specific limitations on this.

[0107] In step S503 of some embodiments, dimensionality reduction processing is performed on the standard data to select data with a small number of dimensions to represent the overall data. The selected data can be used as key data. The method of dimensionality reduction processing can be the Principal Component Analysis (PCA).

[0108] In step S504 of some embodiments, clustering processing is performed on the corresponding cells based on the key data and the clustering method, and the cell types corresponding to the cells are obtained according to the clustering results. It can be understood that the clustering method can be louvain clustering (louvain clustering is a clustering algorithm for graph data, and the basic idea of the louvain clustering algorithm is to divide nodes into different clusters through iteration), leiden clustering (leiden clustering is an improvement and expansion of louvain clustering, and leiden clustering introduces a local optimization strategy), etc. The embodiments of the present application do not make specific limitations on this.

[0109] In step S402 of some embodiments, it can be understood that by arranging the preset colors in different orders, different preset color palettes can be obtained. Different preset color palettes can correspond to different coloring scores. The preset color palettes are screened according to the coloring scores, and the preset color palette with the best coloring effect is used as the target color palette. The target cell region corresponding to the cell type is colored and displayed according to the coloring order of the target color palette, thereby realizing the visualization processing of the target cell region.

[0110] The advantages of steps S401 to S402 are that the preset color palette can be filtered according to the coloring score. When the target cell area is displayed according to the coloring order corresponding to the filtered target color palette, the situation where the colors of the target cell areas corresponding to adjacent cell types are similar can be reduced, and the distinguishability of cell types can be improved.

[0111] It can be understood that when the calculation method of the coloring score is different, the criteria for filtering the target color palette according to the coloring score are also different. Two different calculation methods of the coloring score, and the corresponding filtering methods, are introduced below. It can be understood that the method for calculating the coloring score in the embodiments of the present application is not limited to the following two methods.

[0112] First, the first method is introduced.

[0113] Referring to Figure 6 , in some embodiments, step S401 includes but is not limited to steps S601 to S604.

[0114] Step S601, perform a slicing process on the target slice according to a preset slice window to obtain sub-slices;

[0115] Step S602, determine the neighboring cells of the cells corresponding to each cell type in the sub-slice, and obtain the number of adjacent cells according to the cell types corresponding to the neighboring cells;

[0116] Step S603, perform data splicing on the number of adjacent cells according to the slice window to obtain adjacent cell data;

[0117] Step S604, calculate the coloring score according to the adjacent cell data and the coloring order.

[0118] In step S601 of some embodiments, the preset slice window is a window that is preset for performing a slicing process on the target slice to slice the target slice into multiple sub-slices. The size of the preset slice window can be adaptively set according to the size of the target slice, and the embodiments of the present application do not make specific limitations thereto. For example, the size of the preset slice window can be set to 100.

[0119] In step S602 of some embodiments, each sub-slice may include multiple cell types, and each cell type may correspond to at least one cell. For example, the sub-slice includes three cell types B, D, and E. Cell type B includes cells B1 and B2, cell type D includes cells D1, D2, and D3, and cell type E includes cell E1. Taking cell B1 as an example, the neighbor cells of cell B1 refer to the cells that are adjacent to cell B1 among cells B2, D1, D2, D3, and E1. It can be understood that it can be determined whether two cells are adjacent according to distance, morphological characteristics of the cells, etc. Determine the cell type corresponding to each neighbor cell, so that the number of adjacent cells of cell B1 can be obtained. It can be seen that the number of adjacent cells is used to represent the number of neighbor cells belonging to a certain cell type.

[0120] It can be understood that according to the number of adjacent cells, the adjacent matrix M shown in the following formula (1) can be constructed k :

[0121]

[0122] where N ij represents the total number of neighbor cells of each cell in the i-th cell type belonging to the j-th cell type. For example, N n1 represents the total number of neighbor cells of each cell in the n-th cell type belonging to the 1st cell type. Therefore, N ij can be calculated according to the number of adjacent cells.

[0123] In step S603 of some embodiments, different adjacent matrices M can be obtained according to different sub-slices k . According to the slice window, multiple adjacent matrices M k are data-stitched to obtain the adjacent cell data corresponding to the entire target slice. It can be seen that the adjacent cell data can reflect the number of neighbor cells between different cell types in the target slice.

[0124] In step S604 of some embodiments, the corresponding coloring score is calculated according to the adjacent cell data and the coloring order.

[0125] The advantages of steps S601 to S604 are that the coloring score calculated according to the adjacent cell data and the coloring order can reflect the relationship between the number of neighbor cells between different cell types and the coloring order, so that when the target color palette is screened and displayed according to the coloring score later, the distinguishability between different cell types can be improved.

[0126] Referring to Figure 7 , in some embodiments, step S604 includes but is not limited to steps S701 to S702.

[0127] Step S701: Calculate the distance between the preset color and the origin in the preset color space to obtain the first color difference of the preset color.

[0128] Step S702: Calculate the coloring score based on the first color difference, the coloring order, and the adjacent cell data.

[0129] In step S701 of some embodiments, the preset color space refers to a pre-set model for describing and representing colors. The preset color space can be an RGB color space, a CMYK color space, an HSV color space, an LAB color space, etc., and the embodiments of the present application do not make specific limitations thereto. For the sake of illustration, in the embodiments of the present application, the LAB color space is taken as an example for illustration. The origin refers to the color where L = 0 (L represents brightness), a = 0 (a represents the component from red to green), and b = 0 (b represents the component from yellow to blue), that is, it represents a completely colorless neutral gray. Calculate the distance between the preset color and the origin, and use this distance as the first color difference of the preset color. It can be understood that the distance between the preset color and the origin can be calculated according to methods such as Euclidean distance, Manhattan distance, Chebyshev distance, etc., and the embodiments of the present application do not make specific limitations thereto.

[0130] In step S702 of some embodiments, the coloring score score is calculated according to the first color difference, the coloring order, the adjacent cell data, and the following formula (2).

[0131] score = M * P i , P i ∈ Permutation(L 1 , L 2 ,..., L n )...... Formula (2)

[0132] where M represents the adjacent cell data, and L 1 to L n all represent the first color difference, and the sorting of L 1 to L n corresponds to the coloring order.

[0133] The advantages of steps S701 to S702 are that the coloring score calculated based on the first color difference can reflect the influence of the attributes of different preset colors in the LAB space (including brightness attributes, attributes of the component from red to green, and attributes of the component from yellow to blue) on cell types with different adjacency relationships.

[0134] Refer to Figure 8, in some embodiments, "screening a target color palette from a preset color palette according to the coloring score" in step S402 includes but is not limited to steps S801 to S802.

[0135] Step S801, comparing multiple coloring scores to obtain a first comparison result;

[0136] Step S802, using the preset color palette corresponding to the coloring score with the largest value in the first comparison result as the target color palette.

[0137] In steps S801 to S802 of some embodiments, compare the coloring scores corresponding to different coloring orders to obtain a first comparison result. Based on the first comparison result, screen out the coloring score with the largest value, and use the preset color palette corresponding to the coloring score with the largest value as the target color palette.

[0138] Secondly, introduce the second method.

[0139] Refer to Figure 9 , in some other embodiments, step S401 includes but is not limited to steps S901 to S903.

[0140] Step S901, calculate the distance between any two cell types to obtain a type distance;

[0141] Step S902, calculate the distance between any two preset colors based on the coloring order to obtain a second color difference;

[0142] Step S903, calculate the coloring score based on the second color difference and the type distance.

[0143] In step S901 of some embodiments, an adjacency matrix about cell types can be constructed according to cell types. In this adjacency matrix, both rows and columns represent cell types, and the values corresponding to rows and columns are used to represent the distance between the corresponding two cell types, that is, the type distance. It can be understood that the type distance can be calculated by methods such as Euclidean distance, Manhattan distance, Chebyshev distance, etc., and the embodiments of the present application do not make specific limitations on this.

[0144] In step S902 of some embodiments, similar to step S901, a color difference matrix about preset colors can be constructed according to the distance between preset colors. In this color difference matrix, both rows and columns represent preset colors, and the values corresponding to rows and columns are used to represent the second color difference between the corresponding two preset colors. Among them, the setting of rows and columns can be related to the coloring order. It can be understood that the distance between the corresponding two preset colors can be used as the second color difference. Similarly, the methods for calculating the distance between two preset colors include Euclidean distance, Manhattan distance, Chebyshev distance, etc., and the embodiments of the present application do not make specific limitations on this.

[0145] In step S903 of some embodiments, distance calculation is performed based on the matrix form of the second color difference (i.e., the color difference matrix) and the matrix form of the type distance (i.e., the adjacency matrix), and the calculated distance is used as the coloring score. It can be understood that the method for calculating the distance between the above two matrices is similar to the method for calculating the second color difference, and the embodiments of the present application will not elaborate on this. It can be understood that when the color difference matrix and the adjacency matrix are more similar, the distance between the two matrices is smaller, that is, the coloring score is smaller.

[0146] It can be understood that one of the matrices can be fully permuted. For example, if the color difference matrix is fully permuted, color difference matrices corresponding to different coloring orders can be obtained.

[0147] The advantages of steps S901 to S903 are that the calculated coloring score can reflect the similarity between the second color difference and the type distance.

[0148] Referring to Figure 10 , in some other embodiments, in some embodiments, "screening the target color palette from the preset color palette according to the coloring score" in step S402 includes but is not limited to steps S1001 to S1002.

[0149] Step S1001: Compare multiple coloring scores to obtain a second comparison result;

[0150] Step S1002: Use the preset color palette corresponding to the coloring score with the smallest value in the second comparison result as the target color palette.

[0151] In steps S1001 to S1002 of some embodiments, the coloring scores calculated according to different coloring orders are compared to obtain a second comparison result. Based on the second comparison result, the coloring score with the smallest value is selected, and the preset color palette corresponding to the coloring score with the smallest value is used as the target color palette.

[0152] Referring to Figure 11 and Figure 12 , in a specific embodiment, taking the spatial transcriptome data of E16.5 downloaded from MOSTA as an example, the recognition effect and visualization effect of the embodiments of the present application are described. It can be understood that in this spatial transcriptome data, the expression data of 281,377 cells and 28,103 genes are included. When performing display processing according to the default parameters in scanpy (scanpy is a library for single-cell RNA sequencing data analysis) in the related art, a spatial distribution map of different cell types as shown in Figure 11 can be obtained. When using the preset colors provided by scanpy and performing display processing according to the coloring order of the target color palette determined by the above method, a result as shown in Figure 12Spatial distribution diagrams of different cell types shown. From Figure 11 and Figure 12 comparison, it can be seen that the method provided by the embodiments of the present application can better reflect the differences between different cell types. For example, the cell regions corresponding to the cell type Ganglion can be compared. For the cell type Ganglion, in Figure 11 , it is almost difficult to visually discover this cluster of cell types located in the cortical epidermis.

[0153] The cell region recognition method provided by the embodiments of the present application, which recognizes cell regions through ssDNA staining images and spatial transcript data, can reduce the adhesion phenomenon and improve the accuracy of cell region recognition. In the recognition of dense regions, the recognition effect of the embodiments of the present application is better than that of the watershed algorithm. Moreover, the method of displaying the target cell regions by the coloring score in the embodiments of the present application can improve the distinguishability between different cell types, thus facilitating subsequent biological analysis.

[0154] Next, the cell region recognition device provided by the embodiments of the present application will be described.

[0155] Referring to Figure 13 , the embodiments of the present application also provide a cell region recognition device, which includes:

[0156] A first recognition module 1301, configured to perform cell region recognition on the target slice based on the spatial transcript data of the target slice and a preset first recognition algorithm to obtain a first cell image;

[0157] A second recognition module 1302, configured to perform cell region recognition on the target slice based on the staining image of the target slice and a preset second recognition algorithm to obtain a second cell image;

[0158] An image matching module 1303, configured to perform image matching on the first cell image and the second cell image to obtain the target cell region of the target slice.

[0159] It can be seen that the contents in the above embodiments of the cell region recognition method are all applicable to the embodiments of this cell region recognition device. The functions specifically implemented by the embodiments of this cell region recognition device are the same as those of the above embodiments of the cell region recognition method, and the beneficial effects achieved are also the same as those of the above embodiments of the cell region recognition method.

[0160] Referring to Figure 14 , Figure 14 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0161] The processor 1401 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0162] The memory 1402 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1402 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1402 and are called by the processor 1401 to execute the cell region recognition method of the embodiments of the present application;

[0163] The input / output interface 1403 is used to implement information input and output;

[0164] The communication interface 1404 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0165] The bus 1405 transmits information between various components of the device (such as the processor 1401, the memory 1402, the input / output interface 1403, and the communication interface 1404);

[0166] Among them, the processor 1401, the memory 1402, the input / output interface 1403, and the communication interface 1404 achieve communication connections with each other inside the device through the bus 1405.

[0167] The embodiments of the present application also provide a computer program product, which includes a computer program. The processor of the computer device reads and executes this computer program, so that the computer device executes to implement the above-mentioned cell region recognition method.

[0168] In the description of the present disclosure and the above-mentioned accompanying drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0169] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a, b, and c", where a, b, and c can be single or multiple.

[0170] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.

[0171] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0172] The unit described as a separate component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0173] In addition, each functional unit in various embodiments of the present disclosure may be integrated in a processing unit, may exist physically separately as individual units, or two or more units may be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0174] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0175] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0176] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. A method for identifying cell regions, characterized in that, the method includes: Performing cell region identification on the target section based on the spatial transcriptomic data of the target section and a preset first identification algorithm to obtain a first cell image; Performing cell region identification on the target section based on the stained image of the target section and a preset second identification algorithm to obtain a second cell image; Performing image matching on the first cell image and the second cell image to obtain the target cell region of the target section.

2. The method according to claim 1, characterized in that, the first cell image includes a first baseline corresponding to the sequencing chip, and the second cell image includes a second baseline corresponding to the sequencing chip; The performing image matching on the first cell image and the second cell image to obtain the target cell region of the target section includes: Aligning the first baseline with the second baseline to obtain the overlapping probability between the first cell region in the first cell image and the second cell region in the second cell image; Based on the overlapping probability and a preset probability threshold, screening the first cell region to obtain the target cell region.

3. The method according to claim 1, characterized in that, the spatial transcriptomic data includes gene expression levels and gene position data; The performing cell region identification on the target section based on the spatial transcriptomic data of the target section and a preset first identification algorithm to obtain a first cell image includes: Constructing an expression grayscale image according to the gene expression levels and the gene position data; Performing cell region identification on the target section according to the expression grayscale image and the first identification algorithm to obtain the first cell image.

4. The method according to claim 1, characterized in that, the method further includes: Calculating a coloring score based on the cell type of the target section and the coloring order of a preset color palette; Screening a target color palette from the preset color palette according to the coloring score, and performing display processing on the target cell region corresponding to the cell type according to the coloring order of the target color palette.

5. The method according to claim 4, characterized in that, The calculating a coloring score based on the cell type of the target section and the coloring order of a preset color palette includes: Performing slicing processing on the target section according to a preset slice window to obtain sub-slices; Determining the neighbor cells of the cells corresponding to each cell type in the sub-slices, and obtaining the number of adjacent cells according to the cell types corresponding to the neighbor cells; Performing data splicing on the number of adjacent cells according to the slice window to obtain adjacent cell data; Calculating the coloring score according to the adjacent cell data and the coloring order.

6. The method according to claim 5, characterized in that, the preset color palette includes preset colors, and the calculating the coloring score according to the adjacent cell data and the coloring order includes: Calculating the distance between the preset color and the origin in a preset color space to obtain the first color difference of the preset color; The coloring score is calculated based on the first color difference, the coloring order, and the adjacent cell data.

7. The method according to claim 6, wherein, the screening of the target color palette from the preset color palette according to the coloring score includes: comparing a plurality of the coloring scores to obtain a first comparison result; using the preset color palette corresponding to the coloring score with the largest value in the first comparison result as the target color palette.

8. The method according to claim 4, wherein, the preset color palette includes preset colors, and the calculation of the coloring score based on the cell type of the target slice and the coloring order of the preset color palette includes: calculating the distance between any two of the cell types to obtain a type distance; calculating the distance between any two of the preset colors based on the coloring order to obtain a second color difference; calculating the coloring score based on the second color difference and the type distance.

9. The method according to claim 8, wherein, the screening of the target color palette from the preset color palette according to the coloring score includes: comparing a plurality of the coloring scores to obtain a second comparison result; using the preset color palette corresponding to the coloring score with the smallest value in the second comparison result as the target color palette.

10. The method according to any one of claims 4 to 9, wherein, before calculating the coloring score based on the cell type of the target slice and the coloring order of the preset color palette, the method further includes: determining the cell type, including: obtaining the expression data corresponding to each cell in the target slice based on the spatial transcript data; screening the cells in the target slice based on the expression data, and performing normalization processing on the expression data corresponding to the screened cells to obtain standard data; performing dimensionality reduction processing on the standard data to obtain key data; performing clustering processing on the cells corresponding to the key data to obtain the cell type.

11. A cell region recognition device, wherein, the device includes: a first recognition module, configured to perform cell region recognition on the target slice based on the spatial transcript data of the target slice and a preset first recognition algorithm to obtain a first cell image; a second recognition module, configured to perform cell region recognition on the target slice based on the staining image of the target slice and a preset second recognition algorithm to obtain a second cell image; an image matching module, configured to perform image matching on the first cell image and the second cell image to obtain the target cell region of the target slice.

12. An electronic device, including a memory and a processor, the memory storing a computer program, wherein, when the processor executes the computer program, the cell region recognition method according to any one of claims 1 to 10 is implemented.

13. A computer-readable storage medium, the storage medium storing a computer program, wherein, when the computer program is executed by a processor, the cell region recognition method according to any one of claims 1 to 10 is implemented.

14. A computer program product, comprising a computer program, which is read and executed by a processor of a computer device, so that the computer device executes the cell region recognition method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Function unit prediction model construction method, prediction method, device and electronic equipment

    CN120431990A