Method for determining cell gene expression data and electronic equipment

By processing stained images and screening target cell regions using a pre-set neural network model, and combining the spatial coordinates of the analysis unit with gene expression data, the problem of inaccurate gene expression data of single cells in existing technologies is solved, achieving higher precision gene expression analysis.

CN121237209APending Publication Date: 2025-12-30BOAO BIOLOGICAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511435397.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing high-resolution spatial transcriptomics (HD) platforms struggle to accurately acquire gene expression data for individual cells when analyzing cell image data, resulting in an inability to precisely analyze the gene expression characteristics of cells.

Method used

The stained images of tissue slices are processed using a pre-defined neural network model. The regions to be identified are determined through preliminary segmentation, and the target cell regions are screened using pre-defined cell staining conditions. The gene expression profile of each target cell region is obtained by combining the spatial coordinates of the analysis unit and gene expression data.

Benefits of technology

This improves the accuracy of cell segmentation and gene expression data, ensuring the precision of gene expression data analysis for individual cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237209A_ABST
    Figure CN121237209A_ABST
Patent Text Reader

Abstract

The invention discloses a cell gene expression data determination method and electronic equipment, and the method comprises the steps: processing a dyeing image of a tissue slice chip through a preset neural network model, and determining each to-be-determined region; screening undetermined regions meeting preset cell staining conditions from the cells to obtain a target cell region; obtaining space coordinates and gene expression data of each analysis unit in the tissue slice chip, wherein each analysis unit comprises expression data of a plurality of genes; determining an analysis unit set associated with each target cell region according to each target cell region and the space coordinates of each analysis unit in the tissue slice chip; and integrating the gene expression data of each analysis unit in each analysis unit set to obtain a gene expression profile of each target cell region on the tissue slice chip. The preset neural network model is adopted to segment the dyed picture to obtain the undetermined area, and the preset cell dyeing condition is utilized to screen the undetermined area, so that the cell segmentation accuracy is improved, and the obtained gene expression profile of the cell is high in accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics analysis, and in particular to a method and electronic device for determining cell gene expression data. Background Technology

[0002] With a deeper understanding of cellular heterogeneity, traditional transcriptome sequencing can no longer reveal the gene expression characteristics of individual cells. Therefore, methods for gene expression analysis in single cells have emerged, such as single-cell sequencing technology. Single-cell sequencing (SCS) can amplify and sequence mRNA (ribonucleic acid) at the single-cell level.

[0003] Because single-cell spatial transcriptomics analyzes the unique gene expression profile of each cell, it not only preserves information about heterogeneity between cells, but also precisely analyzes the spatial location of cells in tissues and their interactions, providing higher-dimensional biological insights. This enables researchers to gain a deeper understanding of disease mechanisms, developmental processes, immune responses, and other fields, and to discover new cell types and states, thereby promoting the development of precision medicine.

[0004] To achieve gene expression at the single-cell level, current high-resolution spatial transcriptomics (HD) technology provides dense capture points, achieving subcellular resolution and facilitating the discovery of new cell types and microenvironment features. While current HD platforms can achieve high spatial resolution—for example, two adjacent square bins are 2 micrometers apart—each bin may only contain a portion of a single cell or different parts of multiple cells, failing to capture the full gene expression of a single cell. Therefore, when analyzing cell image data using current high-resolution HD platforms, it is difficult to obtain accurate data on gene expression in single cells. Summary of the Invention

[0005] The first aspect of this application provides a method for determining cell gene expression data, including:

[0006] Obtain stained images of tissue slice chips, wherein the tissue slice chips contain a number of cells;

[0007] The stained image is processed by a preset neural network model to determine at least two undetermined regions in the stained image, and the at least two undetermined regions can form at least a part of the stained image;

[0008] Regions that meet preset cell staining conditions are selected from at least two undetermined regions to obtain target cell regions; spatial coordinates and gene expression data of each analysis unit in the tissue slice chip are obtained, each analysis unit contains expression data of multiple genes, and each cell on the tissue slice chip can express the multiple genes;

[0009] Based on the spatial coordinates of each target cell region and each analysis unit in the tissue slice chip, a set of analysis units associated with each target cell region is determined, and the set of analysis units contains at least one analysis unit.

[0010] By integrating the gene expression data of each analysis unit in each set of analysis units, the gene expression profile of each target cell region on the tissue slice chip is obtained.

[0011] A second aspect of this application provides an electronic device, comprising:

[0012] An interface is used to obtain stained images of tissue slice chips, and to obtain the spatial coordinates and gene expression data of each analysis unit in the tissue slice chip. Each analysis unit contains the expression data of multiple genes, and each cell on the tissue slice chip can express the multiple genes. The tissue slice chip contains a number of cells.

[0013] A processor is configured to process the stained image using a preset neural network model, identify at least two undetermined regions in the stained image, the at least two undetermined regions being able to form at least a part of the stained image; filter regions from the at least two undetermined regions that meet preset cell staining conditions to obtain target cell regions; determine a set of analysis units associated with each target cell region based on the spatial coordinates of each analysis unit in the tissue slice chip, the set of analysis units containing at least one analysis unit; and integrate the gene expression data of each analysis unit in each set of analysis units to obtain the gene expression profile of each target cell region on the tissue slice chip.

[0014] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the method for determining cell gene expression data as described in the first aspect or any implementation thereof.

[0015] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0016] The memory is used to store computer programs;

[0017] The processor is used to execute the computer program so that the electronic device can implement the method for determining cell gene expression data in the first aspect or any implementation thereof.

[0018] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to determine the cell gene expression data of the first aspect or any implementation thereof described above.

[0019] This application provides a method for determining cell gene expression data. A stained image of a tissue slice chip is processed using a preset neural network model. Multiple undetermined regions are identified within the stained image, and these regions constitute all or part of the stained image. Preset cell staining conditions are used to filter these undetermined regions to obtain target cell regions. The spatial coordinates and gene expression data of each analysis unit in the tissue slice chip are obtained. Each analysis unit contains expression data for multiple genes, and each cell on the tissue slice chip expresses multiple genes. Based on the spatial coordinates of each target cell region and each analysis unit in the tissue slice chip, a set of analysis units associated with each target cell region is determined. Each set of analysis units contains at least one analysis unit. The gene expression data of each analysis unit in each set of analysis units are integrated to obtain the gene expression profile of each target cell region on the tissue slice chip. After initial segmentation of the stained image using a preset neural network model to obtain multiple undetermined regions, the use of preset cell staining conditions to filter these undetermined regions reduces the misjudgment of the preset neural network model and improves the accuracy of cell segmentation. Therefore, the accuracy of determining the gene expression of each cell in the tissue slice chip is higher. Attached Figure Description

[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0021] Figure 1 This is a flowchart illustrating a method for determining cell gene expression data provided in an embodiment of this application;

[0022] Figure 2 This is a flowchart illustrating the process of processing a stained image using a preset neural network model to determine at least two undetermined regions in the stained image, as provided in the embodiments of this application.

[0023] Figure 3This is a schematic diagram of the preset neural network model provided in this application processing a stained image to obtain a region to be determined;

[0024] Figure 4 This is a schematic diagram of the process provided in this application embodiment for selecting regions that meet preset cell staining conditions from at least two undetermined regions to obtain target cell regions;

[0025] Figure 5 This is a schematic diagram of the process provided in this application embodiment for selecting regions that meet preset cell staining conditions from at least two undetermined regions to obtain an initial cell region;

[0026] Figure 6 This is a schematic diagram of the undetermined area provided in the embodiments of this application;

[0027] Figure 7 This is a schematic diagram of the process of morphologically expanding an initial cell region to obtain a target cell region, provided in an embodiment of this application.

[0028] Figure 8 This is a schematic diagram of one iteration of dilation in a stained image provided in an embodiment of this application;

[0029] Figure 9 This is a schematic diagram of the target cell region provided in the embodiments of this application;

[0030] Figure 10 This is a flowchart illustrating how the set of analysis units associated with each target cell region is determined based on the spatial coordinates of each analysis unit in the tissue slice chip, according to an embodiment of this application.

[0031] Figure 11 This is a flowchart illustrating the process of integrating gene expression data from each analysis unit in each analysis unit set to obtain the gene expression profile of each target cell region on a tissue slice chip, as provided in the embodiments of this application.

[0032] Figure 12 This is a flowchart illustrating how, based on genes, the gene expression data of each analysis unit in the set of analysis units associated with each target cell region are added together to obtain the gene expression data of each target cell region.

[0033] Figure 13 This is another flowchart illustrating the method for determining cell gene expression data provided in the embodiments of this application;

[0034] Figure 14 This is a flowchart illustrating the method for determining cell gene expression data provided in this application in an application scenario;

[0035] Figure 15This is a partial schematic diagram of the segmentation of fresh human tonsil tissue cells provided in an embodiment of this application;

[0036] Figure 16 The diagram shows the statistical values ​​of five evaluation indicators for cell segmentation results of a fresh human tonsil tissue cell sample.

[0037] Figure 17 This is a partial schematic diagram of mouse kidney FFPE tissue cell segmentation provided in the embodiments of this application;

[0038] Figure 18 The diagram shows the statistical values ​​of five evaluation indicators for cell segmentation results of mouse kidney FFPE tissue samples.

[0039] Figure 19 This is a schematic diagram of the structure of an electronic device using a method for determining cell gene expression data, as provided in an embodiment of this application.

[0040] Figure 20 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0041] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0042] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0043] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0044] Reference Figure 1 , Figure 1 This is a flowchart illustrating a method for determining cell gene expression data provided in an embodiment of this application, as shown below. Figure 1As shown in the embodiments of this application, a method for determining cell gene expression data may include steps 401 to 403, which are described in detail below.

[0045] 101. Obtain stained images of tissue sections on microarrays;

[0046] Technicians attach tissue slices onto a chip to obtain a tissue slice chip, which serves as a sample for subsequent processes.

[0047] After the sample is stained, a microscopic image is obtained using a microscope, and the stained image is then photographed.

[0048] The stained image can be a H&E (hematoxylin-eosin staining) image. The characteristics of H&E staining are: hematoxylin, an alkaline staining solution, is blue and can stain the acidic structures of tissues blue-purple, such as cell nuclei appearing blue, cartilage matrix appearing dark blue, and mucus appearing gray-blue; eosin is a chemically synthesized acidic dye that dissociates into negatively charged anions in water, which easily combine with the positive charges in the amino groups of proteins, staining the cytoplasm red. For example, cytoplasm, muscle, and connective tissue will be stained with different degrees of red or pink.

[0049] The tissue slice chip can be a formalin-fixed and paraffin-embedded (FFPE) sample, a fresh frozen (FF) sample, or a formaldehyde-fixed and frozen (FxF) sample.

[0050] The stained image can be a TIF (Tagged Image File Format) or BTF (Bio-Formats Tiles Format) file.

[0051] In one possible implementation, the obtained stained image can be analyzed. If the resolution of the stained image is clear and the stained result is clear, the stained image can be directly processed. If the resolution of the stained image is unclear or the stained result is unclear, such as low image contrast, image that is too dark or too bright, or image outliers, the stained image can be normalized to improve the image quality and filter out abnormal signal values ​​in the stained image.

[0052] Processing software can be pre-installed in the electronic device that performs the method for determining cell gene expression data in this application, and the normalization processing of the staining image can be performed using the processing software.

[0053] As an example, this processing software could be the `utils.normalize` function from the `CSBDeep` library, which adjusts the brightness and contrast of an image based on specified percentiles. The lower limit of the pixel value distribution can be set to 5%, and the upper limit to 95%. The function will ignore extreme pixel values ​​below 5% and above 95% of the image's pixel value distribution and perform a linear mapping on the remaining pixel values ​​(5% to 95% of pixel values).

[0054] As an example, the automatic adjustment (Auto) function of brightness / contrast in ImageJ software is used to automatically optimize brightness and contrast based on the histogram distribution of the image to enhance visual effects.

[0055] 102. Process the stained image using a pre-defined neural network model, and determine at least two undetermined regions in the stained image, wherein the at least two undetermined regions can form at least a part of the stained image;

[0056] The preset neural network model can be a trained convolutional neural network (CNN).

[0057] The stained image is used as input information and fed into the preset neural network model to predict the region corresponding to each cell in the stained image. This region is then used as the region to be determined.

[0058] Since a tissue slice chip can contain several, dozens, or even hundreds or thousands of cells, multiple regions can be identified.

[0059] The preset neural network model is trained using the Cellpose3 method, which is an open-source anatomical segmentation algorithm developed based on Python 3, specifically designed for segmentation tasks of cell images.

[0060] This preset neural network model processes stained images and can analyze the location where each pixel in the stained image converges. By classifying pixels that converge at the same location into the same region to be determined, it achieves the initial segmentation of the stained image.

[0061] In theory, pixels belonging to the same cell converge, while pixels not belonging to any cell do not converge. This allows for the division of pixels corresponding to each cell in a stained image into multiple undetermined regions, each of which represents a portion of the stained image. Of course, the theoretical effect described above may have some errors in practical applications. However, the accuracy of training can be improved during the training of a pre-defined neural network to maximize the convergence of pixels belonging to the same cell and prevent pixels not belonging to any cell from converging.

[0062] Since the preset neural network model is trained using the Cellpose3 method, it can complete the analysis and processing process using H&E staining images as input, simplifying the experimental procedure.

[0063] 103. Select regions that meet the preset cell staining conditions from at least two undetermined regions to obtain the target cell region;

[0064] Using the aforementioned pre-defined neural network model, misjudgments may occur in the undetermined regions identified in the stained image due to the similarity in shape between certain air bubbles or tissue mucosal pores and cells.

[0065] Since the staining of cells and non-cells in the stained image is different, and the staining of components such as the cell nucleus in each cell has a special color, the identified multiple undetermined areas can be further screened by combining preset cell staining conditions to remove misjudgments.

[0066] The preset cell staining conditions can be obtained by statistically analyzing a large number of stained cells, or by setting the color, size, and other aspects of the components within the cells.

[0067] By using the preset cell staining conditions, the multiple undetermined regions are compared sequentially, and the regions that meet the preset cell staining conditions are identified as target cell regions, thereby removing the undetermined regions that are misjudged by the preset neural network model.

[0068] 104. Obtain the spatial coordinates and gene expression data of each analysis unit in the tissue slice chip. Each analysis unit contains the expression data of multiple genes, and each cell on the tissue slice chip expresses multiple genes.

[0069] Each analysis unit contains gene expression data for multiple genes, with each cell on the tissue slice chip expressing these multiple genes. The gene expression set corresponding to this tissue slice chip can be obtained by designing capture probes on the chip to bind to the mRAN within the cells, thus correlating gene expression information with spatial location, and determining the location of each gene expression on the chip.

[0070] In one possible implementation, HD spatial gene expression products can be used. By designing and constructing capture probes with spatial location-specific barcode sequences on a chip, these probes can bind to intracellular mRNA, thereby linking gene expression information with spatial location, locating the specific regions on the chip where gene expression occurs, and providing crucial spatial coordinate information for subsequent analysis.

[0071] Products employing HD spatial gene expression have a massive array of dots set across the entire chip, with each dot serving as an analysis unit (bin). Each dot captures the total RNA (ribonucleic acid) from all cells within its region. This analysis unit can be a fixed-size square region, with sides typically measuring 2 micrometers.

[0072] In one possible implementation, the gene expression set can be obtained in the form of a FASTQ file, and quantitative expression analysis can be performed using software (such as SpaceRanger) to obtain gene expression matrices corresponding to analysis units with different resolutions of 2 micrometers, 8 micrometers, and 16 micrometers. In this application, a 2-micrometer gene expression matrix is ​​used.

[0073] Accordingly, in this embodiment, the obtained gene expression data is a collection of gene expression data corresponding to each analysis unit on the tissue slice chip. Each analysis unit is at the 2-micrometer level, with each unit corresponding to a region with a side length of 2 micrometers. Since 2 micrometers offers higher resolution than 8 micrometers and 16 micrometers, the gene expression data from the 2-micrometer analysis unit better represents the captured reality. Therefore, using the gene expression data corresponding to this 2-micrometer analysis unit results in a more accurate integration of gene expression data for the final single cell.

[0074] A gene sequencer sequences the RNA from tissue slices on an array, obtaining FASTQ sequencing data files. In this application, the HD sequencing data is saved in two FASTQ files. One file contains UMIs (Unique Molecular Identifiers, serving as the ID (Identity document) for each RNA segment) and spatial barcodes (spatial location sequence IDs used to locate the captured RNA sequence on the array), while the other file contains the RNA sequence information. A pair of sequences from the same RNA fragment can be identified by their ID information.

[0075] The aforementioned steps 101-103 are used to describe the process of segmenting the target cell region corresponding to each cell in the stained image. This process and the process of obtaining the spatial coordinates and gene expression data of each analysis unit in the tissue slice chip in step 104 can be in any order and are not limited to the order in this embodiment.

[0076] 105. Based on the spatial coordinates of each target cell region and each analysis unit in the tissue slice chip, determine the set of analysis units associated with each target cell region. The set of analysis units contains at least one analysis unit.

[0077] The stained image corresponds to the tissue slice chip, and the cell distribution in both is exactly the same. Therefore, by matching the spatial coordinates of each analysis unit in the tissue slice chip with the target cell regions identified in the stained image, the set of analysis units associated with each target cell region can be determined. The gene expression data of each analysis unit in this set of analysis units belongs to the cell corresponding to the target cell region.

[0078] 106. Integrate the gene expression data of each analysis unit in each analysis unit set to obtain the gene expression profile of each target cell region on the tissue slice chip.

[0079] An analysis unit corresponds to a very small area in a stained image. A target cell region corresponding to a cell may contain one or more analysis units. By integrating the gene expression data of each analysis unit in each set of analysis units, the gene expression of each target cell region on the tissue slice chip can be obtained. The gene expression of these multiple target cell regions constitutes the gene expression profile of each cell on the tissue slice chip.

[0080] In this embodiment, a preset neural network model is used to process the stained image of a tissue section chip. Multiple undetermined regions are identified within the stained image, and these regions constitute all or part of the stained image. Using preset cell staining conditions, these undetermined regions are filtered to obtain target cell regions. The spatial coordinates and gene expression data of each analysis unit in the tissue section chip are obtained. Each analysis unit contains expression data for multiple genes, and each cell on the tissue section chip expresses multiple genes. Based on the spatial coordinates of each target cell region and each analysis unit in the tissue section chip, a set of analysis units associated with each target cell region is determined. Each set of analysis units contains at least one analysis unit. The gene expression data of each analysis unit in each set of analysis units are integrated to obtain the gene expression profile of each target cell region on the tissue section chip. After the preset neural network model performs initial segmentation of the stained image to obtain multiple undetermined regions, the use of preset cell staining conditions to filter these undetermined regions reduces the misjudgment of the preset neural network model and improves the accuracy of cell segmentation. Therefore, the accuracy of the determined gene expression profile of cells in the tissue section chip is higher.

[0081] Figure 2 This is a flowchart illustrating the process of processing a stained image using a preset neural network model to determine at least two undetermined regions in the stained image, provided in an embodiment of this application. It may include steps 201 to 204, which are described in detail below.

[0082] 201. Process the stained image through a preset neural network model to obtain the horizontal gradient field map, vertical gradient field map and binary image of the stained image. The numerical value in the binary image represents whether any pixel belongs to any cell.

[0083] The preset neural network model is trained using the Cellpose3 method. The preset neural network uses the CCN (Convolutional Neural Network) algorithm to process the input stained image, and obtains the horizontal gradient field map, vertical gradient field map and binary image of the stained image.

[0084] In one possible implementation, the pre-defined neural network can be a U-Net architecture neural network, in which the process from input to output is a downsampling to upsampling process.

[0085] The horizontal gradient field map represents the gradient change of the stained image in the horizontal direction, and the vertical gradient field map represents the gradient change of the stained image in the vertical direction. In this image processing, the gradient is calculated through convolution operations. The convolution kernel is divided into horizontal and vertical directions. The horizontal gradient field map is usually calculated using the horizontal convolution kernel, and the vertical gradient field map is usually calculated using the vertical convolution kernel.

[0086] This binary image displays pixels in binary form. The value corresponding to each pixel in the binary image represents whether the corresponding pixel in the stained image belongs to any cell. For example, in this binary image, 0 can represent that the corresponding pixel does not belong to any cell, and 1 can represent that the corresponding pixel belongs to a cell. White can represent 0, and gray can represent 1.

[0087] 202. Merge the horizontal gradient field map, vertical gradient field map, and binary image of the stained image to obtain the gradient vector field map corresponding to the stained image;

[0088] The horizontal gradient field map, vertical gradient field map, and binary image corresponding to the stained image are combined to obtain a gradient vector field map containing all features of the three images.

[0089] This gradient vector field can be used to construct a dynamical system corresponding to the stained image, in which each pixel has a flow direction. When a pixel belongs to a cell, it flows to a location within that cell. Pixels belonging to different cells flow to different locations, and pixels belonging to the same cell flow to the same location within that cell. If a pixel does not belong to any cell, it does not flow in any direction.

[0090] 203. Based on the gradient vector field map, determine the target location where each pixel in the stained image converges, and the target location corresponds to at least one pixel;

[0091] The dynamical system constructed using this gradient vector field can direct each pixel in an image to its corresponding target location.

[0092] Each cell contains a target location, which comprises at least one pixel in the stained image. These target locations in the stained image can be used to determine individual pixels belonging to the same cell.

[0093] 204. Assign pixels that converge to the same target location to the same instance region to form a region to be determined. Each instance region constitutes the label mask of the stained image.

[0094] The instance region is the region in the stained image corresponding to a cell, determined by a preset neural network model.

[0095] Multiple pixels converging at the same target location are assigned to the same instance region. Each instance region contains pixels that correspond to a region to be determined. Multiple instance regions constitute the label mask of the stained image.

[0096] In the field of cell microscopy image segmentation, label mask is a standardized term and format for uniquely identifying each cell region.

[0097] In one possible implementation, after determining the target location where each pixel converges, the pixels that converge to the same target location are assigned to different instance regions according to different target locations. Each pixel that converges at a target location corresponds to an instance region. The range of each pixel in an instance region within the stained image is taken as a region to be determined. The instance regions together form the label mask of the stained image.

[0098] Figure 3 This is a schematic diagram of a pre-defined neural network model processing a stained image to obtain a region to be determined, provided in an embodiment of this application. It includes a stained image 301, a horizontal gradient field map 302, a vertical gradient field map 303, a binary image 304, a gradient vector field map 305, a gradient tracking map 306, and a label mask map 307. The stained image 301 is input into a preset neural network model to obtain a horizontal gradient field image 302, a vertical gradient field image 303, and a binary image 304. In the binary image 304, gray represents 1 and white represents 0. 1 indicates that the corresponding pixel belongs to a certain cell, and 0 indicates that the corresponding pixel does not belong to any cell. The horizontal gradient field image 302, the vertical gradient field image 303, and the binary image 304 are combined to obtain a gradient vector field image 305. The dynamic system constructed using the gradient vector field converges the pixels in the image to obtain a gradient tracking image 306. The pink lines in the gradient tracking image 306 indicate the convergence direction, and pixels belonging to the same instance region converge to the same position. Multiple instance regions are combined to obtain a label mask image 307. Different colors are used in this mask image to distinguish adjacent different instance regions 3071.

[0099] In one possible implementation, during the training of the pre-defined neural network model, the training samples may include the original microscopic image (stained image) and a label image corresponding to the original microscopic image. The label image pre-labels the affiliation (cell or background) of each pixel. Supervised learning is used to train the pre-defined neural network model. During training, a topological map is generated by simulating diffusion starting from the center of the cell mask in the label image. The spatial gradient smoothing gradient field, pointing from the cell boundary to the cell center, is derived, converting the mask into vector flow data that the neural network can recognize and predict.

[0100] In one possible implementation, the training samples may include training samples of different cell types, so that the preset neural network can process multiple cell types and improve its compatibility.

[0101] In this embodiment, a preset neural network model is used to process the stained image, resulting in a horizontal gradient field map, a vertical gradient field map, and a binary image. The numerical values ​​in the binary image represent whether any pixel belongs to any cell. These three images are then merged to obtain a gradient vector field map corresponding to the stained image. Based on the gradient vector field map, the target location where each pixel in the stained image converges is determined, and this target location corresponds to at least one pixel. Pixels converging at the same target location are assigned to the same instance region to form undetermined regions. These instance regions constitute the label mask of the stained image. The process of obtaining multiple undetermined regions by processing the stained image using a preset neural network model is described in detail. By using the gradient vector field to determine the flow direction of each pixel, the shape of the cell can be predicted with greater accuracy.

[0102] Figure 4 This is a flowchart illustrating the process of selecting regions that meet preset cell staining conditions from at least two undetermined regions to obtain target cell regions, as provided in the embodiments of this application. It may include steps 401 to 402, which are described in detail below.

[0103] 401. Select regions that meet the preset cell staining conditions from at least two undetermined regions to obtain the initial cell region;

[0104] After determining multiple undetermined regions in the stained image using a preset neural network model, these undetermined regions are segmented using the Cellpose3 method by the preset neural network model. The initial cell regions obtained by screening multiple undetermined regions using preset cell staining conditions are part or all of the undetermined regions.

[0105] The process of obtaining the initial cell region by screening in the undetermined region will be described in detail in subsequent embodiments, and will not be described in detail here.

[0106] 402. Morphological expansion of the initial cell region is performed to obtain the target cell region.

[0107] Due to the limitations of the preset neural network model training method, the initial cell region obtained by processing the stained image can be the region corresponding to the cell nucleus. However, the cell also contains cytoplasm. Therefore, in order to obtain the target cell region corresponding to the complete cell, it is necessary to determine the region corresponding to the cytoplasm of the cell in other regions of the stained image.

[0108] Correspondingly, morphological expansion can be used to expand the initial cell region to obtain the target cell region, which is a region containing a complete cell. This complete cell includes the nucleus and cytoplasm, etc.

[0109] In one possible implementation, an expansion distance can be set for each cell, and this distance can be used to morphologically expand the initial cell region. This expansion distance can be a fixed physical distance or a distance calculated based on the initial cell region and a set ratio.

[0110] In this embodiment, regions that meet preset cell staining conditions are first selected from at least two undetermined regions to obtain initial cell regions; then, morphological dilation is performed on the initial cell regions to obtain target cell regions. By performing screening and morphological dilation on the undetermined regions obtained from the staining image processed by the preset neural network model in two steps, the resulting target cell regions are complete cell regions containing the cell nucleus, improving cell segmentation accuracy. Furthermore, the inclusion of complete cell regions in the target cell regions provides more accurate support for subsequent determination of the cell's gene expression profile.

[0111] Figure 5 This is a flowchart illustrating the process of selecting regions that meet preset cell staining conditions from at least two undetermined regions to obtain an initial cell region, as provided in the embodiments of this application. It may include steps 501 to 504, which are described in detail below.

[0112] 501. Determine the number of first pixels that satisfy the first pixel value in the first undetermined region. The first pixel value is the pixel value corresponding to the preset cell nucleus staining. The first undetermined region is any one of at least two undetermined regions.

[0113] In this application, preset cell staining conditions are set from two dimensions: the dimension of cell nucleus staining and the dimension of non-tissue region, so as to determine whether the image content of the region to be determined is a cell from two dimensions.

[0114] In sequence, one of the at least two undetermined regions is identified as the first undetermined region, and it is determined whether the first undetermined region is the target cell region, so as to achieve the screening of each undetermined region in the stained image.

[0115] In order to make a judgment based on the dimension of cell nucleus staining, it is necessary to determine the number of pixels of cell nuclei contained in the first region to be determined.

[0116] In the stained image obtained by H&E staining, the cell nucleus is blue. Accordingly, the blue pixels in the first undetermined region are counted.

[0117] As an example, the RGB (red, green, blue) value corresponding to blue is (0, 0, 255). Pixels with RGB values ​​of (0, 0, 255) are selected from the first undefined region of the stained image to obtain the first number of pixels.

[0118] Since the color of cell nucleus staining may vary in depth, the RGB value of the corresponding color of cell nucleus may not be (0,0,255), but may be (0,0,253), (0,1,255), etc. Therefore, the first pixel value can be adjusted according to the actual situation, and this application does not restrict the value of the first pixel value.

[0119] 502. Determine the total number of pixels corresponding to the first region to be determined;

[0120] The total number of pixels is obtained by counting the number of pixels contained in the first undetermined region.

[0121] 503. Determine the number of second pixels that satisfy the second pixel value in the first undetermined region. The second pixel value is the pixel value corresponding to the unstained image. In the stained image, the unstained pixels belong to the pixels in the non-organization region.

[0122] Since non-organic regions (such as bubbles or holes) are not colored, the determination of the non-organic region dimension can be made by determining whether there are non-organic regions in the first undetermined region.

[0123] In one possible implementation, in the stained image obtained by H&E staining, if the unstained color is white, then the white pixels in the first unstained region are counted accordingly; if the unstained color is yellow, then the yellow pixels in the first unstained region are counted accordingly.

[0124] As an example, the RGB value corresponding to white is (255, 255, 255). Pixels with RGB values ​​of (255, 255, 255) are filtered out in the first undetermined area of ​​the stained image to obtain the second number of pixels.

[0125] 504. Based on the fact that the number of first pixels is greater than the preset number of pixels and the ratio of the number of second pixels to the total number of pixels is less than the preset ratio, the first undetermined region is determined to meet the preset cell staining conditions, and the first undetermined region is taken as the initial cell region.

[0126] The determination of whether the first region to be determined meets the preset cell staining conditions is made from two dimensions: whether the number of first pixels is greater than the preset number of pixels, and whether the ratio of the number of second pixels to the total number of pixels is less than the preset ratio.

[0127] The preset number of pixels is the criterion for determining whether there are nuclear staining pixels in the first undetermined region. If the number of first pixels in the first undetermined region is greater than the preset number of pixels, it can be determined that there are nuclear staining pixels in the first undetermined region.

[0128] For example, if the preset number of pixels is 4, then it can be determined whether the number of first pixels is greater than 3. If the number of first pixels is 4, it can be determined that there are cell nucleus-stained pixels in the first undetermined area; otherwise, the blue pixels in the first undetermined area may be staining errors caused by other reasons, and there are no cell nucleus-stained pixels in the first undetermined area.

[0129] Since large areas of non-tissue regions do not exist within normal cells, the ratio of such non-tissue regions to the first undetermined region can be used to determine whether the first undetermined region is a cell.

[0130] The preset ratio is a criterion for determining whether there is a large area of ​​non-organic region in the first undetermined region. If the ratio of the number of second pixels to the total number of pixels in the first undetermined region is less than the preset ratio, it can be determined that there is no large area of ​​non-organic region in the first undetermined region; otherwise, there is a large area of ​​non-organic region, which does not belong to normal cells.

[0131] For example, if the preset ratio is 30%, then it can be determined whether the ratio of the number of second pixels to the total number of pixels is greater than 30%. If the ratio is greater than 30% (e.g., 43%), there is a large area of ​​unorganized region in the first undetermined region; otherwise, the uncolored pixels in the first undetermined region may be due to coloring errors caused by other reasons, and there is no large area of ​​unorganized region in the first undetermined region.

[0132] The determination of whether the first undetermined region meets the preset cell staining conditions is based on two dimensions. Since normal cells contain a nucleus and do not have large non-tissue areas, if both dimensions' conditions are met, the first undetermined region is considered to meet the preset cell staining conditions and can be used as the initial cell region. If either dimension's conditions are not met, the first undetermined region does not meet the preset cell staining conditions and cannot be used as the initial cell region.

[0133] If each undetermined region in the stained image satisfies the above two-dimensional criteria, then each undetermined region is taken as an initial cell region; otherwise, the undetermined region and a portion thereof are taken as initial cell regions.

[0134] Figure 6This is a schematic diagram of the undetermined region provided in the embodiments of this application. The left side is a schematic diagram of the initial cell region, in which cell 601 contains cell nucleus 6011 and does not have a large area of ​​non-tissue region. The right side is a schematic diagram of the bubble, in which unstained pixels account for most of the pixels.

[0135] In this diagram, irregularly shaped areas represent cells, but the shape of the cells is not limited to this and can be any shape. In this diagram, blue represents the color of the cell nucleus, pink represents the color of other tissues, and the pink area in this diagram represents the area to be determined. White is used to represent unstained areas, but in the actual implementation, the color of the unstained pixels is not limited to this.

[0136] In this embodiment, the number of first pixels satisfying the preset cell nucleus staining corresponding first pixel value is determined in the first undetermined region. The first undetermined region is any one of at least two undetermined regions obtained through processing by a preset neural network model. The total number of pixels corresponding to the first undetermined region is determined. The number of second pixels satisfying the second pixel value corresponding to the unstained image is determined in the first undetermined region. Unstained pixels in the stained image belong to non-tissue regions. Based on the fact that the number of first pixels is greater than the preset number of pixels, and the ratio of the number of second pixels to the total number of pixels is less than the preset ratio, the first undetermined region is determined to meet the preset cell staining conditions, and this first undetermined region is used as the initial cell region. Judging whether the undetermined region meets the preset cell staining conditions from two dimensions enables more detailed filtering and screening of the undetermined regions obtained by the preset neural network model segmentation, resulting in more accurate determination of the single cell range in the stained image.

[0137] Figure 7 This is a flowchart illustrating the process of morphologically expanding an initial cell region to obtain a target cell region, provided in an embodiment of this application. It may include steps 701 to 703, which are described in detail below.

[0138] 701. Based on the preset external physical distance and the pixel resolution of the stained image, calculate the number of iterations required for the morphological dilation operation;

[0139] In this application, a preset external expansion physical distance is defined as the distance for morphological expansion of the initial cell region.

[0140] Since the diameter of a single cell in a typical eukaryotic organism is usually between 5 and 30 micrometers, the preset external physical distance can be set to match this diameter. For example, the external physical distance can be set to 2 to 5 micrometers. Of course, the setting of the external physical distance needs to be selected according to the actual situation. For example, it can be selected according to the cell density in the tissue image. This application does not limit the specific value of the external physical distance.

[0141] Since stained images are composed of pixels, morphological dilation of the initial cellular regions in a stained image is performed on a pixel-by-pixel basis. Furthermore, morphological dilation can be performed iteratively. Each iteration performs one dilation operation until a predetermined number of iterations is reached to complete the dilation process.

[0142] Accordingly, based on the preset external physical distance and the pixel resolution of the stained image, the number of iterations required for the morphological dilation operation is determined.

[0143] In one possible implementation, the pixel resolution of the stained image can be used to determine the physical distance of each pixel in the stained image. Based on the physical distance of each pixel and the preset physical distance, the number of pixel loops for expansion can be determined. In the morphological dilation operation, the number of pixel loops can be used as the number of iterations.

[0144] The number of iterations is the ratio of the preset extended microscopic distance to the physical distance corresponding to each pixel. If the ratio is an integer, it is used as the number of iterations; if the ratio is a decimal, it is rounded to the nearest integer and used as the number of iterations.

[0145] In one possible implementation, the number of iterations N can be calculated using the following formula (1):

[0146] N= round_decimal (D / R) (1)

[0147] Where D is the preset external physical distance, R is the physical distance corresponding to a single pixel in the stained image, and round_decimal is the rounding function.

[0148] As an example, if the outward physical distance is chosen to be 2 micrometers, then the number of iterations is obtained by dividing 2 micrometers by the physical distance corresponding to each pixel.

[0149] 702. Based on the initial cell region, perform morphological dilation in the stained image with pixels as the dilation unit for an iterative number of iterations to obtain the dilated region, which surrounds the initial cell region.

[0150] The initial cell region can be the region corresponding to the cell nucleus. In subsequent steps, the pixels at the edge of the cell nucleus region are morphologically dilated according to the number of iterations.

[0151] In the process of morphological expansion, expansion is carried out in an iterative manner.

[0152] In one possible implementation, starting from the edge of the initial cell region, expansion is performed outside the initial cell region in the stained image. In each iteration of expansion, the expansion area is obtained by taking a ring of pixels outside the initial cell region.

[0153] As an example, the number of iterations is N (N is an integer greater than 1). Outside the initial cell region, a ring of pixels adjacent to the edge of that initial cell region is identified as the expansion region corresponding to the first expansion. Outside this expansion region, a ring of pixels adjacent to the edge of the first expansion region is identified as the expansion region corresponding to the second expansion, and so on. Outside the expansion region corresponding to the (N-1)th expansion, a ring of pixels adjacent to the edge of the initial cell region is identified as the expansion region corresponding to the Nth expansion. The total expansion region obtained from these N expansions is taken as the expansion region obtained from these N iterations.

[0154] The initial cell regions obtained after the aforementioned screening correspond to cells, with one initial cell region corresponding to one cell. During the expansion process based on the initial cell regions, each initial cell region can be expanded simultaneously or sequentially in a specific order. This specific order can be from top to bottom, from left to right, or any other order, which is not limited in this application.

[0155] In one possible implementation, during the morphological dilation process performed sequentially according to the number of iterations, if any pixel belongs to at least two cells, then no cell assignment is performed on that pixel.

[0156] Because two or more cells are located close to each other, there may be a certain gap between the initial regions of each cell, and they may cross each other during the expansion process. This situation can occur at any time during the expansion outside the initial cell region. The crossover occurs when any pixel belongs to two or more cells during a certain expansion.

[0157] Since cells in a tissue slice chip are distributed in a single layer and do not overlap, if two or more cell regions intersect at a certain pixel, it can be considered an error. The pixel at the intersection is not assigned to any cell, and the remaining pixels form the expansion region for this expansion.

[0158] Figure 8This is a schematic diagram of iterative dilation in a stained image provided in this application embodiment. The diagram shows the dilation regions of two iterative dilation cycles. The first iterative dilation region of the initial cell region 801 is the set of pixels surrounding the initial cell region 801, as shown in the light gray area in the diagram. The second iterative dilation region is the set of pixels surrounding the first iterative dilation region, as shown in the dark gray area in the diagram. The first and second iterative dilation regions of the initial cell region 802 are similar. One pixel 803 in the ring of pixels outside the initial cell region belongs to the intersection of the initial cell region 801 and the adjacent initial cell region 802 during the dilation process. In this schematic diagram, pixel 802 is the intersection of the two cell regions during the second dilation process. Accordingly, the iterative dilation region includes all pixels in the ring except for pixel 803. In this schematic diagram, the initial cell region is the white area outlined by dashed lines, and the pixels in the iterative dilation region are in the area outlined by thin solid lines. Pixel 803 is highlighted in yellow.

[0159] 703. Merge the expanded region with the initial cell region to obtain the target cell region.

[0160] After performing morphological dilation on the initial cell region in the stained image with a certain number of iterations, the resulting dilated region surrounds the initial target cell region. The initial target cell region is then merged with the dilated region to obtain the morphologically dilated target cell region.

[0161] Figure 9 This is a schematic diagram of the target cell region provided in an embodiment of this application. The target cell region 901 is the initial cell region 902 after two iterations of dilation. In this schematic diagram, the initial cell region is the area outlined by a dashed line, and the target cell region is the area outlined by a thin solid line. The dilated area of ​​the first iteration is the set of pixels surrounding the initial cell region 902, as shown in the light gray area in the figure. The dilated area of ​​the second iteration is the set of pixels surrounding the dilated area of ​​the first iteration, as shown in the dark gray area in the figure.

[0162] Should Figure 9 A pixel is represented by a square, but the shape of the pixel is not limited to this in the actual implementation.

[0163] Should Figure 9 In the diagram, during each expansion process, the edge of a square (one pixel) is used as the reference, and each pixel is expanded outward by one pixel in sequence, which is considered as one expansion circle.

[0164] Should Figure 9 The initial cell region and target cell region are used only to illustrate the relative relationship between the initial cell region and the target cell region; their shapes are not limited to these. Figure 9 The rectangle shown can be of any shape.

[0165] In this embodiment, the number of iterations required for the morphological dilation operation is calculated based on the preset external physical distance and the pixel resolution of the stained image. Using the initial cell region as a reference, morphological dilation is performed in the stained image with the number of iterations as the dilation unit to obtain the dilated region, which surrounds the initial cell region. The dilated region is then merged with the initial cell region to obtain the target cell region. This achieves morphological dilation of the initial region obtained through processing and screening by the preset neural network model, resulting in a target cell region containing complete cells. This improves the accuracy of the corresponding cell region and provides a more accurate basis for subsequent determination of cell gene expression.

[0166] Figure 10 This is a flowchart illustrating how the set of analysis units associated with each target cell region is determined based on the spatial coordinates of each analysis unit in the tissue slice chip, according to an embodiment of this application. It may include steps 1001 to 1002, which are described in detail below.

[0167] 1001. Determine the polygonal spatial coordinate range of the target cell region based on the target cell region in the stained image, and determine the polygonal spatial coordinate range of each analysis unit based on the spatial coordinates of each analysis unit in the tissue slice chip.

[0168] After identifying each target cell region in the stained image, the polygonal spatial coordinate range corresponding to each target cell region can be determined based on the coordinate system set in the stained image.

[0169] In the stained image, the coordinate range corresponding to each target cell region is determined. The coordinate range of the target cell region is the range corresponding to each cell in the stained image, which can be defined by the outermost pixels of the target cell region forming the boundary of the target cell region. The boundary of the target cell region can be regarded as a polygon boundary.

[0170] Furthermore, the polygonal spatial coordinate range of each analysis unit can be determined based on the spatial coordinates of each analysis unit in the tissue slice chip.

[0171] In one possible implementation, after quantitative analysis of the FASTQ file using spatial transcriptomics analysis software (such as Spaceranger), the results include the spatial coordinates (e.g., x and y coordinates of a two-dimensional square centroid) of each analysis unit. Correspondingly, the coordinates of the four vertices of the analysis unit can be calculated based on its side length, and the polygonal boundary of the analysis unit can be constructed based on these coordinates. The polygonal boundary of each analysis unit defines its polygonal spatial coordinate range.

[0172] In one possible implementation, during the analysis of RNA from tissue slices on an MCU, the tissue slice can be divided into several regions corresponding to analysis units in a specific order (e.g., from left to right and from top to bottom), and the corresponding gene expression can be measured for each region. Each analysis unit has spatial coordinates and gene expression data.

[0173] In one possible implementation, since the analysis software requires inputting stained image files and tissue capture region image files, and calibrates the stained images and capture region images, the spatial coordinate systems of the stained images and tissue slice chips are corresponding.

[0174] 1002. Perform spatial intersection calculation on the polygonal spatial coordinate range of the target cell region and the polygonal spatial coordinate range of the analysis unit to obtain at least one target analysis unit that intersects with the target cell region. At least one target analysis unit constitutes a set of analysis units associated with the target cell region.

[0175] The stained image is obtained by taking a microscopic image of the tissue slice chip through a microscope. Accordingly, the distribution of cells in the stained image is completely consistent with the distribution of cells in the tissue slice chip.

[0176] In this application, by taking advantage of the fact that the cell distribution in the stained image and the tissue slice chip is completely consistent, the analysis unit corresponding to each cell in the stained image can be obtained by spatially intersecting the stained image and the tissue slice chip.

[0177] In one possible implementation, taking advantage of the fact that the cell distribution in the stained image and the tissue slice chip is completely consistent, the analysis unit corresponding to the polygonal spatial coordinate range of each target cell region is obtained by calculating the spatial intersection of the polygonal spatial coordinate range of the target cell region and the polygonal spatial coordinate range of each analysis unit. The polygonal spatial coordinate range of the analysis unit belongs to the polygonal spatial coordinate range of the target cell region.

[0178] The term "belongs to" can refer to either partial or complete belonging. Accordingly, if the entire polygonal spatial coordinate range of the analysis unit belongs to the target cell region, the expression data of all genes in the analysis unit can be used as the gene expression of the corresponding cell in the target cell region; if only a portion of the polygonal spatial coordinate range of the analysis unit belongs to the target cell region, the expression data of only a portion of the genes in the analysis unit can be used as the gene expression of the corresponding cell in the target cell region.

[0179] In one possible implementation, the process of determining the set of analysis units corresponding to each target cell region can be implemented using software.

[0180] For example, the geopandas package (a Python library for processing spatial data) can be used to calculate the polygonal spatial coordinate range of the target cell region and the polygonal spatial coordinate range of the analysis unit. The sjoin function in the geopandas package can calculate whether there is an intersection of spatial coordinates between two polygonal boundaries. Analysis units with spatial connections will be assigned to the corresponding cells.

[0181] If the polygonal spatial coordinate range corresponding to an analysis unit intersects with the polygonal spatial coordinate range of a unique cell, it means that the analysis unit belongs exclusively to that cell; if the polygonal spatial coordinate range corresponding to an analysis unit intersects with the polygonal spatial coordinate ranges of multiple cells, it means that the analysis unit belongs to those multiple cells.

[0182] In this embodiment, the polygonal spatial coordinate range corresponding to each target cell region is determined in the stained image, and the polygonal spatial coordinate range of each analysis unit is determined in the tissue slice chip. By performing spatial intersection calculation on the polygonal spatial coordinate range of the target cell region on the stained image and the tissue slice chip with the polygonal spatial coordinate range of the analysis unit, at least one target analysis unit that intersects with the target cell region is obtained. This target analysis unit constitutes the set of analysis units associated with the target cell region. This realizes the process of obtaining the set of analysis units corresponding to each target cell region by intersecting the stained image and the tissue slice chip. Taking advantage of the fact that the cell distribution in the stained image and the tissue slice chip is completely consistent, the set of analysis units corresponding to each target cell region can be determined by comparing the stained image and the tissue slice chip, and the data processing burden of the device is small.

[0183] Figure 11 This is a flowchart illustrating the process of integrating gene expression data from each analysis unit in each set of analysis units to obtain the gene expression profile of each target cell region on a tissue slice chip, as provided in the embodiments of this application. It may include steps 1101 to 1102, which are described in detail below.

[0184] 1101. Based on genes, add the gene expression data corresponding to each analysis unit in the set of analysis units associated with each target cell region to obtain the gene expression data of each target cell region.

[0185] Each analysis unit in an analysis unit set corresponds to the same target cell region, and each target cell region corresponds to one cell. Correspondingly, the gene expression data of each analysis unit in the same analysis unit set belongs to the gene expression data of the same cell. By adding the gene expression data of each analysis unit belonging to the same cell, the gene expression data of that cell is obtained, thus realizing the determination of single-cell gene expression data.

[0186] After determining the set of analytical units associated with each target cell region, the gene expression data of each analytical unit in the set of analytical units are added together to obtain the gene expression data of the corresponding target cell region. The gene expression data of the target cell region is the gene expression data of the corresponding single cell.

[0187] This analysis unit corresponds to a 2-millimeter square area. For tissue slice chips, it produces a gene expression matrix with the analysis unit as the resolution. Each analysis unit contains expression data of multiple genes, and each cell on the tissue slice chip can express multiple genes.

[0188] Since each analysis unit contains the expression data of each gene, when summing the gene expression of each analysis unit in the target cell region, the gene expression can be summed separately to obtain the expression data of each gene in each cell.

[0189] As an example, a cell corresponds to a target cell region in a stained image. This target cell region contains 100 analysis units. The expression data of each gene in these 100 analysis units are summed separately, gene by gene, to obtain the expression data of each gene. For example, if the cell contains 5 genes, then when summing the expression data of the 100 analysis units corresponding to the target cell region, the expression data of the 5 genes in each of the 100 analysis units are summed separately to obtain the expression data of the 5 genes corresponding to the cell.

[0190] The gene expression matrix of this tissue slice chip can be recorded in the form of Table 1 below.

[0191] Table 1

[0192]

[0193] In Table 1, the expression data in each table item is the expression data of each gene in the corresponding analysis unit.

[0194] In one possible implementation, the expression data in each table entry may include information about whether the gene exists, and if the corresponding gene data is included in any analysis unit, the table entry may include the gene data obtained from sequencing the gene; if the corresponding gene data is not included in any analysis unit, the expression data in the table entry may be used only to indicate the absence of the gene.

[0195] For example, a value of 2 in a table entry indicates that the corresponding analysis unit contains gene data of the corresponding gene, while a value of 0 in a table entry indicates that the corresponding analysis unit does not contain gene data of the corresponding gene.

[0196] As an example, analysis unit 1 contains data for genes A, B, and D. In the table, expression data a is the expression data of gene A in analysis unit, expression data b is the expression data of gene B in analysis unit, expression data d is the expression data of gene D in analysis unit; expression data c indicates that there is no expression data for gene C in analysis unit 1.

[0197] Table 1 is only used to illustrate the form of the gene expression matrix. The tissue slice chip contains several analysis units. Moreover, the number of genes contained in a cell is not limited to the four types in Table 1 (A, B, C, D). The number of genes contained in a cell can be determined according to the actual situation. Other values ​​can also be used to represent the gene data of an analysis unit that contains a certain gene or does not contain a certain gene.

[0198] If a target cell region is determined to contain analysis unit 1 and analysis unit 2, the gene expression data of the target cell region is obtained by adding the gene data of analysis unit 1 and analysis unit 2 based on genes. The gene expression data of the target cell region includes: gene A (expression data a + expression data e), gene B (expression data b + expression data f), gene C (expression data c + expression data g), and gene D (expression data d + expression data h).

[0199] In one possible implementation, the gene expression data contained in each analysis unit is the gene data of each gene in the cell to which it belongs. Each analysis unit may contain the gene expression of all genes in the cell to which it belongs, or it may contain the gene expression of some genes in the cell to which it belongs.

[0200] If an analytical unit belongs to only one set of analytical units, that is, if the analytical unit belongs to a unique cell, then the gene expression of the analytical unit belongs entirely to that unique cell.

[0201] If an analytical unit belongs to two or more sets of analytical units, meaning it belongs to two or more cells, then the gene expression of that analytical unit is allocated to those two or more cells. This allocation can be weighted according to the overlap area between the analytical unit and each cell, and the proportion of the analytical unit's area.

[0202] 1102. Gene expression data of each target cell region are used to form the gene expression profile of the target cell region on the tissue slice chip.

[0203] Through the aforementioned step 1101, gene expression data of each target cell region in the stained image are obtained, and the gene expression data of each target cell region in the stained image are collected to form the gene expression profile of each cell on the tissue slice chip.

[0204] This gene expression profile records gene expression data for each cell at the single-cell level.

[0205] It should be noted that, due to various factors, the gene expression data of each cell in this gene expression map may not be completely identical, or they may be completely identical. This application does not impose any restrictions on the relationship between the gene expression data of each cell in the gene expression map.

[0206] In this embodiment, based on genes, the gene expression data corresponding to each analysis unit in the set of analysis units associated with each target cell region are added together to obtain the gene expression data of each target cell region. The gene expression data of each target cell region are then used to form the gene expression profile of the target cell region on the tissue slice chip. By adding the gene expression data corresponding to each analysis unit belonging to the same target cell region, the gene expression data of the cell corresponding to that target cell region can be obtained. Thus, the gene expression data of each cell on the tissue slice chip are used to form the gene expression profile of each cell on the tissue slice chip, making the data processing process simple.

[0207] Figure 12 This is a flowchart illustrating how, based on genes, the gene expression data of each analysis unit in the set of analysis units associated with each target cell region are added together to obtain the gene expression data of each target cell region. It may include steps 1201 to 1204, which are described in detail below.

[0208] 1201. When any analysis unit intersects with the polygonal spatial coordinate range of at least two target cell regions, calculate the overlap area between the analysis unit and each intersecting target cell region.

[0209] By intersecting the polygonal spatial coordinate range of each analysis unit on the tissue slice chip with the polygonal spatial coordinate range of the target cell region in the stained image, we can obtain the analysis units corresponding to the polygonal spatial coordinate range of each target cell region.

[0210] Since the polygonal spatial coordinate range of an analysis unit corresponds to multiple pixels in the stained image, it is possible that an analysis unit intersects with the polygonal spatial coordinate range of only one target cell region. In this case, the entire analysis unit belongs to the polygonal spatial coordinate range of one target cell region, and the gene expression data in the analysis unit belongs to a unique cell. Alternatively, it is possible that an analysis unit intersects with the polygonal spatial coordinate ranges of multiple target cell regions. In this case, the gene expression data in the analysis unit belongs to the cells corresponding to the multiple target cell regions.

[0211] The size of the overlap area between the analysis unit and the intersecting target cell region indicates how close the analysis unit is to the target cell region. The greater the closeness, the more likely the gene expression data in the analysis unit belongs to the target cell region. Conversely, the smaller the closeness, the less likely the gene expression data in the analysis unit belongs to the target cell region.

[0212] Therefore, the overlap area between the analysis unit and each intersecting target cell region can be determined sequentially. This overlap area can be determined in terms of physical size or in terms of pixels.

[0213] As an example, the set of pixels in the stained image corresponding to the polygonal spatial coordinates of the analysis unit can be determined. Then, the number of pixels belonging to each intersecting target cell region within this set can be determined. Finally, the overlapping area can be calculated using this number and the physical size of the pixels. Of course, this is just one example of determining the overlapping area; other methods can be used in specific implementations.

[0214] 1202. Determine the weight of each intersecting target cell region based on the ratio of the overlapping area to the total area of ​​the analysis unit;

[0215] The ratio of the overlapping area of ​​the analysis unit and each intersecting target cell region to the total area of ​​the analysis unit is calculated to obtain the ratio of the overlapping area to the total area of ​​the analysis unit.

[0216] For example, if the overlap area between the analysis unit and cell A accounts for 10% of the area of ​​the analysis unit, the overlap area between the analysis unit and cell B accounts for 60% of the area of ​​the analysis unit, and the remaining 30% of the area of ​​the analysis unit is empty and does not overlap with any cell, then for this analysis unit, the weight of cell A is 10% and the weight of cell B is 60%.

[0217] For example, if the overlap area between the analysis unit and cell A accounts for 8% of the area of ​​the analysis unit, the overlap area between the analysis unit and cell B accounts for 72% of the area of ​​the analysis unit, and the overlap area between the analysis unit and cell C accounts for 20% of the area of ​​the analysis unit, then for this analysis unit, the weight of cell A is 8%, the weight of cell B is 72%, and the weight of cell C is 20%.

[0218] 1203. Allocate the gene expression data of the analysis unit to each intersecting target cell region according to the weight of each intersecting target cell region;

[0219] According to the weight of each intersecting target cell region determined in the aforementioned steps, the gene expression data in the analysis unit is allocated to achieve the determination of the gene expression data corresponding to the multiple intersecting target cell regions of the analysis unit.

[0220] In one possible implementation, each target cell region contains fully overlapping analysis units and partially overlapping analysis units. The fully overlapping analysis units have all gene expression data belonging to the target cell region, and the partially overlapping analysis units have some gene expression data belonging to the target cell region, allocated according to a determined weight.

[0221] For example, cell A has a weight of 10%, cell B has a weight of 60%, and 10% of the gene expression data in the analysis unit is allocated to cell A, while 60% of the gene expression data in the analysis unit is allocated to cell B.

[0222] 1204. Based on genes, add the gene expression data corresponding to each analysis unit in the set of analysis units associated with each target cell region and the gene expression data assigned to the target cell region by the analysis unit to obtain the gene expression data of the corresponding target cell region.

[0223] Based on genes, the gene expression data of each analysis unit associated with the target cell region are added together to obtain the gene expression data of the target cell region.

[0224] Specifically, if the polygonal spatial coordinate range of the associated analysis unit completely overlaps with the polygonal spatial coordinate range of the target cell region, then all gene expression data of that analysis unit are considered as part of the gene expression data of the target cell region. If the polygonal spatial coordinate range of the associated analysis unit partially overlaps with the polygonal spatial coordinate range of the target cell region, then the portion of gene expression data allocated to the target cell region by that analysis unit is considered as part of the gene expression data of the target cell region. The gene expression data of each analysis unit associated with each target cell region are weighted and summed to obtain the gene expression data of that target cell region.

[0225] As an example, if a target cell region is determined to contain analysis unit 1 and analysis unit 2, wherein the coordinate range of analysis unit 1 completely overlaps with that of the target cell region, and the weight of the target cell region relative to analysis unit 2 is 0.5, referring to the records in Table 1 above, the gene expression data of the target cell region includes: gene A (expression data a + 0.5 × expression data e), gene B (expression data b + 0.5 × expression data f), gene C (expression data c + 0.5 × expression data g), and gene D (expression data d + 0.5 × expression data h).

[0226] In this embodiment, when any analysis unit intersects with the polygonal spatial coordinate range of at least two target cell regions, the weight assigned to the gene expression data of the analysis unit for each intersecting target cell region is determined by the overlap area between the analysis unit and the intersecting target cell regions and the total area of ​​the analysis unit. The gene expression data of the corresponding analysis unit is then allocated to the corresponding target cell region according to this weight. When determining the gene expression data of each target cell region, the gene expression data in the corresponding analysis unit and the gene expression data allocated to the target cell region are added together. The gene expression data in the analysis units corresponding to different target cell regions is allocated according to the overlap area ratio, thereby achieving refined allocation of gene expression data and improving the accuracy of cell gene expression data.

[0227] Figure 13 This is another flowchart illustrating the method for determining cell gene expression data provided in the embodiments of this application, which may include steps 1301 to 1302. The steps in this flowchart may be... Figure 1 After step 103 shown, steps 104-105 are executed in parallel. These steps are described in detail below.

[0228] 1301. Based on a preset number of preset evaluation indicators, the segmentation quality of the target cell region determined in the stained image is evaluated to obtain the evaluation results;

[0229] After segmenting each cell in the stained image into target cell regions, the gene expression data of each cell determined by these target cell regions can be used as input data for data processing of the tissue slice microarray in certain scenarios. By evaluating the quality of the target cell regions obtained from cell segmentation in the stained image, the accuracy of this data processing can be improved.

[0230] Preset evaluation indicators can be set in the electronic device that performs the method for determining single-cell gene expression data. These preset evaluation indicators can be set as needed.

[0231] In this stained image, one target cell region corresponds to one cell. After screening by a preset neural network model and preset cell staining conditions, the quality of each target cell region in the stained image is evaluated by preset evaluation indicators, and an evaluation result can be obtained. This evaluation result can represent the region segmentation effect of the corresponding cell in the stained image.

[0232] In one possible implementation, the evaluation metrics include at least two of the following: the number of cells per unit area in the stained image, the coverage of the target cell region in the stained image, the coverage of cell pixels in the target cell region, the similarity of pixels outside the target cell region within the tissue region in the stained image, and the standard deviation of each target cell region.

[0233] The number of cells per unit area can be used to standardize cell density measurements between different image masks, providing a unified benchmark for comparison.

[0234] The number of cells per unit area can be calculated using the following formula:

[0235] (2)

[0236] Where, N cell This represents the number of cells per unit area; s0 represents the unit area; n c This indicates the number of cells identified in the stained image (number of target cell regions), where c represents a cell; s p p represents the area of ​​one pixel; n represents the area of ​​one pixel. p This represents the total number of pixels in the tissue region on the tissue slice chip, where p stands for pixel.

[0237] For example, when s0 is 100 square micrometers, formula (2) determines the number of cells within 100 square micrometers.

[0238] The coverage of the target cell region in the stained image represents the proportion of effective tissue region pixels on the tissue slice chip. The higher the coverage value, the higher the utilization rate of the tissue slice chip.

[0239] The coverage of the target cell region in the stained image can be calculated using the following formula:

[0240] (3)

[0241] Where FIT represents the coverage of the target cell region in the stained image; N ptThis represents the number of effective tissue region pixels on the tissue slice chip in the stained image. In this pt, p represents pixel and t represents tissue region; N pi This represents the total number of pixels on the tissue slice chip in the stained image, where p stands for pixel and i stands for image.

[0242] The effective tissue area is the area covered by tissue, including tissue areas where cells are present and the aforementioned non-tissue areas such as bubbles or cavities.

[0243] The coverage rate of cell pixels in the target cell region represents the proportion of cell pixels in the effective tissue region on the tissue slice chip. The larger the coverage rate, the more pixels are identified as cells.

[0244] The coverage of cell pixels in the target cell region can be calculated using the following formula:

[0245] (4)

[0246] Where FTC represents the coverage of cell pixels in the target cell region; N pc This represents the total number of pixels in the target cell region corresponding to each cell in the stained image, where p in pc stands for pixel and c stands for cell; N pt This represents the total number of pixels in the effective tissue region on the tissue slice chip in the stained image. In this pt, p represents pixel and t represents tissue region.

[0247] The similarity between pixels outside the target cell region within the tissue region of the stained image is determined by the coefficient of variation (CV) of the extracellular pixels in the tissue region of the stained image. The coefficient of variation is the ratio of the standard deviation to the mean. The coefficient of variation is used to evaluate the uniformity of the image. A lower coefficient of variation indicates that the pixel signal values ​​are more uniformly distributed, that is, the composition of pixels outside the predicted cell region within the tissue region is similar.

[0248] The coefficient of variation of extracellular pixels in the tissue region of this stained image can be calculated using the following formula:

[0249] (5)

[0250] Where ACVT represents the coefficient of variation of extracellular pixels in the tissue region of the stained image; n ch σ represents the number of image channels, where ch represents the image channel. jtμ represents the standard deviation of pixel intensity in the extracellular region of the tissue region in the j-th channel, where t represents the tissue region. jt represents the average pixel intensity of the extracellular region of the tissue region in the j-th channel, where t represents the tissue region.

[0251] The similarity between pixels outside the target cell region within a stained image is evaluated using 1 / (CV+1). A small coefficient of variation indicates a uniform distribution of extracellular pixels, meaning that pixels not classified within the target cell region correspond to the extracellular matrix, indicating good segmentation of the target cell region in the stained image. The upper limit of this similarity value is 1.

[0252] The standard deviation of each target cell region is used to measure the degree of variation in cell size, since the target cell region is the region corresponding to the cell in the stained image.

[0253] (6)

[0254] Where CSSD represents the standard deviation of the target cell region, s i μ represents the size of the i-th cell (the previously defined target cell region); s This represents the average size of all target cell regions, where 's' represents the target cell region; n c This indicates the number of cells identified in the stained image, where 'c' stands for cell.

[0255] Generally speaking, the cell size of the same tissue slice should be uniform. Therefore, 1 / (ln(CSSD)+1) can be used as an evaluation metric. The smaller the CSSD value, the more uniform the cell size. The higher the corresponding 1 / (ln(CSSD)+1) value, the better the effect of identifying the target cell region for each cell.

[0256] Currently, other indicators can also be used as evaluation indicators, such as the total number of genes expressed in all cells, the average number of genes expressed in each cell, and the average UMI expressed in each cell. This application does not limit the specific content of the evaluation indicator, and other indicators may be used in addition to the five mentioned above.

[0257] In one possible implementation, the five indicators mentioned above can be combined to determine whether the target cell region corresponding to the divided cells is accurate, and then the gene expression data of each cell determined by the target cell region can be combined to determine whether the calculation basis conditions are met.

[0258] 1302. If the evaluation results meet the calculation criteria, the cell gene expression profile of the tissue slice chip will be used as the input data for downstream data processing.

[0259] Based on the input data requirements of different downstream data processing, set the corresponding calculation basis conditions.

[0260] In one possible implementation, data from multiple samples are compared horizontally. For example, multiple samples from the same batch are compared to determine the effectiveness of cell segmentation in the staining images of this batch of samples. This effectiveness can be used to evaluate the consistency and comparability of this batch of samples.

[0261] Tissue slides from similar tissues are treated as samples from the same batch. The corresponding stained images from this batch are processed, resulting in each stained image being divided into multiple target cell regions. The division results for each stained image are evaluated, and the evaluation results are compared uniformly to determine if they meet the calculation criteria. Since the evaluation indicators for cells in similar tissues should be the same or similar, the evaluation indicators for the corresponding tissue slides are also the same or similar.

[0262] In one possible implementation, prior knowledge can be combined to set different computational basis conditions for different cell types and different data processing needs.

[0263] For example, if the size of type A cells is much larger than that of type B cells, then the number of type A cells per unit area will be much smaller than the number of type B cells.

[0264] In this embodiment, the segmentation quality of the target cell regions identified in the stained image is evaluated based on a preset number of preset evaluation indicators to obtain evaluation results. If the evaluation results meet the calculation criteria, the cell gene expression profile of the tissue slice chip is used as input data for downstream data processing. By evaluating the quality of the target cell regions corresponding to each cell obtained from the segmentation of the stained image, the level of segmentation quality can be determined, which helps to assess the accuracy and reliability of the cell-corresponding region division. Furthermore, the gene expression data of each cell determined using highly reliable segmentation results also has higher accuracy.

[0265] Figure 14 This is a flowchart illustrating the application scenario of the method for determining cell gene expression data provided in this application, including:

[0266] 1401. Input the FASTQ file;

[0267] In this scenario, HD's product stains the tissue slice chip sample and captures a microscopic image to obtain a H&E image. Then, it sequences the sample to obtain a FASTQ file of the sequence. Here, the input is the FASTQ file of the sample.

[0268] 1402. Process FASTQ files;

[0269] In this scenario, SpaceRanger (version 3.0.0) software was used to perform quantitative expression analysis on the FASTQ file of the input sample, resulting in a gene expression data matrix and spatial coordinates containing square bins with a resolution of 2 micrometers.

[0270] Steps 1401-1402 are executed in parallel with steps 1403-1406.

[0271] 1403. Input H&E image;

[0272] The input stained image is a H&E image obtained by microscopic imaging of tissue section chips.

[0273] 1404. Cell segmentation yields multiple regions to be determined;

[0274] A pre-defined neural network model is used to process and segment the H&E image, obtaining a cell label mask. This label mask consists of multiple instance regions, each corresponding to a region to be determined. Each region to be determined represents a cell, thus achieving initial segmentation of the H&E image. However, due to various influencing factors, some instance regions (regions to be determined) in the label mask obtained by the pre-defined neural network model may not be actual cell regions but rather misidentified. Therefore, a subsequent filtering process is required.

[0275] 1405. Filter the target area;

[0276] Preset cell staining conditions are used to filter the target area to remove misjudgments caused by non-tissue areas such as air bubbles and pores.

[0277] 1406. Expand the filtered region of interest;

[0278] The target cell region is obtained by expanding the region to be determined, and the target cell region corresponds to the instance region of the complete cell.

[0279] 1407. Obtain the cell gene expression matrix;

[0280] By intersecting the spatial coordinates of the analysis unit and the target cell region, and using the gene expression data corresponding to each analysis unit, the gene expression data of each analysis unit corresponding to the same cell are added together to obtain a single-cell level cell gene expression data matrix.

[0281] 1408. Evaluate the quality of segmentation.

[0282] The segmentation results of the H&E image are evaluated using a preset evaluation index by using the label mask of multiple complete cells obtained from the segmentation of the H&E image, and the evaluation index value corresponding to each evaluation index is obtained.

[0283] This application also provides analysis results of a practical application scenario using a method for determining cell gene expression data.

[0284] In one application scenario, it involves determining single-cell gene expression data from fresh human tonsil tissue cell samples.

[0285] You can use publicly available HD test data of fresh human tonsil tissue from certain official websites / platforms, download the TIF image file of this sample, and the output results directory obtained after analysis software, which mainly contains gene expression data of 2-micron square bins.

[0286] The sample's TIF image file serves as the input staining file, and the outs result directory serves as the gene expression set corresponding to the input tissue slice chip.

[0287] Figure 15 This is a partial schematic diagram of fresh human tonsil tissue cell segmentation provided in an embodiment of this application. The diagram is obtained by analyzing a stained image of a fresh human tonsil tissue section. In this diagram, the area within a red circle represents one cell.

[0288] Using the steps in the method for determining cellular gene expression data in the aforementioned embodiments, a gene expression matrix at the single-cell level is obtained. The segmented cell regions are obtained, and a local view reference is obtained from the H&E image segmentation of a fresh human tonsil tissue cell sample. Figure 15 The quality of cell segmentation was evaluated using the five evaluation metrics recorded in the aforementioned examples.

[0289] Figure 16 The diagram shows the statistical values ​​of five evaluation indicators for cell segmentation results of fresh human tonsil tissue cell samples.

[0290] In this application scenario, the five evaluation metrics include: cell count per 100 square micrometers, tissue coverage on the chip, cell pixel coverage within the tissue region, coefficient of variation of extracellular pixels in the tissue image, and standard deviation of cell size, with values ​​of 1.5107602743921082, 0.9905654438749714, 0.8528573636899542, 0.8650876040615006, and 0.1480680152318473. The tissue coverage on the chip refers to the region corresponding to each cell, and cell size is determined using the target cell region corresponding to each cell. This sample yielded 11,210,791 2-micrometer analysis units, with an average of 31 genes and 34 UMIs detected per analysis unit. A total of 677,520 cells were obtained. On average, 625.6834027039793 genes and 503.3300878636719 UMIs were detected per cell, improving data resolution while ensuring data quality.

[0291] In another application scenario, it is used to determine single-cell gene expression data from mouse kidney FFPE tissue samples.

[0292] You can use publicly available HD test data of mouse kidney FFPE tissue from certain official websites / platforms, download the TIF image file of this sample, and the output results directory obtained after analysis software, which mainly contains gene expression data of 2-micron square bins.

[0293] The sample's TIF image file serves as the input staining file, and the outs results directory serves as the gene expression data set corresponding to the input tissue slice chip.

[0294] Figure 17 This is a partial schematic diagram of mouse kidney FFPE tissue cell segmentation provided in the embodiments of this application. The schematic diagram is obtained by analyzing the stained image of mouse kidney FFPE tissue section sample. In the schematic diagram, the area within a red circle belongs to one cell.

[0295] Using the steps in the method for determining cellular gene expression data in the aforementioned embodiments, a gene expression data matrix at the single-cell level is obtained. The segmented cell regions are then obtained, and a local view reference is obtained from the H&E image segmentation of mouse kidney FFPE tissue samples. Figure 17 The quality of cell segmentation was evaluated using the five evaluation metrics recorded in the aforementioned examples.

[0296] Figure 18The diagram shows the statistical values ​​of five evaluation indicators for cell segmentation results of mouse kidney FFPE tissue samples. In this application scenario, these five evaluation indicators include: cell number per 100 square micrometers, tissue coverage on the chip, cell pixel coverage in the tissue region, coefficient of variation of extracellular pixels in the tissue image, and standard deviation of cell size. Their values ​​are: 0.9778193465673501, 0.7885825438539193, 0.7022752945956752, 0.880871905890176, and 0.14383937563001883. The tissue coverage on the chip refers to the region corresponding to each cell, and the cell size is determined using the target cell region corresponding to each cell. This sample yielded 7,998,466 2-micrometer analysis units, with an average of 27 genes and 29 UMIs detected per analysis unit. A total of 312,983 cells were obtained. On average, 578.9120016103111 genes and 578.5492485172548 UMIs were detected per cell, which improved data resolution while ensuring data quality.

[0297] The above describes a method for determining cell gene expression data provided by embodiments of this application. The following will describe an electronic device for performing the above method for determining cell gene expression data.

[0298] Please see Figure 19 , Figure 19 This is a schematic diagram of the structure of an electronic device using a method for determining cell gene expression data, as provided in an embodiment of this application. Figure 19 As shown, the electronic device 1900 includes:

[0299] Interface 1901 is used to obtain stained images of tissue slice chips, as well as spatial coordinates and gene expression data of each analysis unit in the tissue slice chip. Each analysis unit contains expression data of multiple genes. Each cell on the tissue slice chip can express multiple genes, and the tissue slice chip contains several cells.

[0300] The processor 1902 is used to process a stained image using a preset neural network model, identify at least two undetermined regions in the stained image, and the at least two undetermined regions can form at least a part of the stained image; select regions that meet preset cell staining conditions from the at least two undetermined regions to obtain target cell regions; determine a set of analysis units associated with each target cell region based on the spatial coordinates of each analysis unit in the tissue slice chip, the set of analysis units containing at least one analysis unit; and integrate the gene expression data of each analysis unit in each set of analysis units to obtain the gene expression profile of each target cell region on the tissue slice chip.

[0301] It should be noted that the specific steps and explanations for the functions performed by each component in this electronic device can be found in the explanations in the foregoing method embodiments, and will not be repeated here.

[0302] In this embodiment, a stained image of a tissue section chip is obtained, along with the spatial coordinates and gene expression data of each analysis unit within the tissue section chip. The processor processes the stained image of the tissue section chip using a preset neural network model, identifying multiple undetermined regions within the stained image. These undetermined regions constitute all or part of the stained image. Using preset cell staining conditions, these undetermined regions are filtered to obtain target cell regions. The spatial coordinates and gene expression data of each analysis unit within the tissue section chip are obtained; each analysis unit contains the expression data of multiple genes, and each cell on the tissue section chip expresses multiple genes. Based on each target cell region and the spatial coordinates of each analysis unit in the tissue section chip, a set of analysis units associated with each target cell region is determined, and each set of analysis units contains at least one analysis unit. The gene expression data of each analysis unit in each set of analysis units is integrated to obtain the gene expression profile of each target cell region on the tissue section chip. After initial segmentation of the stained image using a preset neural network model to obtain multiple undetermined regions, the use of preset cell staining conditions to filter these undetermined regions reduces the misjudgment of the preset neural network model and improves the accuracy of cell segmentation. Therefore, the accuracy of the determined gene expression profile of cells in the tissue section chip is higher.

[0303] This application also provides an electronic device in its embodiments. (See reference...) Figure 20 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 20 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0304] like Figure 20As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 2001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2002 or a program loaded from a storage device 2008 into a random access memory (RAM) 2003. When the electronic device is powered on, the RAM 2003 also stores various programs and data required for the operation of the electronic device. The processing unit 2001, ROM 2002, and RAM 2003 are interconnected via a bus 2004. An input / output (I / O) interface 2005 is also connected to the bus 2004.

[0305] Typically, the following devices can be connected to the I / O interface 2005: input devices 2006 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 2007 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 2008 including, for example, memory cards, hard drives, etc.; and communication devices 2009. The communication device 2009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although... Figure 20 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0306] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the cell gene expression data determination methods provided in this application.

[0307] In one possible implementation, the computer program product can be software developed using the Python programming language and the conda integrated software environment.

[0308] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the methods for determining cell gene expression data provided in this application.

[0309] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0310] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0311] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0312] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0313] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A method of determining cell gene expression data, characterized by, The method comprises the following steps: obtain a staining image of a tissue section chip, the tissue section chip comprising a plurality of cells; determine at least two uncertain regions in the staining image by processing the staining image through a preset neural network model, the at least two uncertain regions being capable of constituting at least part of the staining image; screen regions satisfying a preset cell staining condition in the at least two uncertain regions to obtain target cell regions; obtain spatial coordinates and gene expression data of each analysis unit in the tissue section chip, each analysis unit comprising expression data of a plurality of genes, each cell on the tissue section chip expressing the plurality of genes; determine a set of analysis units associated with each target cell region according to each target cell region and the spatial coordinates of each analysis unit in the tissue section chip, the set of analysis units comprising at least one analysis unit; integrate gene expression data of each analysis unit in each set of analysis units to obtain a gene expression profile of each target cell region on the tissue section chip.

2. The method of determining cellular gene expression data according to claim 1, wherein, The step of determining at least two uncertain regions in the staining image by processing the staining image through a preset neural network model comprises the following steps: processing the staining image through a preset neural network model to obtain a horizontal gradient field map, a vertical gradient field map and a binary image of the staining image, the value in the binary image representing whether any pixel belongs to any cell or not; merging the horizontal gradient field map, the vertical gradient field map and the binary image of the staining image to obtain a gradient vector field map corresponding to the staining image; determining a target position to which each pixel in the staining image converges according to the gradient vector field map, the target position corresponding to at least one pixel; allocating each pixel converging to the same target position to the same instance region to form the uncertain region, each instance region constituting a label mask of the staining image.

3. The method for determining cell gene expression data according to claim 1, characterized in that, The step of screening regions satisfying a preset cell staining condition in the at least two uncertain regions to obtain target cell regions comprises the following steps: screening regions satisfying a preset cell staining condition in the at least two uncertain regions to obtain initial cell regions; performing morphological dilation on the initial cell regions to obtain target cell regions.

4. The method for determining cell gene expression data according to claim 1, characterized in that, The step of screening regions satisfying a preset cell staining condition in the at least two uncertain regions to obtain initial cell regions comprises the following steps: determining a first number of pixels satisfying a first pixel value in a first uncertain region, the first pixel value being a pixel value corresponding to a preset cell nucleus staining, the first uncertain region being any one of the at least two uncertain regions; determining a total number of pixels corresponding to the first uncertain region; determining a second number of pixels satisfying a second pixel value in the first uncertain region, the second pixel value being a pixel value corresponding to an unstained image, the pixels not stained in the staining image belonging to non-tissue region pixels; determining that the first uncertain region satisfies a preset cell staining condition according to that the first number of pixels is greater than a preset number of pixels and that a ratio of the second number of pixels to the total number of pixels is less than a preset ratio, the first uncertain region being an initial cell region.

5. The method for determining cell gene expression data according to claim 3, characterized in that, The morphological dilation on the initial cell region is performed to obtain a target cell region, including: Based on a preset physical dilation distance and a pixel resolution of the staining image, an iteration number required by the morphological dilation operation is calculated; Taking the initial cell region as a reference, morphological dilation is performed on the staining image by pixels as a dilation unit for the iteration number to obtain a dilated region, and the dilated region surrounds the initial cell region; The dilated region and the initial cell region are merged to obtain the target cell region.

6. The method for determining cell gene expression data according to claim 1, characterized in that, The analysis unit set associated with each target cell region is determined according to the spatial coordinates of each target cell region and each analysis unit in the tissue section chip, including: The polygonal spatial coordinate range of each target cell region is determined according to each target cell region in the staining image, and the polygonal spatial coordinate range of each analysis unit is determined according to the spatial coordinates of each analysis unit in the tissue section chip; The polygonal spatial coordinate range of the target cell region and the polygonal spatial coordinate range of the analysis unit are subjected to spatial intersection calculation to obtain at least one target analysis unit intersecting with the target cell region, and the at least one target analysis unit constitutes the analysis unit set associated with the target cell region.

7. The method of determining cellular gene expression data according to claim 6, wherein, The gene expression data of each analysis unit in each analysis unit set is integrated to obtain the gene expression profile of each target cell region on the tissue section chip, including: The gene expression data of each analysis unit in the analysis unit set associated with each target cell region is added to obtain the gene expression data of each target cell region; The gene expression data of each target cell region constitutes the gene expression profile of the target cell region on the tissue section chip.

8. The method for determining cell gene expression data according to claim 7, characterized in that, The gene expression data of each analysis unit in the analysis unit set associated with each target cell region is added to obtain the gene expression data of each target cell region, including: In the case that any analysis unit intersects with the polygonal spatial coordinate range of at least two target cell regions, the overlapping area of the analysis unit and each intersecting target cell region is calculated; According to the proportion of the overlapping area to the total area of the analysis unit, the weight of each intersecting target cell region is determined; According to the weight of each intersecting target cell region, the gene expression data of the analysis unit is distributed to each intersecting target cell region; The gene expression data of each analysis unit in the analysis unit set associated with each target cell region is added to obtain the gene expression data of each target cell region.

9. The method for determining cell gene expression data according to claim 1, characterized in that, After the gene expression profile of each single cell is obtained, further including: According to a preset evaluation index of a preset number, the segmentation quality of the target cell region determined in the staining image is evaluated to obtain an evaluation result; If the evaluation result meets the calculation condition, the cell gene expression profile of the tissue section chip is taken as input data for downstream data processing.

10. An electronic device, comprising: Including: An interface is configured to obtain a staining image of a tissue section chip and obtain spatial coordinates and gene expression data of each analysis unit in the tissue section chip, each analysis unit containing expression data of a plurality of genes, each cell on the tissue section chip being capable of expressing the plurality of genes, and the tissue section chip containing a plurality of cells; a processor is configured to process the staining image by using a preset neural network model, determine at least two to-be-determined regions in the staining image, and the at least two to-be-determined regions being capable of constituting at least part of the staining image; screen the at least two to-be-determined regions to obtain a target cell region satisfying a preset cell staining condition; determine, according to each target cell region and the spatial coordinates of each analysis unit in the tissue section chip, an analysis unit set associated with each target cell region, the analysis unit set containing at least one analysis unit; and integrate the gene expression data of each analysis unit in each analysis unit set to obtain a gene expression profile of each target cell region on the tissue section chip.