Gene image data correction method and system, electronic equipment and storage medium

CN120322796APending Publication Date: 2025-07-15SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102396.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-05
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the existing technology, the correction algorithm of RNA cell labeling relies on cell outline images and is weak in robustness, resulting in large errors in gene image data. Many gene molecules are not classified and the accuracy is not high.

Method used

By determining the area occupied by the gene molecules of the target cell in the gene image, the spatial coordinates and gene expression characteristic signal values ​​of the unknown background gene molecules are obtained, and the probability distribution model is input to calculate the probability score of the cell to which the background gene molecules belong, and correction is performed based on the probability score. , reducing data errors caused by gene molecule diffusion.

Benefits of technology

It improves the accuracy and robustness of genetic image data correction, reduces errors caused by gene molecule diffusion, enhances correction efficiency, and does not rely on high-precision cell segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120322796A_ABST
    Figure CN120322796A_ABST
Patent Text Reader

Abstract

The invention provides a gene image data correction method and system, electronic equipment and a storage medium. The gene image data correction method comprises the following steps: inputting a space coordinate and a gene expression characteristic signal value into a probability distribution model to obtain a probability score of a target cell to which a background gene molecule belongs; and determining cells to which the background gene molecules belong according to the probability score. On the premise of roughly knowing the target area of the target cell, the probability score of the background gene molecule belonging to the unknown belonging cell in the preset correction range belongs to the target cell is calculated, and the belonging cell of the background gene molecule is determined according to the probability score. The situation of large gene image data error caused by gene molecule diffusion is reduced, and the correction efficiency is improved; and meanwhile, the method does not depend on a high-precision cell segmentation result, so that the robustness of a correction scheme is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Gene image data correction method, system, electronic device and storage medium Technical Field

[0001] The present application relates to the technical field of gene image processing, and in particular to a gene image data correction method, system, electronic device and storage medium. Background Art

[0002] RNA cell labeling (a technology for labeling genetic information within living cells) assigns genes to cells based on aligned cell outline images and gene expression data corresponding to the cells. However, due to issues such as the small number of genes within cells, nuclear segmentation, and the diffusion of RNA (ribonucleic acid, a genetic information carrier present in biological cells and some viruses and viroids), the number of genes assigned to cells is far smaller than the number of genes actually contained in the cells. Therefore, the results of RNA cell labeling need to be corrected to improve the accuracy of RNA cell labeling. RNA cell labeling is a very important step in the downstream analysis of spatial transcriptome sequencing technology such as Stereo-seq (Spatio-Temporal Enhanced REsolution Omics-sequencing) and is the basis for other downstream analyses. The current correction algorithm for RNA cell labeling relies on the cell morphology of the cell outline image. The algorithm is less robust, the final result is not very accurate, and a large number of RNA gene molecules are not classified. Therefore, there is still much room for exploration and optimization to address this issue.

[0003] Summary of the Invention

[0004] The main purpose of this application is to provide a gene image data correction method, system, electronic device and storage medium to improve the defects in the existing technology that gene image data has large errors due to the diffusion of gene molecules and a large number of gene molecules are not classified.

[0005] This application solves the above technical problems through the following technical solutions:

[0006] In a first aspect, the present application provides a method for correcting gene image data, the method comprising:

[0007] Determine the target area occupied by the gene molecules belonging to the target cells in the gene image;

[0008] Determine the spatial coordinates and gene expression characteristic signal values ​​of background gene molecules of unknown cells within a preset calibration range of the target area;

[0009] Inputting the spatial coordinates and the gene expression characteristic signal values ​​into a probability distribution model to obtain a probability score of the background gene molecule belonging to the target cell; the probability distribution model is fitted based on the spatial coordinates and gene expression characteristic signal values ​​of the gene molecule sample belonging to the target cell;

[0010] The cell to which the background gene molecule belongs is determined according to the probability score.

[0011] Preferably, the step of determining the cell to which the background gene molecule belongs according to the probability score comprises:

[0012] The probability score is compared with a probability threshold; if the probability score is greater than or equal to the probability threshold, the background gene molecules are corrected to the gene molecules belonging to the target cell.

[0013] Preferably, the step of determining the spatial coordinates and gene expression characteristic signal values ​​of background gene molecules of unknown cells within a preset calibration range of the target area includes:

[0014] Obtaining coordinate information of the gene molecule belonging to the target cell;

[0015] Determining the center point of the target area according to the coordinate information;

[0016] The preset correction range of the target area is determined with the center point as the center.

[0017] Preferably, the step of obtaining the coordinate information of the gene molecule belonging to the target cell includes:

[0018] Obtain microscopic and genetic images of biological samples;

[0019] Perform image registration on microscope images and gene images of biological samples;

[0020] performing cell segmentation on the microscope image to obtain a cell segmentation result;

[0021] Determining a gene image of a target cell based on the cell segmentation result and the image registration result of the gene image;

[0022] The coordinate information of the gene molecules belonging to the target cells is determined according to the gene image of the target cells.

[0023] Preferably, the step of determining the center point of the target area according to the coordinate information includes:

[0024] Counting the number of molecules of the gene molecules belonging to the target cells;

[0025] Calculating the sum of the horizontal coordinate values ​​and the sum of the vertical coordinate values ​​of the gene molecules belonging to the target cell;

[0026] Determine the value of the abscissa of the center point of the target area as a first ratio; the first ratio is the ratio of the sum of the abscissa values ​​to the number of molecules;

[0027] The value of the vertical coordinate of the center point of the target area is determined to be a second ratio; the second ratio is the ratio of the sum of the vertical coordinate values ​​to the number of molecules.

[0028] Preferably, the step of comparing the probability score with the probability threshold comprises:

[0029] determining a third quartile between the first probability score and the second probability score as a probability threshold;

[0030] The first probability score is the probability score with the highest score value among the probability scores; and the second probability score is the probability score with the lowest score value among the probability scores.

[0031] Preferably, the gene image data correction method further comprises:

[0032] If, according to the comparison result, the background gene molecules belong to different target cells within a preset distance range, then the probability scores corresponding to the background gene molecules in the different target cells are obtained;

[0033] The probability scores are compared, and the background gene molecules are corrected to the gene molecules belonging to the target cell with the highest probability score.

[0034] Preferably, before the step of determining the preset correction range of the target area with the center point as the center, the step includes:

[0035] Constructing a two-dimensional point cloud map of the target area according to the center point;

[0036] The preset correction range is determined according to the two-dimensional point cloud image.

[0037] Preferably, the step of constructing a two-dimensional point cloud map of the target area according to the center point includes:

[0038] constructing a digital image based on the microscope image within the target area, wherein the grayscale values ​​of the pixels of the digital image are the same, namely, the first grayscale value;

[0039] Adjusting a center point in the digital image to a second grayscale value, wherein the first grayscale value is different from the second grayscale value;

[0040] The adjusted digital image is determined as a two-dimensional point cloud image of the target area.

[0041] Preferably, the step of determining the cell to which the background gene molecule belongs according to the probability score comprises:

[0042] Generate target cell gene expression information based on the corrected results of the gene molecules belonging to the target cell.

[0043] In a second aspect, the present application provides a gene image data correction system, the gene image data correction system comprising:

[0044] A determination module is used to determine the target area occupied by the gene molecules belonging to the target cells in the gene image; and is also used to determine the spatial coordinates and gene expression characteristic signal values ​​of the background gene molecules of the unknown cells within a preset correction range of the target area;

[0045] an acquisition module, configured to input the spatial coordinates and the gene expression characteristic signal values ​​into a probability distribution model to obtain a probability score of the background gene molecule belonging to the target cell; the probability distribution model is fitted based on the spatial coordinates and gene expression characteristic signal values ​​of the gene molecule sample belonging to the target cell;

[0046] A correction module is used to determine the cell to which the background gene molecule belongs according to the probability score.

[0047] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned gene image data correction method when executing the computer program.

[0048] In a fourth aspect, the present application provides a computer-readable medium having computer instructions stored thereon, which, when executed by a processor, implement the gene image data correction method as described above.

[0049] The positive progress of this application is:

[0050] On the premise of roughly understanding the target area of ​​the target cell, this application calculates the probability score of the background gene molecules of the unknown belonging cells within the preset correction range belonging to the target cell, and determines the cell to which the background gene molecules belong based on the probability score, thereby reducing the situation where the gene image data has large errors due to the diffusion of gene molecules and improving the correction efficiency; at the same time, it does not rely on high-precision cell segmentation results, thereby enhancing the robustness of the correction scheme. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] FIG1 is a first flow chart of the gene image data correction method according to Example 1 of the present application;

[0052] FIG2 is a second flow chart of the gene image data correction method of Example 1 of the present application;

[0053] FIG3 is a third flow chart of the gene image data correction method of Example 1 of the present application;

[0054] FIG4 is a fourth flow chart of the gene image data correction method of Example 1 of the present application;

[0055] FIG5 is a fifth flow chart of the gene image data correction method of Example 1 of the present application;

[0056] FIG6 is a schematic diagram of a two-dimensional point cloud image of a target region of a target cell in the gene image data correction method according to Example 1 of the present application;

[0057] FIG7 is a sixth flow chart of the gene image data correction method of Example 1 of the present application;

[0058] FIG8 is a diagram showing the morphological changes of cells after correction using the gene image data correction method of Example 1 of the present application;

[0059] FIG9 a is a diagram showing the first spatial clustering effect of the mouse brain before correction of the gene image data correction method of Example 1 of the present application;

[0060] FIG9 b is a diagram showing the second spatial clustering effect of the mouse brain before correction of the gene image data correction method of Example 1 of the present application;

[0061] FIG9 c is a diagram showing the first spatial clustering effect of the mouse brain after correction using the gene image data correction method of Example 1 of the present application;

[0062] FIG9 d is a diagram showing the second spatial clustering effect of the mouse brain after correction using the gene image data correction method of Example 1 of the present application;

[0063] FIG10 shows the change in the average number of genes in rat brain cells before and after correction using the gene image data correction method according to Example 1 of the present application;

[0064] FIG11a is a diagram showing the effect of target cells before correction using the UMAP algorithm (a dimensionality reduction algorithm) in the gene image data correction method of Example 1 of the present application;

[0065] FIG11b is a diagram showing the effect of target cells before correction using the Leiden algorithm (a clustering algorithm) in the gene image data correction method of Example 1 of the present application;

[0066] FIG11c is a diagram showing the target cell effect after correction using the UMAP algorithm according to the gene image data correction method of Example 1 of the present application;

[0067] FIG11d is a diagram showing the effect of target cells after correction using the Leiden algorithm in the gene image data correction method of Example 1 of the present application;

[0068] FIG12 is a first structural diagram of a gene image data correction system according to Example 2 of the present application;

[0069] FIG13 is a second structural diagram of the gene image data correction system according to Example 2 of the present application;

[0070] FIG14 is a schematic structural diagram of an electronic device according to Example 3 of the present application. DETAILED DESCRIPTION

[0071] The present application is further described below by way of examples, but the present application is not limited to the scope of the examples.

[0072] Example 1

[0073] During high-throughput sequencing analysis, Stereo-seq spatial transcriptome data are aligned with ssDNA (single-stranded DNA) images, and cell segmentation is performed on the ssDNA images to establish a correspondence between pixel coordinates and cells. This correspondence is then used to convert the Stereo-seq spatial transcriptome data into cellular gene expression data for downstream analysis.

[0074] However, because the cell size obtained based on ssDNA image segmentation is only the size of the cell nucleus, there is a discrepancy between the actual cell size and the actual cell size. Furthermore, due to issues such as insufficient gene capture within the cell during sequencing experiments, nuclear segmentation, and RNA diffusion, the number of genes assigned to the cell is far smaller than the number of genes contained in the actual cell. This results in low data accuracy throughout the process, errors in the resulting single-cell gene expression data, and consequently distortion in downstream single-cell analysis. Sequencing distortion occurs when acquiring spatial transcriptome data.

[0075] The spatiotemporal data correction algorithm based on the Gaussian mixture model further corrects the results of RNA cell labeling. Specifically, the spatiotemporal data correction algorithm uses a Gaussian mixture model (GMM) to calculate the probability score of RNA molecules belonging to neighboring cells. The parameters of the GMM are fitted according to the RNA distribution in the current cell (mainly relying on the spatial coordinates of ssDNA (single-stranded DNA, single-stranded deoxyribonucleic acid) and the UMI (unique molecular identifiers, unique molecular identifier) ​​value. The probability values ​​of the RNA gene molecules adjacent to the current cell boundary are calculated using the fitted GMM parameters. However, the spatiotemporal data correction algorithm depends on the cell morphology of the cell contour image, the algorithm is less robust, the final effect is not accurate, and a large number of RNA gene molecules are not classified. In order to overcome the above-mentioned defects currently existing, this embodiment provides a gene image data correction method, see Figure 1, the gene image data correction method includes:

[0076] S1. Determine the target area occupied by the gene molecules belonging to the target cells in the gene image.

[0077] S2. Determine the spatial coordinates and gene expression characteristic signal values ​​of the background gene molecules of the unknown cells within the preset calibration range of the target region.

[0078] In this embodiment, the gene expression signature is DNB. By mapping each DNB in ​​the Stereo-seq spatial transcriptome data, the cell information to which the DNB belongs can be obtained. The remaining DNBs outside the cell are considered background DNBs. The spatial coordinates of the background gene molecules can be represented by the coordinates of the background DNBs.

[0079] S3. Input the spatial coordinates and gene expression characteristic signal values ​​into the probability distribution model to obtain the probability score of the background gene molecules belonging to the target cell.

[0080] When fitting a GMM (Gaussian mixture model), there are three independent variables: spatial coordinates (x, y) and the signal value. The dependent variable, or output, is the probability score for each gene molecule belonging to the target cell. The specific implementation process involves introducing the Gaussian mixture function from Scikit-learn (a machine learning library), automatically fitting the GMM model parameters to the input data, and obtaining the final probabilistic model.

[0081] The probability distribution model is obtained by fitting the spatial coordinates of the gene molecule samples and the gene expression characteristic signal value samples belonging to the target cells. As a preferred embodiment, in this embodiment, the probability distribution model can be a Gaussian mixture model.

[0082] S4. Determine the cell to which the background gene molecule belongs based on the probability score.

[0083] In an optional embodiment, referring to FIG2 , step S4 includes:

[0084] S41. Compare whether the probability score is greater than or equal to the probability threshold.

[0085] If the probability score is greater than or equal to the probability threshold, step S42 is executed. If the probability score is less than the probability threshold, the background gene molecules are still background molecules of unknown cells.

[0086] In an optional embodiment, the third quartile between the first probability score and the second probability score is determined as the probability threshold.

[0087] The first probability score is the probability score with the highest score value among the probability scores; the second probability score is the probability score with the lowest score value among the probability scores.

[0088] S42. Correct the background gene molecules to the gene molecules belonging to the target cells.

[0089] In this embodiment, on the premise of roughly understanding the target area of ​​the target cell, the probability score of the background gene molecules of the unknown belonging cells within the preset correction range belonging to the target cell is calculated, and the cell to which the background gene molecules belong is determined based on the probability score, thereby reducing the situation where the gene image data has large errors due to the diffusion of gene molecules and improving the correction efficiency; at the same time, it does not rely on high-precision cell segmentation results, thereby enhancing the robustness of the correction scheme.

[0090] In an optional embodiment, referring to FIG2 , the process before step S2 includes:

[0091] S21. Obtain coordinate information of the gene molecule belonging to the target cell.

[0092] In an optional embodiment, referring to FIG3 , step S21 includes:

[0093] S211. Obtain microscope images and gene images of biological samples.

[0094] Among them, the microscope image and gene image contain several target cells.

[0095] A microscope image can be an image obtained by photographing a biological sample using a microscope and can include the spatial information of the biological sample. A gene image can include both the spatial information and the genetic data of the biological sample. Pixels in a gene image correspond to elements of the gene matrix, and the pixel values ​​of the pixels correspond to the values ​​of the elements. Therefore, a gene image can include the spatial information of the same biological sample.

[0096] In this embodiment, the microscope image can be represented as an ssDNA image, but the type of microscope image is not limited and can be adjusted and selected according to actual needs. The gene image can be represented as stereo-seq spatial transcriptome data, spatial transcriptome data, or gene expression data, etc.

[0097] S212: Perform image registration on the microscope image and the gene image of the biological sample.

[0098] For example, ssDNA images (i.e., microscope images) are image-aligned with Stereo-seq spatial transcriptome data (i.e., gene images) to generate cellular gene expression information.

[0099] During image registration, the microscope image is rotated and / or scaled based on its markers. A preliminary offset is calculated using the center of gravity of the microscope image and the genetic image to perform a preliminary offset. The offset of the microscope image is corrected using the markers of the two images, and image registration is performed using the offset-corrected microscope image.

[0100] S213. Perform cell segmentation on the microscope image to obtain a cell segmentation result.

[0101] That is, cell segmentation is performed on the microscope image to obtain the correspondence between the spatial position and the cells, and the cell segmentation result is generated according to the correspondence between the spatial position and the cells.

[0102] S214 : Determine the gene image of the target cell based on the cell segmentation result and the image registration result of the gene image.

[0103] Based on the gene image of the target cell, the gene molecules can be divided into the gene molecules belonging to the target cell and the background gene molecules of the unknown cell.

[0104] S215. Determine the coordinate information of the gene molecules belonging to the target cell according to the gene image of the target cell.

[0105] In this embodiment, by determining the gene image of the target cell, the coordinate information of the gene molecule belonging to the target cell is determined, thereby improving the accuracy of determining the coordinate information of the gene molecule belonging to the target cell.

[0106] S22. Determine the center point of the target area according to the coordinate information.

[0107] In an optional embodiment, referring to FIG4 , step S22 includes:

[0108] S221. Count the number of gene molecules belonging to the target cell.

[0109] S222. Calculate the sum of the horizontal coordinate values ​​and the sum of the vertical coordinate values ​​of the gene molecules belonging to the target cells.

[0110] S223: Determine the value of the horizontal coordinate of the center point of the target area as a first ratio.

[0111] The first ratio is the ratio of the sum of the horizontal coordinate values ​​to the number of molecules.

[0112] S224. Determine the value of the vertical coordinate of the center point of the target area as the second ratio.

[0113] The second ratio is the ratio of the sum of the vertical coordinate values ​​to the number of molecules.

[0114] In this embodiment, the center point of the target area of ​​the target cell is scientifically determined through the above calculation, thereby improving the correction accuracy.

[0115] S23. Determine a preset correction range of the target area with the center point as the center.

[0116] In an optional embodiment, referring to FIG5 , before step S23, the following steps are included:

[0117] S231. Construct a two-dimensional point cloud map of the target area based on the center point.

[0118] In an optional embodiment, step S231 includes:

[0119] S2311. Construct a digital image based on the microscope image within the target area.

[0120] The grayscale values ​​of the pixels of the digital image are the same, which are all first grayscale values;

[0121] S2312: Adjust the center point of the digital image to a second grayscale value.

[0122] The first grayscale value is different from the second grayscale value to distinguish the center point of the cell from other pixel points of the digital image.

[0123] S2313. The adjusted digital image is determined as a two-dimensional point cloud image of the target area. A two-dimensional point cloud image is an image with a black background (grayscale value 0) and a foreground of discrete pixels (grayscale value is any value greater than 0 and less than or equal to 255). Each pixel represents the center point of a cell. (The cell center point is determined based on the horizontal and vertical coordinates of the gene molecules within the cell.)

[0124] The following is an example of constructing a two-dimensional point cloud map:

[0125] First, a completely black digital image of the same size as the original microscope image is generated (ie, the grayscale values ​​of the image are all 0, and the first grayscale value is 0).

[0126] Next, the center points of all cells are calculated in step S22.

[0127] Finally, in the previously generated all-black digital image, the center points of all the calculated cells are lit up in turn (that is, the grayscale value is set to an arbitrary number in the range of (0, 255), which is the second grayscale value). The final image obtained is the two-dimensional point cloud image of this embodiment.

[0128] FIG6 is a schematic diagram of a two-dimensional point cloud image of a target area of ​​a target cell.

[0129] S232. Determine a preset correction range based on the two-dimensional point cloud image.

[0130] In this embodiment, the correction step can be performed only by obtaining the approximate position coordinates of the target area of ​​the target cell, thereby further enhancing the robustness of this embodiment and avoiding the dependence of this embodiment on high-precision cell segmentation results.

[0131] In this embodiment, each target cell is abstracted as a coordinate point, and with the coordinate point of the target cell as the center, the spatial coordinates and gene expression characteristic signal values ​​of the background gene molecules of the unknown cells within the preset correction range are obtained. This reduces the large errors in gene image data caused by the diffusion of gene molecules without relying on high-precision cell segmentation results, further improves the correction efficiency, and at the same time more scientifically determines the preset correction range, thereby further improving the accuracy of the correction.

[0132] In an optional embodiment, referring to FIG7 , the gene image data correction method further includes:

[0133] S6. If, according to the comparison result, the background gene molecules belong to different target cells within the preset distance range, the probability scores corresponding to the background gene molecules in the different target cells are obtained.

[0134] S7. Compare the probability scores and correct the background gene molecules to the gene molecules belonging to the target cell with the highest probability score.

[0135] In practice, using the gene image data correction method of this embodiment, background gene molecules from an unknown cell may be simultaneously attributed to multiple neighboring cells. In this embodiment, the probability scores of these background gene molecules belonging to several neighboring cells are further calculated. These background gene molecules are then corrected to the gene molecules belonging to the target cell with the highest probability score, further improving the accuracy of the correction results.

[0136] In an optional embodiment, referring to FIG7 , step S5 includes:

[0137] S8. Generate target cell gene expression information based on the corrected results of the gene molecules belonging to the target cells.

[0138] The following sets of experimental effect diagrams illustrate the beneficial effects of the gene image data correction method of this embodiment.

[0139] FIG8 is a diagram showing the morphological changes of cells after correction by the gene image data correction method of this embodiment, where the black color represents the newly added gene molecules that determine the cell to which they belong.

[0140] Figures 9a and 9b show the spatial clustering of the mouse brain before correction; Figures 9c and 9d show the clustering after correction. Comparison shows that the spatial locations of the corrected mouse brain cell clusters are clearer and that there are fewer noise points within the clusters. The x-axis (horizontal axis) and y-axis (vertical axis) of Figures 9a, 9b, 9c, and 9d are all in pixels.

[0141] Figure 10 shows the change in the average gene number per cell in mouse brain cells before and after correction. The solid line represents the average gene number per cell after correction; the dashed line represents the average gene number per cell before correction. The horizontal axis represents the gene num per cell, and the vertical axis represents the ratio of the average gene number per cell. Figure 10 clearly shows that the average gene number per cell in mouse brain cells increased significantly after correction.

[0142] Figure 11a is the effect diagram before correction by the UMAP algorithm (a dimensionality reduction algorithm), and Figure 11b is the effect diagram before correction by the Leiden algorithm (a clustering algorithm); Figure 11c is the effect diagram after correction by the UMAP algorithm (a dimensionality reduction algorithm), and Figure 11d is the effect diagram after correction by the Leiden algorithm (a clustering algorithm).

[0143] Figures 11a and 11c are visualizations of the dimensionality reduction results of the original data. Different numbers represent different categories after dimensionality reduction. The more dispersed the points of different categories are, and the more concentrated the points of the same category are, the better the clustering effect is, which proves the effectiveness of the correction algorithm. Compared to Figure 11a, Figure 11c shows that the clustering effect after correction is even better.

[0144] Figures 11b and 11d show the results of clustering using the Leiden algorithm. In the corrected data, if the distribution of genes within the same cluster exhibits a more distinct spatial distribution pattern, the correction algorithm is considered effective. Compared to Figure 11b, Figure 11d shows that the corrected clustering results exhibit a more distinct spatial distribution pattern (the black dots on the tissue in the figure).

[0145] Example 2

[0146] This embodiment provides a gene image data correction system. Referring to FIG12 , the gene image data correction system includes:

[0147] Determination module 1 is used to determine the target area occupied by the gene molecules belonging to the target cells in the gene image; it is also used to determine the spatial coordinates and gene expression characteristic signal values ​​of the background gene molecules of the unknown cells within the preset correction range of the target area.

[0148] Acquisition module 2 is used to input the spatial coordinates and gene expression characteristic signal values ​​into the probability distribution model to obtain the probability score of the background gene molecule belonging to the target cell; the probability distribution model is fitted based on the spatial coordinates and gene expression characteristic signal values ​​of the gene molecule sample belonging to the target cell.

[0149] Correction module 3 is used to determine the cell to which the background gene molecules belong based on the probability score.

[0150] In an optional embodiment, the correction module 3 is further configured to compare the probability score with a probability threshold; if the probability score is greater than or equal to the probability threshold, the background gene molecules are corrected to the gene molecules belonging to the target cell.

[0151] In an optional embodiment, the acquisition module 2 is further configured to acquire coordinate information of gene molecules belonging to the target cell.

[0152] The determination module 1 is further used to determine the center point of the target area according to the coordinate information; and is further used to determine a preset correction range of the target area with the center point as the center.

[0153] In an optional embodiment, the acquisition module 2 is further configured to acquire microscope images and gene images of biological samples.

[0154] Referring to FIG13 , the gene image data correction system further includes:

[0155] The registration module 4 is used to perform image registration on the microscope image and the gene image of the biological sample.

[0156] The acquisition module 2 is further used to perform cell segmentation on the microscope image to obtain a cell segmentation result.

[0157] The determination module 1 is further used to determine the gene image of the target cell based on the cell segmentation result and the image registration result of the gene image; and is further used to determine the coordinate information of the gene molecule belonging to the target cell according to the gene image of the target cell.

[0158] In an optional embodiment, referring to FIG13 , the gene image data correction system further includes:

[0159] The statistical module 5 is used to count the number of gene molecules belonging to the target cells.

[0160] The calculation module 6 is used to calculate the sum of the horizontal coordinate values ​​and the sum of the vertical coordinate values ​​of the gene molecules belonging to the target cells.

[0161] Determination module 1 is also used to determine that the value of the horizontal coordinate of the center point of the target area is a first ratio; the first ratio is the ratio of the sum of the horizontal coordinate values ​​to the number of molecules; and is also used to determine that the value of the vertical coordinate of the center point of the target area is a second ratio; the second ratio is the ratio of the sum of the vertical coordinate values ​​to the number of molecules.

[0162] In an optional embodiment, the determination module 1 is further configured to determine the third quartile between the first probability score and the second probability score as the probability threshold; wherein the first probability score is the probability score with the highest score value among the probability scores; and the second probability score is the probability score with the lowest score value among the probability scores.

[0163] In an optional embodiment, the acquisition module 2 is further configured to obtain probability scores corresponding to the background gene molecules in different target cells if, according to the comparison result, the background gene molecules belong to different target cells within a preset distance range.

[0164] The correction module 3 is further used to compare the probability scores and correct the background gene molecules to the gene molecules belonging to the target cell with the highest probability score.

[0165] In an optional embodiment, referring to FIG13 , the gene image data correction system further includes:

[0166] The construction module 7 is used to construct a two-dimensional point cloud map of the target area based on the center point.

[0167] The determination module 1 is further configured to determine a preset correction range based on the two-dimensional point cloud image.

[0168] In an optional embodiment, the construction module 7 is further configured to construct a digital image based on the microscope image within the target area, wherein the grayscale values ​​of the pixels in the digital image are the same, namely, the first grayscale value;

[0169] The correction module 3 is further configured to adjust a center point in the digital image to a second grayscale value, wherein the first grayscale value is different from the second grayscale value;

[0170] The determination module 1 is further configured to determine the adjusted digital image as a two-dimensional point cloud image of the target area.

[0171] In an optional embodiment, referring to FIG13 , the gene image data correction system further includes:

[0172] The generating module 8 is used to generate target cell gene expression information according to the corrected results of the gene molecules belonging to the target cells.

[0173] It should be noted that the implementation principles and technical effects of each module in this embodiment can refer to the corresponding parts of Example 1 and will not be repeated here.

[0174] Example 3

[0175] This embodiment provides an electronic device. Figure 14 is a schematic diagram of the modules of the electronic device. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable by the processor. When the processor executes the computer program, it implements the genetic image data correction method of Example 1. The electronic device 30 shown in Figure 14 is merely an example and should not limit the functionality or scope of use of the embodiments of the present invention.

[0176] As shown in FIG14 , the electronic device 30 may be a general-purpose computing device, such as a server device. Components of the electronic device 30 may include, but are not limited to, the at least one processor 31, the at least one memory 32, and a bus 33 connecting various system components (including the memory 32 and the processor 31).

[0177] The bus 33 includes a data bus, an address bus, and a control bus.

[0178] The memory 32 may include a volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322 , and may further include a read-only memory (ROM) 323 .

[0179] The memory 32 may also include a program / utility 325 having a set (at least one) of program modules 324, such program modules 324 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0180] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32 , such as the gene image data correction method of embodiment 1 of the present invention.

[0181] The electronic device 30 can also communicate with one or more external devices 34 (e.g., a keyboard, pointing device, etc.). This communication can occur via an input / output (I / O) interface 35. Furthermore, the model-generating device 30 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 36. As shown in FIG14 , the network adapter 36 communicates with other modules of the model-generating device 30 via a bus 33. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the model-generating device 30, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.

[0182] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more units / modules described above may be embodied in a single unit / module. Conversely, the features and functions of a single unit / module described above may be further divided and embodied by multiple units / modules.

[0183] Example 4

[0184] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the gene image data correction method of embodiment 1 is implemented.

[0185] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0186] In a possible implementation manner, the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the gene image data correction method of embodiment 1.

[0187] The program code for executing the present invention may be written in any combination of one or more programming languages, and may be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on the remote device.

[0188] Although the above describes specific embodiments of the present invention, it should be understood by those skilled in the art that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention. Therefore, the scope of protection of the present invention is defined by the appended claims.

Claims

1. A gene image data correction method, characterized in that: The gene image data correction method comprises: Determine the target area occupied by the gene molecules belonging to the target cells in the gene image; Determine the spatial coordinates and gene expression characteristic signal values ​​of background gene molecules of unknown cells within a preset calibration range of the target area; Inputting the spatial coordinates and the gene expression characteristic signal values ​​into a probability distribution model to obtain a probability score of the background gene molecule belonging to the target cell; the probability distribution model is fitted based on the spatial coordinates and gene expression characteristic signal values ​​of the gene molecule sample belonging to the target cell; The cell to which the background gene molecule belongs is determined according to the probability score.

2. The gene image data correction method according to claim 1, wherein: The step of determining the cell to which the background gene molecule belongs according to the probability score comprises: The probability score is compared with a probability threshold; if the probability score is greater than or equal to the probability threshold, the background gene molecules are corrected to the gene molecules belonging to the target cell.

3. The gene image data correction method according to claim 2, wherein: The step of determining the spatial coordinates and gene expression characteristic signal values ​​of background gene molecules of unknown cells within the preset correction range of the target area includes: Obtaining coordinate information of the gene molecule belonging to the target cell; Determining the center point of the target area according to the coordinate information; The preset correction range of the target area is determined with the center point as the center.

4. The gene image data correction method according to claim 3, wherein: The step of obtaining the coordinate information of the gene molecule belonging to the target cell includes: Obtain microscopic and genetic images of biological samples; Perform image registration on microscope images and gene images of biological samples; performing cell segmentation on the microscope image to obtain a cell segmentation result; Determining a gene image of a target cell based on the cell segmentation result and the image registration result of the gene image; The coordinate information of the gene molecules belonging to the target cells is determined according to the gene image of the target cells.

5. The gene image data correction method according to claim 3, wherein: The step of determining the center point of the target area according to the coordinate information includes: Counting the number of molecules of the gene molecules belonging to the target cells; Calculating the sum of the horizontal coordinate values ​​and the sum of the vertical coordinate values ​​of the gene molecules belonging to the target cell; Determine the value of the abscissa of the center point of the target area as a first ratio; the first ratio is the ratio of the sum of the abscissa values ​​to the number of molecules; The value of the vertical coordinate of the center point of the target area is determined to be a second ratio; the second ratio is the ratio of the sum of the vertical coordinate values ​​to the number of molecules.

6. The gene image data correction method according to claim 2, wherein: The step of comparing the probability score with the probability threshold comprises: determining a third quartile between the first probability score and the second probability score as a probability threshold; The first probability score is the probability score with the highest score value among the probability scores; and the second probability score is the probability score with the lowest score value among the probability scores.

7. The gene image data correction method according to claim 1, wherein: The gene image data correction method further comprises: If, according to the comparison result, the background gene molecules belong to different target cells within a preset distance range, then the probability scores corresponding to the background gene molecules in the different target cells are obtained; The probability scores are compared, and the background gene molecules are corrected to the gene molecules belonging to the target cell with the highest probability score.

8. The gene image data correction method according to claim 3, wherein: The step of determining the preset correction range of the target area with the center point as the center includes: Constructing a two-dimensional point cloud map of the target area according to the center point; The preset correction range is determined according to the two-dimensional point cloud image.

9. The gene image data correction method according to claim 8, wherein: The step of constructing a two-dimensional point cloud map of the target area according to the center point includes: constructing a digital image based on the microscope image within the target area, wherein the grayscale values ​​of the pixels of the digital image are the same, namely, the first grayscale value; Adjusting a center point in the digital image to a second grayscale value, wherein the first grayscale value is different from the second grayscale value; The adjusted digital image is determined as a two-dimensional point cloud image of the target area.

10. The gene image data correction method according to claim 1, wherein: The step of determining the cell to which the background gene molecule belongs according to the probability score comprises: Generate target cell gene expression information based on the corrected results of the gene molecules belonging to the target cell.

11. A gene image data correction system, characterized in that: The gene image data correction system includes: A determination module is used to determine the target area occupied by the gene molecules belonging to the target cells in the gene image; and is also used to determine the spatial coordinates and gene expression characteristic signal values ​​of the background gene molecules of the unknown cells within a preset correction range of the target area; an acquisition module, configured to input the spatial coordinates and the gene expression characteristic signal values ​​into a probability distribution model to obtain a probability score of the background gene molecule belonging to the target cell; the probability distribution model is fitted based on the spatial coordinates and gene expression characteristic signal values ​​of the gene molecule sample belonging to the target cell; A correction module is used to determine the cell to which the background gene molecule belongs according to the probability score.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the gene image data correction method according to any one of claims 1 to 10 is implemented.

13. A computer-readable medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the method for correcting genetic image data according to any one of claims 1 to 10 is implemented.