Data correction method and system based on Gaussian smoothing, related equipment and medium

CN120359497APending Publication Date: 2025-07-22SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102680.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The barcode-based multiple sequence alignment method in Stereo-seq spatial transcriptome sequencing technology has problems of high noise and high sparsity, which affects the accuracy of expression information and subsequent analysis.

Method used

The data correction method based on Gaussian smoothing model is used to obtain the gene expression data and position data of the sample, determine the nearest neighbor samples and calculate the weight, and then correct the gene expression data, reduce noise interference and enhance features.

Benefits of technology

Effectively denoises and enhances expression matrix data features, provides more accurate cell expression information, improves the correction effect of spatial biological data, and provides reliable data for downstream analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359497A_ABST
    Figure CN120359497A_ABST
Patent Text Reader

Abstract

The invention discloses a data correction method and system based on a Gaussian smoothing model, related computer equipment and a storage medium. The data correction method comprises the following steps: acquiring gene expression data and position data of a plurality of samples; processing neighbor samples of the plurality of samples, determining neighbor samples of individual samples according to the gene expression data of the plurality of samples, and calculating weights of the neighbor samples corresponding to the individual samples according to the position data of the individual samples and the position data of the corresponding neighbor samples; and calculating a correction value of the gene expression data of the individual sample according to the gene expression data of the neighbor sample corresponding to the individual sample and the weight. The data correction method can solve the problems of high noise and high sparsity in data and has the advantage of enhancing features, so that noise interference is reduced, and reliable data is provided for downstream analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Data correction method, system, and related equipment and medium based on Gaussian smoothing Technical Field

[0001] The present invention belongs to the field of biotechnology, and more specifically, relates to a data correction method and system based on a Gaussian smoothing model, and related computer equipment and storage media. Background Art

[0002] High-throughput analysis of in situ gene expression represents a major advance in our systematic understanding of tissue complexity. Application of this technology with sufficient capture area and high sample throughput has helped define the spatiotemporal dynamics of gene expression in tissues and organisms. Stereo-seq spatial transcriptome sequencing is a high-throughput spatiotemporal transcriptome sequencing technology. This technology uses two sequencing steps to independently determine the spatial location of mRNA sequences and their corresponding expression levels. This technology enables high-throughput transcriptome analysis of tissue sections at nanometer-scale resolution, scalable to centimeter-scale areas, with high sensitivity and uniform capture rate. In this technique, DNA nanoballs (DNBs) containing random barcode sequences are first deposited onto a photolithographically etched modified silicon surface (chip). This surface is engraved with an array of spots in a grid pattern, into which DNBs are effectively docked. Each spot has a diameter of approximately 220 nm and a center-to-center distance of 500 or 715 nm. The array is then micrographed, incubated with primers, and sequenced to generate a data matrix containing coordinate identifiers (CIDs) for each etched DNB. Then, by hybridization with CID, molecular identifiers (MIDs) or unique molecular identifiers (UMIs) and oligonucleotides containing polyT sequences are attached to each point, where UMI can distinguish whether the measured reads are derived from different amplified copies of mRNA molecules or from independent mRNA molecules for more accurate quantification. The next step includes capturing tissue polyA-tailed RNA, which is achieved by loading fresh nitrogen-frozen tissue sections onto the chip surface in situ, followed by fixation, permeabilization, and finally reverse transcription and amplification. The collected amplified cDNA is used as a template for library preparation and sequenced together with the CID. Computational analysis of sequencing data allows spatially resolved transcriptomics with a resolution of 500 or 715 nm and a minimum bin (the basic unit for analyzing data statistics in spatiotemporal omics technology) space.

[0003] Stereo-seq spatial transcriptome sequencing technology has been applied in multiple biological and medical fields and has become one of the mainstream methods for spatial transcriptome sequencing. However, this barcode-based multiple sequence alignment method has long been plagued by high noise and high dropout (e.g., high sparsity). True single-cell expression information cannot be perfectly presented in the expression matrix, which affects subsequent analysis and data mining.

[0004] Therefore, it is necessary to provide a data correction method and a corresponding system for performing data denoising and feature enhancement on an expression matrix.

[0005] Summary of the Invention

[0006] The purpose of the present invention is to propose a solution to at least partially solve or alleviate the above-mentioned problems in the prior art.

[0007] According to a first aspect of the present invention, a method for data correction based on a Gaussian smoothing model is provided, the method comprising: acquiring gene expression data and position data of a plurality of samples; processing neighboring samples of the plurality of samples, wherein the neighboring samples of an individual sample are determined based on the gene expression data of the plurality of samples, and the weights of the corresponding neighboring samples of the individual sample are calculated based on the position data of the individual sample and the position data of the corresponding neighboring samples; and calculating a correction value of the gene expression data of the individual sample based on the gene expression data of the neighboring samples corresponding to the individual sample and the weights.

[0008] According to a second aspect of the present invention, a system for data correction based on a Gaussian smoothing model is provided, the system comprising: a data acquisition module, the data acquisition module being configured to acquire gene expression data and position data thereof of a plurality of samples; a neighbor processing module, the neighbor processing module comprising a neighbor selection module and a weight calculation module, wherein the neighbor selection module is configured to determine the neighbor samples of an individual sample according to the gene expression data of the plurality of samples, and the weight calculation module is configured to calculate the weight of the corresponding neighbor samples of the individual sample according to the position data of the individual sample and the position data of the corresponding neighbor samples; and a data correction module, the data correction module being configured to calculate a correction value of the gene expression data of the individual sample according to the gene expression data of the neighbor samples corresponding to the individual sample and the weight.

[0009] According to the third aspect of the present invention, a system for correcting spatial group expression matrices based on a Gaussian smoothing model is provided, comprising a memory and a processor, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the method for correcting data based on the Gaussian smoothing model described above is executed.

[0010] According to a fourth aspect of the present invention, a computer device is provided, comprising a memory and a processor, wherein the memory stores computer instructions, which, when executed by the processor, cause the above-mentioned method for data correction based on the Gaussian smoothing model to be executed.

[0011] According to a fifth aspect of the present invention, a non-transitory computer-readable storage medium is provided, on which computer instructions are stored, which, when executed by a processor, cause the above-mentioned method for data correction based on the Gaussian smoothing model to be executed.

[0012] According to the sixth aspect of the present invention, a method for correcting a spatial group expression matrix based on a Gaussian smoothing model is provided, the method comprising inputting an expression matrix of a single cell having spatial coordinates into the system for correcting a spatial group expression matrix based on a Gaussian smoothing model according to the third aspect of the present invention or the computer device according to the fourth aspect of the present invention.

[0013] The present invention's data correction method based on a Gaussian smoothing model can address the problems of high noise and high sparsity in expression matrices, while also offering the advantage of enhanced features. Specifically, the present invention employs a novel expression matrix denoising algorithm based on similar cell and spatial information. Using a Gaussian smoothing model to enhance the features of expression matrix data (e.g., marker gene features), the algorithm accurately corrects cell expression information and reduces noise interference, enabling high-resolution correction of spatially derived biological data (e.g., gene expression data), providing reliable data for downstream analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0015] FIG1 shows a flow chart of a method for data correction based on a Gaussian smoothing model according to an embodiment of the present invention.

[0016] FIG2 shows a flow chart of a method for data correction based on a Gaussian smoothing model according to another embodiment of the present invention.

[0017] FIG3 shows a comparison diagram of the expression matrix of a cell before and after Gaussian smoothing according to another embodiment of the present invention.

[0018] FIG4 shows a before-and-after comparison of mouse brain data using the Gaussian smoothing algorithm of the present invention.

[0019] FIG5 shows a comparison of annotation results of two cell types before and after smoothing using the Gaussian smoothing algorithm of the present invention.

[0020] FIG6 shows the clustering results of the UMAP dimensionality reduction before and after the smoothing process using the Gaussian smoothing algorithm of the present invention.

[0021] FIG7 shows a schematic block diagram of a system for data correction based on a Gaussian smoothing model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the above and other features and advantages of the present invention clearer, the present invention is further described below in conjunction with the accompanying drawings. It should be understood that the specific embodiments provided herein are for the purpose of explaining to those skilled in the art and are merely exemplary and non-restrictive. For those of ordinary skill in the art, the specific meanings of terms in the present invention can be understood according to specific circumstances, unless otherwise clearly defined. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless otherwise clearly defined, without contradicting each other.

[0023] The Stereo-seq single-cell expression matrix is ​​segmented by ssDNA images, the cell nucleus range is circled, and then aligned to the DNB expression spectrum, which is divided into the expression levels of each cell for downstream analysis. Due to the limitations of the technology (the gene capture rate is not high, mRNA has diffusion problems, the circled cell size is only the cell nucleus, etc.), the expression matrix obtained is a highly sparse, high-noise data, which directly affects the downstream feature extraction and bioinformatics mining. The present invention is based on the basis of "similar cells" and the correction weight of spatial distance information. The expression matrix is ​​subjected to data denoising and feature enhancement through a Gaussian model, thereby correcting the data preprocessing method of technical errors in the sequencing process. The smoothed expression matrix can be stored for downstream analysis.

[0024] In the present invention, the data correction (for example, spatial group expression matrix correction or denoising) system based on the Gaussian smooth model can be hardware or software form, thereby is used in conjunction with or merges with other hardware or software systems. The hardware or software system can be a computer (for example desktop, workstation or server), or can be a sequencing system. When the sequencing system comprises the spatial group expression matrix denoising system based on the Gaussian smooth model of the present invention, the spatial group expression matrix denoising system based on the Gaussian smooth model of the present invention can be integrated into the sequencing system in the form of software or hardware, thereby the spatiotemporal group expression data measured by the sequencing system are directly processed, and the expression matrix after smoothing is obtained. Therefore, in the invention, software modules or hardware modules can be used to carry out the method steps listed in the first aspect of the present invention.

[0025] 1 shows a flow chart of a method 100 for data correction based on a Gaussian smoothing model according to an embodiment of the present invention. The method 100 includes steps S102, S104, and S106.

[0026] In step S102, obtain the gene expression data and positional data thereof of multiple samples. In one embodiment, the gene expression data of described multiple samples derive from space group data. This space group data is the transcriptome data actually measured utilizing spatial sequencing technology. The expression matrix in the space group data can be selected in whole or in part to carry out denoising. In one embodiment, this space group data can be a certain value, a certain feature or feature expression, a certain cell or cell expression or other field-related value, object or its expression amount. In another embodiment, this positional data is the positional data at the location before forming the gene expression data of multiple samples. Alternatively, in one embodiment, this positional data can also be the positional data at the location when or after forming the gene expression data of multiple samples.

[0027] In step S104, neighboring samples of the plurality of samples are processed. Step S104 includes step S1041 and step S1042.

[0028] In step S1041 , neighboring samples of an individual sample are determined based on the gene expression data of the multiple samples.

[0029] In step S1042 , the weights of the neighboring samples corresponding to the individual sample are calculated based on the position data of the individual sample and the position data of the corresponding neighboring samples.

[0030] In step S106 , a correction value of the gene expression data of the individual sample is calculated based on the gene expression data of the neighboring samples corresponding to the individual sample and the weight.

[0031] In one embodiment, the individual sample is each of one or more samples, or each of all samples.For example, data correction can be performed on the first sample included in the individual samples.

[0032] Specifically, for a first sample, a neighboring sample of the first sample is obtained from the plurality of samples based on the gene expression data of the plurality of samples. A weight of the neighboring sample is calculated based on the position data of the first sample and the position data of its neighboring sample. A correction value of the gene expression data of the first sample is calculated based on the gene expression data of the neighboring sample corresponding to the first sample and the weight of the neighboring sample.

[0033] In one embodiment, the gene expression data of the plurality of samples are dimensionality reduced, and the neighboring samples of the plurality of samples are processed in the dimensionality reduced space, as described in FIG2 below.

[0034] 2 shows a flow chart of a method 200 for data correction based on a Gaussian smoothing model according to another embodiment of the present invention. The method 200 includes steps S202, S203, S204, S206, and S208.

[0035] In step S202 , similar to step S102 , gene expression data and position data of multiple samples are obtained.

[0036] In step S203 , dimension reduction is performed on the gene expression data of the multiple samples to eliminate noise.

[0037] In step S204, neighboring samples of the plurality of samples are processed in the dimensionally reduced space. Step S204 includes step S2041, step S2041' and step S2042.

[0038] In step S2041 and step S2041', the neighboring samples of the individual sample are determined based on the gene expression data of the multiple samples. In one embodiment, in step S2041, the distance between the individual sample and other samples is calculated based on the gene expression feature vector of the individual sample obtained from the gene expression data of the individual sample; and in step S2041', for the individual sample, the neighboring samples of the individual sample are selected from the other samples based on the distance. In another embodiment, in step S2041, for the individual sample, other samples are obtained from the multiple samples as neighboring samples, for example, other samples are selected within a certain radius with the individual sample as the center; and in step S2041', the distance between the individual sample and the other samples obtained is calculated.

[0039] In another embodiment, the characteristic distance (i.e., 1-similarity) between the individual sample and other samples is calculated based on the gene expression feature vector of the individual sample obtained from the gene expression data of the individual sample, and the positional distance between the individual sample and other samples is calculated based on the position data of the individual sample and other samples, and a unified measure is used for the characteristic distance and the positional distance; for the individual sample, a neighboring sample of the individual sample is selected from other samples based on the distance measured by the unified measure.

[0040] In step S2042 , similar to step S1042 , the weights of the neighboring samples corresponding to the individual sample are calculated based on the position data of the individual sample and the position data of the corresponding neighboring samples.

[0041] In step S206 , similar to step S106 , a correction value of the gene expression data of the individual sample is calculated based on the gene expression data of the neighboring samples corresponding to the individual sample and the weight.

[0042] In step S208, the correction value is output.

[0043] In one embodiment, each sample in the plurality of samples comprises at least one or more cells.

[0044] In one embodiment, the gene expression data of the individual sample is further corrected by the correction value. For example, the gene expression data of a corrected individual sample can be used to correct the gene expression data of another individual sample.

[0045] With reference to Figures 1 and 2, in one embodiment, based on the gene expression characteristics of the multiple samples, the similarity between the individual sample and other samples is calculated, and the k other individual samples with high similarity to the gene expression characteristics of the individual sample are used as the nearest neighbor samples of the individual sample. That is, for the distance in the PCA space based on the gene expression characteristics, the closer the distance, the higher the similarity, and the farther the distance, the lower the similarity. In the PCA space, the k samples with the closest distance are selected as the nearest neighbors of the individual sample. In addition, the Pearson similarity between each sample can be calculated and sorted, and the k samples with the maximum similarity are selected as the nearest neighbors.

[0046] In one embodiment, determining the neighboring samples of an individual sample according to the gene expression data of the multiple samples includes: calculating the distance between the individual sample and other samples based on the position data associated with the gene expression data of the multiple samples, and obtaining the top k samples closest to the individual sample based on the distance as k neighbors similar to the individual sample and therefore as the neighboring samples.

[0047] For example, in step S1041, for an individual sample, its k neighbors are obtained, and the k neighbors come from the spatial group data comprising the multiple samples. Specifically, the acquisition of k neighbors can be directly obtained or obtained later by relevant calculations. In one embodiment, the Euclidean distances of all other samples except the individual sample relative to the individual sample are calculated, and the first k nearest other samples are selected as the neighbors of the individual sample. In one embodiment, the most similar k sample data can be selected as nearest neighbors by calculating the similarity (for example, Pearson similarity) between the sample data. Preferably, as described in step S203 and steps S2041 and S2041 ', by carrying out dimensionality reduction (for example, PCA dimensionality reduction) to the acquired spatial group data, k neighbors of the individual sample are obtained in the space after dimensionality reduction as nearest neighbor samples.

[0048] 1 and 2 , in one embodiment, the step of calculating the weight of the neighboring samples corresponding to the individual sample based on the location data of the individual sample and the location data of the corresponding neighboring samples includes:

[0049] The weights of the k neighbors as the nearest neighbor samples are calculated based on the Gaussian smoothing model using the following formula:

[0050] w i =GS(R(x,i))i∈N x (k),

[0051] Where x represents the individual sample, N x (k) represents the k neighbors corresponding to the individual sample, w i represents the Gaussian weight of the i-th neighbor among the k neighbors for correcting the gene expression data of the individual sample,

[0052] And R is the Euclidean distance between the position data of the individual sample and the position data of the corresponding neighboring samples, GS is the Gaussian formula, and the formula is as follows:

[0053]

[0054] Where a, b, and c are three hyperparameters in the formula.

[0055] For example, in step S1042, the weights of the k neighbors are calculated based on the Gaussian smoothing model. Alternatively, in one embodiment, the weights of the k neighbors may be calculated based on a Gamma distribution model or an inverse function model.

[0056] 1 and 2 , in one embodiment, the step of calculating the corrected value of the gene expression data of the individual sample based on the gene expression data of the neighboring samples corresponding to the individual sample and the weight includes:

[0057] The correction value of the gene expression data of the individual sample is calculated based on the gene expression data of the k neighbors and the Gaussian weights of the k neighbors, and is implemented by the following smoothing formula:

[0058]

[0059] Where x represents the individual sample, N x (k) represents the k neighbors corresponding to the individual sample, w i represents the Gaussian weight of the i-th neighbor among the k neighbors correcting the gene expression data of the individual sample, exp i represents the gene expression data of the i-th neighbor among the k neighbors, exp′ x Represents the corrected value of the gene expression data of the individual sample after Gaussian smoothing.

[0060] In one embodiment, in step S106 , the method further comprises: calculating a correction value of the gene expression data of the individual sample according to the gene expression data of the i-th neighbor and the Gaussian weight of the i-th neighbor.

[0061] As an example, when the individual samples are cells, the expression matrix data of the cells is corrected. In one embodiment, by performing dimensionality reduction on the expression matrix data of multiple samples including the expression matrix data of the cell, k similar neighbor cells of the cell are obtained in the reduced dimensionality space, as described in detail below with reference to FIG2 .

[0062] In step S202 , the spatial transcriptome expression matrix and the spatial coordinates of the cells are input.

[0063] In step S203, the expression matrix is ​​preprocessed and subjected to principal component analysis (PCA) dimensionality reduction, where dimensionality reduction is performed to eliminate noise and identify the most similar cells. In one embodiment, without dimensionality reduction, as indicated by the dotted arrow in Figure 2 (directly from step S202 to step S204), the k most similar cells can be selected as nearest neighbors by calculating similarities between cells (e.g., Pearson similarity), and subsequent analysis remains unchanged.

[0064] In step S2041 , the distance between each cell (eg, single cell x) and other cells is calculated in the PCA space.

[0065] In step S2041 ′, k neighboring cells of each cell (eg, single cell x) are obtained based on the calculated distance.

[0066] Optionally, in step S2041, k neighboring cells of each cell (eg, single cell x) are obtained in the PCA space. In step S2041', the distances from the k cells to, eg, single cell x are calculated or extracted.

[0067] In step S2042, the Gaussian weights of k cells are calculated. Specifically, this is implemented using the following formula:

[0068] w i =GS(R(x,i))i∈N x (k),

[0069] Where x represents a single cell to be corrected, N x (k) represents the k nearest neighbors or similar cells of a single cell x, w i represents the Gaussian weight of the single cell x modified by the i-th cell among the k neighboring or similar cells, and wherein R is the Euclidean distance and GS is the Gaussian formula (as described herein).

[0070] In step S206, the expression of each cell is recalculated. Specifically, it is implemented by the following smoothing formula:

[0071]

[0072] Where x represents a single cell to be corrected, N x (k) represents the k nearest neighbors or similar cells of a single cell x, w i represents the Gaussian weight of the i-th cell among the k neighboring or similar cells to modify the single cell x, exp i represents the original expression level of the i-th cell among k neighboring or similar cells, exp′ x represents the expression level of single cell x after Gaussian smoothing (i.e., as a corrected value).

[0073] In step S208 , the smoothed expression matrix of the cell (eg, single cell x) is output (as shown in the right part of FIG. 3 below).

[0074] Figure 3 shows a comparison of cell expression matrices before and after Gaussian smoothing according to another embodiment of the present invention. The original data contains two matrices: an expression matrix and a cell coordinate matrix. The rows of the expression matrix represent cells, the columns represent genes, and the numbers in the matrix are the measured UMI counts. The UMI counts represent the absolute number of observed transcripts (e.g., for each gene, cell, or sample) (e.g., an integer value representing the UMI type). After Gaussian smoothing, the structure of the expression matrix remains unchanged, but all the values ​​in it are replaced by exp′. x (decimal), the result can be directly used as input for subsequent analysis.

[0075] For example, in combination with what is described in FIG2 , in the present invention, the spatial group expression matrix denoising method based on the Gaussian smoothing model may include the following steps:

[0076] 1) Using conventional Stereo-seq methods, we obtain the expression matrix and spatial coordinates of single cells (cell bins). Specifically, Stereo-seq performs two sequencing runs to obtain chip coordinates and the number of UMI counts of different genes for captured DNBs. Combined with the cell ssDNA image, the gene expression of each cell on the chip can be selected using the image and organized into an expression matrix and a cell coordinate matrix, as shown on the left side of Figure 3.

[0077] 2) Normalize and log-process the expression matrix of single cells with spatial coordinates;

[0078] In the present invention, the expression matrix is ​​normalized, log-normalized, and other preprocessing methods are performed. Here, the preprocessing method of scanpy can be used, as shown below:

[0079] https: / / scanpy-tutorials.readthedocs.io / en / latest / pbmc3k.html#Preprocessing,

[0080] That is, the conventional steps of filtering low-expression cells and genes, normalizing all genes in a single cell, and taking the UMI count value as log(UMI+1).

[0081] 3) PCA is then used to reduce the dimensionality of the preprocessed expression matrix, retaining principal components with dimensions of 50-200, and obtaining the k nearest neighbors in the PCA space. The dimension of the PCA, which is an important factor affecting downstream analysis, is determined by the actual bioanalysis requirements. Preferably, in experiments conducted by this group, a PCA component size of 50-100 is generally optimal. The k value is an empirical value, and its size affects the degree of correction. Preferably, the k value is 10, 50, or 100.

[0082] Alternatively, in neighbor selection, participating objects can be selected in other ways. For example, "neighbors" within a radius of r in the PCA space can be calculated as neighbor objects; for another example, an adaptive Gauss kernel can be used to select neighbor objects with a specific threshold.

[0083] 4) Calculate the Gaussian weights of the k neighbors using the actual physical distance within the DNA nanoball (DNB) chip (i.e., the Euclidean distance calculated from two cell coordinates, which is related to tissue structure and cell density) and substitute into the following formula:

[0084] w i =GS(R(x,i))i∈N x (k)

[0085] w i is the calculated Gaussian weight, N x (k) represents the k nearest neighbor cells of cell x determined above, R is the Euclidean distance, and GS is the Gaussian formula, as follows:

[0086]

[0087] a, b, and c are the three hyperparameters in this formula. Adjusting these parameters affects the extreme values, center point, and slope of the Gaussian function curve. Preferably, a = 1, b = 0, and c = 12500 are used as reference values. The c value is a hyperparameter empirically measured in mouse brain tissue using a 500nm chip.

[0088] In the Gaussian weight formula, other distribution models can be adopted, as mentioned above.

[0089] 5) For cell x, the smoothed expression matrix can be obtained using the following smoothing formula:

[0090]

[0091] exp′ x represents the expression of x cells after Gaussian smoothing (i.e., the corrected value), N x (k) represents the k nearest neighbor cells of cell x determined above, w i represents the weight of cell i repairing x calculated above, exp i Represents the MID counts (original expression value, that is, the UMI count value of the original expression matrix) of cell i.

[0092] The methods or algorithms of the present invention can adjust thresholds and set correction ranges. In one embodiment, the algorithm defaults to smoothing the entire expression matrix. Users can then select cells or genes to smooth specific rows and columns of the expression matrix, creating a user-defined smoothing range. For example, the k value or hyperparameters a, b, and c can be adjusted.

[0093] The above examples of the present invention are algorithms developed for the inherent characteristics of Stereo-seq spatial transcriptome data, but in principle the algorithm can be applied to the pre-processing of other spatial group data (e.g., 10X, slide-seq, etc.), and its scope of application is not limited to high-resolution correction based on spatial gene expression data. The present invention can also be applied to other fields where data denoising is required, and any idea of ​​using the above characteristics to perform data correction falls within the scope of the present invention.

[0094] Figure 4 shows a before and after comparison of mouse brain data using the Gaussian smoothing algorithm of the present invention. The algorithm of the present invention is applied to the preprocessing part of Stereo-seq spatial transcriptome data analysis. Figure 4 shows that after smoothing using this algorithm, the marker genes in the mouse brain data have the effect of enhancing features and reducing noise. The figure illustrates the comparison of the effects of three marker genes after Gaussian smoothing. The in situ hybridization results on the right can be regarded as a reference for the true expression of the gene in the brain region. As can be seen from the figure, the expression characteristics of the marker genes after Gaussian smoothing are more obvious, the noise in other positions is lower, and their expression pattern is closer to the results of the in situ hybridization experiment.

[0095] Figure 5 shows a comparison of the annotation results of two cell types before and after smoothing. The upper figure shows the annotation of two cell types before smoothing, and the lower figure shows the annotation results of the same cell types after smoothing. After smoothing, more cell types can be annotated, and it is consistent with biological evidence. Specifically, the algorithm of the present invention is applied to the preprocessing part of Stereo-seq spatial transcriptome data analysis. Subsequent annotation using the spatialID annotation method is as follows:

[0096] https: / / www.biorxiv.org / content / 10.1101 / 2022.05.26.493527v1).

[0097] Comparing the impact of smoothing on subsequent annotation results with and without the smoothing algorithm. Figure 3 shows an example where smoothing allowed for the annotation of two additional cell types, whose distributions, when compared to the mouse brain atlas, are consistent with biological understanding.

[0098] Figure 6 shows the clustering results for UMAP dimensionality reduction. Uniform Manifold Approximation and Projection (UMAP) is a nonlinear dimensionality reduction algorithm based on Riemannian geometry and algebraic topology. The left figure shows the UMAP result before smoothing, and the right figure shows the UMAP result after smoothing. Different shapes represent different clusters. After smoothing, the distinction between clusters is more distinct.

[0099] In one embodiment, the algorithm of the present invention is applied to the preprocessing portion of Stereo-seq spatial transcriptome data analysis. PCA dimensionality reduction and clustering results under UMAP coordinates are subsequently performed. In one embodiment, a comparison of UMAP results using Gaussian smoothing preprocessing and not using smoothing preprocessing shows a significant change in shape, with more data features applied to cluster analysis. This demonstrates that smoothing can make the differences between classes more distinct, allowing for the clustering of more classes.

[0100] 7 shows a schematic block diagram of a system 700 for data correction based on a Gaussian smoothing model according to an embodiment of the present invention, wherein the system 700 includes a data acquisition module 702 , a neighbor processing module 704 , and a data correction module 706 .

[0101] In one embodiment, data acquisition module 702 is configured to obtain gene expression data and positional data thereof of multiple samples.Neighbor processing module 704 comprises neighbor selection module 7041 and weight calculation module 7042, wherein said neighbor selection module is configured to determine wherein the nearest neighbor sample of individual sample according to the gene expression data of described multiple samples, and described weight calculation module is configured to calculate the weight of the corresponding nearest neighbor sample of described individual sample according to the positional data of described individual sample and the positional data of corresponding nearest neighbor sample.Data correction module 706 is configured to calculate the corrected value of the gene expression data of described individual sample according to the gene expression data of the nearest neighbor sample corresponding to described individual sample and described weight.

[0102] The features mentioned in the description of each embodiment of the data correction method can also be extended and applied to the data correction system of this article and the space group expression matrix correction system and space group expression matrix correction method based on the Gaussian smoothing model in the context (see the detailed description in the context). For the sake of simplicity, they are not described here in detail.

[0103] According to another aspect of the present invention, a system for correcting spatial group expression matrices based on a Gaussian smoothing model is also provided, comprising a memory, a processor and corresponding components, wherein computer instructions are stored on the memory, and when the computer instructions are executed by the processor, the method described in the context is executed.

[0104] The present invention also provides a computer device comprising a memory and a processor, wherein the memory stores computer instructions, which, when executed by the processor, cause the method described above and below to be performed.

[0105] The present invention also provides a non-transitory computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, result in the method described above and below being performed.

[0106] According to another aspect of the present invention, a method for correcting a spatial group expression matrix based on a Gaussian smoothing model is also provided, the method comprising inputting an expression matrix of a single cell with spatial coordinates into the system for correcting a spatial group expression matrix based on a Gaussian smoothing model or the computer device described above, as described in detail in the context.

[0107] Those skilled in the art will appreciate that the embodiments of the present application can be provided as a system or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete module embodiment, or an embodiment combining modules and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0108] The application is described with reference to the flow chart and / or block diagram according to the method, equipment (system) and computer program product of embodiment of the present application.It should be understood that each flow process and / or box in the flow chart and / or block diagram and the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of general-purpose computer, special-purpose computer, embedded processing machine or other programmable data processing equipment to produce a machine, so that the instruction executed by the processor of computer or other programmable data processing equipment produces the device for realizing the function specified in flow chart one flow chart or multiple flow charts and / or one box or multiple boxes of block diagram.

[0109] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0111] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0112] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0113] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0114] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0115] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for data correction based on a Gaussian smoothing model, characterized in that: The method comprises: Obtain gene expression data and location data of multiple samples; Processing neighboring samples of the plurality of samples, wherein: Determine neighboring samples of the individual sample according to the gene expression data of the multiple samples; and Calculating the weight of the neighboring samples corresponding to the individual sample according to the position data of the individual sample and the position data of the corresponding neighboring samples; and The correction value of the gene expression data of the individual sample is calculated according to the gene expression data of the neighboring samples corresponding to the individual sample and the weight.

2. The method according to claim 1, characterized in that The individual sample is each of one or more samples, or each of all samples.

3. The method according to claim 1, characterized in that The method further includes: performing dimensionality reduction on the gene expression data of the multiple samples, and processing neighboring samples of the multiple samples in the dimensionality reduced space.

4. The method according to claim 1, characterized in that: Each sample in the plurality of samples includes at least one or more cells.

5. The method according to any one of claims 1 to 4, characterized in that The gene expression data of the individual sample is further corrected by the correction value.

6. The method according to any one of claims 1 to 4, characterized in that Determining the neighboring samples of the individual samples according to the gene expression data of the multiple samples comprises: Based on the gene expression features of the multiple samples, the similarity between the individual sample and other individual samples is calculated, and k other individual samples having high similarity with the gene expression features of the individual sample are used as neighboring samples of the individual sample.

7. The method according to claim 1, characterized in that The step of calculating the weight of the neighboring sample corresponding to the individual sample according to the position data of the individual sample and the position data of the corresponding neighboring sample comprises: The weights of the k neighbors as the nearest neighbor samples are calculated based on the Gaussian smoothing model by the following formula: w i =GS(R(x,i))i∈N x (k), Where x represents the individual sample, N x (k) represents the corresponding k neighbors of the individual sample, w i represents the Gaussian weight of the i-th neighbor among the k neighbors correcting the gene expression data of the individual sample, And R is the Euclidean distance between the location data of the individual sample and the location data of the corresponding neighboring sample, GS is the Gaussian formula, and the formula is as follows: Where a, b, and c are three hyperparameters in the formula.

8. The method according to claim 1, characterized in that The step of calculating the correction value of the gene expression data of the individual sample according to the gene expression data of the neighboring sample corresponding to the individual sample and the weight comprises: The correction value of the gene expression data of the individual sample is calculated according to the gene expression data of the k neighbors and the Gaussian weights of the k neighbors, and is implemented by the following smoothing formula: Where x represents the individual sample, N x (k) represents the corresponding k neighbors of the individual sample, w i represents the Gaussian weight of the i-th neighbor among the k neighbors for correcting the gene expression data of the individual sample, exp i represents the gene expression data of the i-th neighbor among the k neighbors, exp′ x Represents the corrected value of the gene expression data of the individual sample after Gaussian smoothing.

9. The method according to claim 8, characterized in that The method further includes: calculating a correction value of the gene expression data of the individual sample according to the gene expression data of the i-th neighbor and the Gaussian weight of the i-th neighbor.

10. A data correction system based on a Gaussian smoothing model, characterized in that: The system comprises: A data acquisition module, wherein the data acquisition module is configured to acquire gene expression data and position data of multiple samples; A neighbor processing module, the neighbor processing module includes a neighbor selection module and a weight calculation module, wherein: The neighbor selection module is configured to determine neighbor samples of the individual samples according to the gene expression data of the multiple samples, and The weight calculation module is configured to calculate the weight of the neighboring samples corresponding to the individual sample according to the position data of the individual sample and the position data of the corresponding neighboring samples; and A data correction module is configured to calculate a correction value of the gene expression data of the individual sample according to the gene expression data of the neighboring samples corresponding to the individual sample and the weight.

11. A system for space group expression matrix correction based on Gaussian smoothing model, characterized in that: The invention comprises a memory and a processor, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the method for data correction based on a Gaussian smoothing model according to any one of claims 1 to 9 is executed.

12. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the method for data correction based on a Gaussian smoothing model according to any one of claims 1 to 9 is executed.

13. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which, when executed by a processor, result in the execution of the method for data correction based on a Gaussian smoothing model according to any one of claims 1 to 9.

14. A method for correcting a space group expression matrix based on a Gaussian smoothing model, characterized in that: The method comprises inputting the expression matrix of a single cell having spatial coordinates into the system for spatial group expression matrix correction based on Gaussian smoothing model according to claim 11 or the computer device according to claim 12.