A method and system for spatially variable gene identification for spatial transcriptome data

By employing semi-pooling processing and combined testing methods, the accuracy and speed issues of traditional models in identifying spatially variable genes were resolved, achieving efficient and accurate identification of spatially variable genes.

CN116453597BActive Publication Date: 2026-01-02SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310369928.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2026-01-02
Estimated Expiration
2043-04-07

AI Technical Summary

Technical Problem

Traditional spatial statistical models fail when dealing with large-scale, structurally complex, high-dimensional, and sparse spatial transcriptomics data, making it difficult to accurately identify spatially variable genes.

Method used

The original gene expression data were processed using a semi-pooling method. The stability test results were combined with the Box-Pierce test and the Stouffer combined method. The holm method was used to correct the false positive rate and identify spatially variable genes.

Benefits of technology

It improves the accuracy and computational speed of identifying spatially variable genes, effectively identifies genes with spatial expression patterns, and controls the false positive rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116453597B_ABST
    Figure CN116453597B_ABST
Patent Text Reader

Abstract

The application relates to a spatial variable gene identification method for spatial transcriptomics data, which comprises the following steps: data conversion and feature extraction on original data through a semi-pooling method; stability test on output data obtained through the semi-pooling process; and combination test on the stability test result, so as to identify the spatial variable gene. Compared with the prior art, the application has the advantages of high identification accuracy and fast calculation speed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and in particular to a spatial variable gene identification method and system for spatial transcriptome data. BACKGROUND

[0002] The rapid development of spatial transcriptomics technology has promoted the research on tissue structure reconstruction, development and disease, and large-scale spatial transcriptomics research has become increasingly popular. One of the most important and unique problems in spatial transcriptomics analysis methods is to identify spatial variable genes. The specific meaning of spatial variable genes is that the gene expression has a certain spatial pattern in the spatial distribution of the tissue. From the data, the expression count of the spatial variable gene has a certain relationship with the spatial position.

[0003] Traditional spatial statistical models often fail to deal with spatial transcriptomics data with large number, complex structure, high dimension and sparsity, so it is necessary to develop a spatial variable gene identification method suitable for the characteristics of spatial transcriptomics data. SUMMARY

[0004] The purpose of the present application is to overcome the defects of the prior art and provide a spatial variable gene identification method and system for spatial transcriptome data with high accuracy and fast calculation speed.

[0005] The purpose of the present application can be achieved by the following technical solutions:

[0006] According to a first aspect of the present application, a spatial variable gene identification method for spatial transcriptomics data is provided, which comprises the following steps:

[0007] Step S1, performing semi-pooling processing on the original gene expression data of each gene;

[0008] Step S2, performing stability test on the output data after semi-pooling processing;

[0009] Step S3, combining test is performed on a plurality of stability test results;

[0010] Step S4, judging whether it is a spatial variable gene according to the combination test result.

[0011] Preferably, the semi-pooling processing in step S1 is specifically: performing average value calculation on the spatial transcriptomics data according to the given K sets of semi-pooling parameters, and rearranging the obtained output data into a one-dimensional sequence according to the spatial position; wherein the semi-pooling parameters include direction parameters and step parameters.

[0012] Preferably, the semi-pooling processing includes four different semi-pooling parameters, which are:

[0013] 1) Direction: row direction, step size: n row ;

[0014] 2) Direction: row direction, step size:

[0015] 3) Direction: column direction, step size: n col ;

[0016] 4) Direction: column direction, step size:

[0017] wherein n col is the number of columns contained in the spatial transcriptomic data, n row is the number of rows contained in the spatial transcriptomic data, and [·] represents the integer part.

[0018] Preferably, the stability test in step S2 is a Box-Pierce test, which is used to perform a stability test on the output data processed for different semi-pooling parameters respectively.

[0019] Preferably, the parameter setting in the Box-Pierce test comprises: a maximum delay order parameter m = [ln(T)], wherein T is the length of the output data after semi-pooling processing, and [·] represents the integer part.

[0020] Preferably, the combination test in step S3 adopts a Stouffer combination method, and the specific calculation method is:

[0021]

[0022] wherein Φ -1 (·) is the inverse function of the cumulative distribution function of the standard normal distribution, K is the number of groups of semi-pooling parameters, and N(0, 1) is the standard normal distribution.

[0023] Preferably, the step S4 further comprises holm method correction on the combination test result.

[0024] According to a second aspect of the present application, a spatial variable gene recognition system based on spatial transcriptomic data is provided, which comprises:

[0025] a semi-pooling processing module, configured to perform semi-pooling processing on the original gene expression data of each gene;

[0026] a stability test module, configured to perform a stability test on the output data after semi-pooling processing;

[0027] a combination test module, configured to perform a combination test on a plurality of stability test results;

[0028] The spatial variable gene judging module is configured to judge whether the spatial variable gene exists according to the combination test result.

[0029] According to a third aspect of the present application, there is provided an electronic device comprising a memory and a processor, the memory having stored thereon a computer program, the processor being configured to implement any of the methods described.

[0030] According to a fourth aspect of the present application, there is provided a computer readable storage medium having stored thereon a computer program, the program being configured to implement any of the methods described when executed by a processor.

[0031] Compared with the prior art, the present application has the following advantages:

[0032] 1) The present application performs data conversion and feature extraction on the original data by the semi-pooling method, performs stability test on the output data obtained by the semi-pooling processing, and performs combination test on the stability test results, thereby identifying the spatial variable gene, which has the advantages of high identification accuracy and fast calculation speed;

[0033] 2) The present application uses the semi-pooling method containing direction parameters and step parameters for data conversion and feature extraction, which is used for large-scale spatial transcriptome data with large quantity, complex structure, high dimension and sparsity;

[0034] 3) The Box-Pierce test is used to perform stability test on the output data after semi-pooling processing, which has high accuracy;

[0035] 4) The Stouffer combination method is used to perform combination test on multiple stability test results, which improves the accuracy of the test results;

[0036] 5) The P value of the combination test is corrected by the holm method, which can effectively control the false positive rate and improve the identification accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 The figure is a flow chart of the spatial variable gene identification method of the present application.

[0038] Figure 2 The figure is a specific implementation schematic diagram of the semi-pooling processing step of the present application.

[0039] Figure 3 The figure is a part of the results of the embodiment of the present application, which shows the top 20 spatial variable genes identified in the embodiment. DETAILED DESCRIPTION

[0040] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the protection scope of the present application.

[0041] Embodiments

[0042] The present application relates to a spatially variable gene identification method for large-scale spatial transcriptome data, which aims to identify spatially variable genes in large-scale spatial transcriptome data, and the method comprises the following steps:

[0043] Step S1, semi-pooling processing is performed on the original gene expression data of each gene;

[0044] Step S2, stability test is performed on the output data after semi-pooling processing;

[0045] Step S3, combination test is performed on the plurality of stability test results;

[0046] Step S4, it is judged whether it is a spatially variable gene according to the combination test result.

[0047] Next, the specific implementation of the method of the present embodiment will be described in detail.

[0048] The present embodiment uses a spatial transcriptome data of a colorectal cancer tissue, and the data set is a publicly available data set (http: / / www.cancerdiversity.asia / scCRLM / ).

[0049] 1, data set preprocessing

[0050] Filtering genes not expressed or lowly expressed in the original data set, the filtering standard used in the present embodiment is: filtering out genes with expression ratio less than 1% in all spots. The filtered data set includes 15427 genes and 4124 spots, including 78 rows and 128 columns.

[0051] 2, semi-pooling processing

[0052] The expression data of each gene in space is calculated according to the given direction parameter and step parameter, and the specific four groups of parameters are as follows:

[0053] 1) direction: row direction, step: 78;

[0054] 2) direction: row direction, step:

[0055] 3) Direction: column direction, step: 128;

[0056] 4) Direction: column direction, step:

[0057] where [·] denotes the integer part, and the semi-pooling processing schematic diagram is shown in FIG. 2. Figure 2

[0058] 3. Stability test

[0059] According to the four sets of semi-pooling parameters given above, each gene obtains four new output sequences after semi-pooling processing, and the lengths of the four output sequences are 128, 584, 78, and 390 respectively. For each semi-pooling processed output sequence r = (r1,..,r t ,…,r T ) T , a stability test, i.e. Box-Pierce test, is performed. The Box-Pierce test is a test method for testing the autocorrelation of sequence data, and the test statistic Q m obeys the χ 2 distribution with the degree of freedom m. The calculation method of the Box-Pierce test statistic is as follows:

[0060]

[0061]

[0062]

[0063] where denotes the autocorrelation coefficient, denotes the autocovariance, r = (r1,..,r t ,…,r T ) T is the output data after semi-pooling processing, is the mean of r, m = [ln(T)], T is the length of the semi-pooling processed output sequence, and [·] denotes the integer part.

[0064] The degrees of freedom of the Box-Pierce test of the four output sequences are 4, 5, 6, and 6 respectively. The stability test is performed on the four semi-pooling output sequences of each gene, and the corresponding P values are obtained.

[0065] 4. Combination test

[0066] ​The P values of the 4 stability tests for each gene are combined using the Stouffer's combination method, which converts the p values of multiple independent hypothesis tests into one p value. If there are h p values in total, the combination is calculated as:

[0067]

[0068] where Φ -1 (·) is the inverse function of the cumulative distribution function of the standard normal distribution, which is specifically: where erf -1 (x) is the inverse function of the error function, which is defined as finding a number y such that erf(y) = x. The inverse function of the error function does not have a simple analytical expression and is usually calculated using numerical methods.

[0069] The P value of the combined test is calculated as:

[0070] p c = 1 - Φ(z stouffer )

[0071] 5. P value correction

[0072] In order to control the false positive rate, the P value of the combined test is corrected using the holm method. The genes with P value less than 0.05 after holm correction are considered as spatially variable genes. The holm method is a commonly used multiple comparison correction method for controlling the error rate. Specifically, first, sort all P p values from small to large as: p (1) ,p (2) ,…,p (rank) ,…,p P , and calculate the correction factor for each sorted p value: correction factor (rank) = P-rank+1, and the corrected p value is:

[0073] In this example, 8020 spatially variable genes are identified, and the top 20 genes selected in this example are given in the attached table Figure 3 It can be seen that there is a clear spatial expression pattern, which shows that this example can effectively identify genes with spatial expression patterns.

[0074] Specifically, taking gene B2M as an example, after the spatial expression data of the gene B2M in the example is semi-pooled according to the given four groups of parameters, four new output sequences can be obtained, and the P values of the stability test of each output sequence are <2.2e-16, <2.2e-16, <2.2e-16, and <2.2e-16 respectively. The P value of the combination test is 0, and the corrected P value is 0, so B2M is a spatial variable gene, and the Figure 3 It can also be seen from the above description that the gene expression of B2M has a clear spatial pattern.

[0075] The electronic device of the present application includes a central processing unit (CPU) that can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded into a random access memory (RAM) from a storage unit. In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0076] A plurality of components in the device are connected to the I / O interface, including: an input unit such as a keyboard, a mouse, etc.; an output unit such as various types of displays, a speaker, etc.; a storage unit such as a magnetic disk, an optical disk, etc.; and a communication unit such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0077] The processing unit performs the various methods and processes described above, such as methods S1-S4. For example, in some embodiments, methods S1-S4 can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1-S4 described above can be performed. Alternatively, in other embodiments, the CPU can be configured to perform methods S1-S4 by any other appropriate means (e.g., by means of firmware).

[0078] The functions described above herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip systems (SOCs), complex programmable logic devices (CPLDs), etc.

[0079] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / operations specified in the flowchart diagrams and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.

[0080] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage medium can include, but are not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.

[0081] The above descriptions are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A spatially variable gene identification method for spatial transcriptomic data, characterized in that, The method comprises the following steps: Step S1, semi-pooling processing is performed on original gene expression data of each gene; Step S2, stability test is performed on output data after semi-pooling processing; Step S3, combination test is performed on multiple stability test results; Step S4, whether the gene is a spatial variable gene is determined according to the combination test result; The semi-pooling processing in the step S1 is specifically: average value calculation is performed on spatial transcriptomics data according to given K sets of semi-pooling parameters, and output data obtained is rearranged into a one-dimensional sequence according to spatial positions; wherein, the semi-pooling parameters include direction parameters and step parameters; The stability test in the step S2 is Box-Pierce test, which is used for stability test on output data processed by different semi-pooling parameters respectively; The combination test in the step S3 adopts Stouffer combination method, and the specific calculation method is: wherein is the inverse of the cumulative distribution function of the standard normal distribution, is the number of groups of semi-pooling parameters, is the standard normal distribution.

2. The method for spatially variable gene identification of spatial transcriptomic data according to claim 1, wherein, The semi-pooling processing includes four sets of different semi-pooling parameters, which are respectively: 1) Direction: row direction, step size: ; 2) Direction: row direction, step size: ; 3) Direction: Column direction, step size: ; 4) Direction: Column direction, step size: ; wherein, the number of columns contained by the spatial transcriptome data, the number of rows contained by the spatial transcriptome data, denotes the integer part.

3. The method for spatially variable gene identification of spatial transcriptomic data according to claim 1, wherein, The parameter setting in the Box-Pierce test includes a maximum delay order parameter wherein T is the length of the output data after half-pooling processing, denotes rounding.

4. The method for spatially variable gene identification of spatial transcriptomic data according to claim 1, wherein, The step S4 further includes holm method correction on the combination test result.

5. A spatially variable gene identification system for spatial transcriptomic data, characterized in that, The system comprises: A semi-pooling processing module, which is used for semi-pooling processing on original gene expression data of each gene; A stability test module, which is used for stability test on output data after semi-pooling processing; A combination test module, which is used for combination test on multiple stability test results; A spatial variable gene determination module, which is used for determining whether the gene is a spatial variable gene according to the combination test result. 6.An electronic device comprising a memory and a processor, the memory having stored thereon a computer program, characterized in that, The processor executes the program to implement the method according to any one of claims 1-4.

7. A computer readable storage medium having stored thereon a computer program, characterized in that The program is executed by the processor to implement the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Computer-implemented method and computer system for rank normalization for differential expression analysis of transcriptome sequencing data

    CN103377317A

  • Spatial mapping of nucleic acid sequence information

    CN108138225A