Method and System for Converting Spatial Transcriptome Data across Language Platforms
By using mapping files on R and Python language platforms to realize cross-language conversion of spatial transcriptome data, the problem of data structures not being lost-free is solved, tool development efficiency and storage space utilization are improved, and workflow construction is simplified.
Patent Information
- Application Number
- CN202210550264.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-20
AI Technical Summary
The existing data structures cannot meet the lossless conversion of spatial transcriptome data on R and Python platforms, and cannot easily integrate newly developed tools into existing workflows, resulting in developers spending a lot of time on data format conversion.
It provides a method and system for spatial transcriptome data conversion across language platforms. It realizes mutual conversion of data structures on R and Python language platforms through mapping files, including initializing mapping files, and reading and storing spatial transcriptome data on different platforms, supporting multiple interfaces and compatible with existing data structures.
It realizes the conversion of data structures on different platforms, is compatible with read and write interfaces of multiple sequencing data, saves storage space, improves the development efficiency of downstream analysis tools, and simplifies the construction process of workflows.
Smart Images

Figure CN115206439B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics transcriptome sequencing data analysis, and particularly to a method and system for converting spatial transcriptome data across language platforms. Background Art
[0002] The statements in this section only mention the background art related to the present invention and do not necessarily constitute prior art.
[0003] In recent years, with the development of spatial transcriptome sequencing technology and analysis technology, more and more third-party tools for analyzing spatial gene expression data have been developed under two platforms: the R language and the Python language. Certain achievements have been made in research on the spatial heterogeneity of cell expression, construction of spatial transcriptome maps of cells, etc.
[0004] The inventors found that the existing data structures cannot meet the lossless conversion between two spatial transcriptome software analysis platforms: the R language and Python, cannot meet the analysis requirements of a relatively long spatial transcriptome workflow, and cannot easily integrate newly developed tools into the existing workflow, and developers need to spend a lot of time on data format conversion issues. Summary of the Invention
[0005] To solve the deficiencies of the prior art, the present invention provides a method and system for converting spatial transcriptome data across language platforms; provides multiple interfaces for tools on the R language and Python, reduces the consumption of storage space at the same time, and provides read-write strategies compatible with existing data structures, making the construction of the workflow simpler, so as to give full play to the advantages of analysis tools on different platforms.
[0006] In the first aspect, the present invention provides a method for converting spatial transcriptome data across language platforms;
[0007] The method for converting spatial transcriptome data across language platforms includes:
[0008] Reading and storing the spatial transcriptome data, single-cell transcriptome reference data or intermediate results generated by spatial transcriptome analysis tools on the first language platform using a mapping file;
[0009] On the second language platform, reading the stored results and continuing to run;
[0010] Among them, the reading and storing using the mapping file specifically includes: initializing the mapping file;
[0011] When the first language platform is the R language platform, on the R language platform, implementing the mutual conversion operation between the mapping file and the data structure in the R language;
[0012] When the first language platform is the Python language platform; on the Python platform, implement the mutual conversion operation between the mapping file and the data structure in Python.
[0013] In a second aspect, the present invention provides a spatial transcriptome data conversion system across language platforms;
[0014] The spatial transcriptome data conversion system across language platforms includes:
[0015] A reading and storage module, which is configured to: read and store the spatial transcriptome data, single-cell transcriptome reference data, or intermediate results generated by spatial transcriptome analysis tools on the first language platform by using a mapping file;
[0016] An operation module, which is configured to: on the second language platform, read the stored results and continue to run;
[0017] Wherein, the reading and storage by using the mapping file specifically includes: initializing the mapping file;
[0018] When the first language platform is the R language platform; on the R language platform, implement the mutual conversion operation between the mapping file and the data structure in R language;
[0019] When the first language platform is the Python language platform; on the Python platform, implement the mutual conversion operation between the mapping file and the data structure in Python.
[0020] Compared with the prior art, the beneficial effects of the present invention are:
[0021] The solution described in the present disclosure can achieve the conversion of different data structures on the Python platform and the R platform, so that the intermediate results generated on one platform can be read and continued to run by the other platform; it can be compatible with multiple read and write interfaces of the original sequencing data, that is, it can read the mapping file like reading the original "h5" format file; it provides a unified file format for the original data generated by multiple sequencing technologies, which can save storage space while improving the development efficiency of downstream analysis tools; in addition, the data structure of the solution can also be used as the main data read and write center of the spatial transcriptome data analysis workflow, improving the simplicity, usability and scalability of the workflow. Description of the Drawings
[0022] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0023] Figure 1 It is the storage flowchart of the conversion intermediate file described in the embodiments of the present disclosure;
[0024] Figure 2 This is the structural type diagram of the conversion intermediate file described in the embodiments of the present disclosure;
[0025] Figure 3 This is a storage example of the sequencing data of the axial section of the mouse brain tissue obtained by the 10X Visium sequencing technology described in the embodiments of the present disclosure;
[0026] Figure 4 This is an example of the conversion of the Seurat structure in R language to the AnnData structure in Python through an intermediate file described in the embodiments of the present disclosure. Detailed implementation manners
[0027] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0028] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0029] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0030] All data acquisition in this embodiment is a legal application of the data on the basis of compliance with laws, regulations, and user consent.
[0031] Term explanation:
[0032] The R language is a computer language used for statistical analysis and drawing. Some important spatial transcriptome analysis tools, such as Seurat, Giotto, SpatialExperiment, etc., are deployed using this language;
[0033] The Python language is a computer programming language. Some important spatial transcriptome analysis tools, such as Scanpy, SpaGCN, stLearn, etc., are deployed using this language.
[0034] Embodiment 1
[0035] This embodiment provides a method for converting spatial transcriptome data across language platforms;
[0036] As Figure 1 shown, the method for converting spatial transcriptome data across language platforms includes:
[0037] S101: Read and store the spatial transcriptome data of the first language platform, the single-cell transcriptome reference data, or the intermediate results generated by the spatial transcriptome analysis tool using a mapping file;
[0038] S102: On the second language platform, read the stored results and continue to run;
[0039] Among them, the reading and storing using the mapping file specifically includes:
[0040] Initialize the mapping file;
[0041] When the first language platform is the R language platform; then on the R language platform, implement the mutual conversion operation between the mapping file and the data structure in the R language;
[0042] When the first language platform is the Python language platform; then on the Python platform, implement the mutual conversion operation between the mapping file and the data structure in Python.
[0043] Furthermore, the first language platform is the R language platform; the second language platform is the Python language platform.
[0044] Or, the first language platform is the Python language platform; the second language platform is the R language platform.
[0045] Furthermore, the spatial transcriptome data specifically includes:
[0046] An expression matrix composed of gene expressions (Features) at each sampling point (Barcodes), tissue section images (Images) at different resolutions, the specific positions (coordinates) of each sampling point in the tissue section image, and the scale factor (scaleFactors) between the original high-resolution image and the low-resolution image.
[0047] Furthermore, the single-cell transcriptome reference data includes:
[0048] A single-cell expression matrix instance composed of gene expressions (scFeatures) of each cell (Cell) in the scRNA-seq sequencing technology;
[0049] Single-cell highly variable genes (HVGs) calculated based on the single-cell expression matrix instance;
[0050] The clustering (clusters) of all cells in the expression matrix; and
[0051] Cell type (CellType).
[0052] Furthermore, the intermediate results generated by the spatial transcriptome analysis tool are specifically:
[0053] The spatial transcriptome expression matrix obtained after preprocessing the spatial expression matrix instance by noise reduction, decontamination, normalization, and gene filtering;
[0054] The spatially highly variable genes (HVGs) calculated based on the spatial expression matrix instance;
[0055] The expression values of each dimension after dimensionality reduction of the spatial expression matrix instance;
[0056] The clustering analysis results of all Barcodes in the spatial expression matrix instance and the differential genes of each cluster; and
[0057] The cell type composition of each Barcode obtained by combining single-cell reference data.
[0058] Furthermore, the initialization mapping file specifically includes:
[0059] Using the "rhdf5" package to create the structure of the mapping file and import the spatial transcriptome data into the mapping file.
[0060] Among them, the "rhdf5" package is an R language toolkit containing a series of functions for reading and writing h5 files; import refers to the assignment operation in the R language or Python language. Importing A to B means assigning the value of A to B.
[0061] Furthermore, the use of the "rhdf5" package to create the structure of the mapping file specifically includes:
[0062] Call the rhdf5::h5createGroup() function to create the basic structure of the mapping file for storing the expression matrix and tissue image. The "::" connector represents a function under a certain toolkit in the R language, that is, "toolkit::function".
[0063] Furthermore, the basic structure of the mapping file consists of two classes: the spatial image class (sptimages) and the expression matrix class (sptmatrix).
[0064] Among them, the spatial image class is used to store spatial information, namely histological images at different resolutions, the specific coordinates of each sampling point in the tissue section image, and the scaling factor between the original high-resolution image and the low-resolution image.
[0065] The expression matrix class is used to store the expression matrix and the intermediate results generated by the spatial transcriptomics analysis tool. Among them, the expression matrix is a sparse matrix in column storage format, including:
[0066] data, which is used to store the expression values of matrix elements;
[0067] indices, which is used to store the row numbers of elements;
[0068] indptr, which is used to represent the starting position of each row of elements;
[0069] shape, which is used to record the dimensions of the sparse matrix;
[0070] name, which is used to record the name of the expression matrix class instance.
[0071] The intermediate results with the same dimensions as Barcodes and Features are appended under the barcode item and the feature item respectively, and other data are appended in the unstructured information.
[0072] Furthermore, the importing of the spatial transcriptomics data into the mapping file specifically includes:
[0073] When the user needs to import 10X Visium spatial transcriptomics data, in the R language environment, call the rhdf5::h5read() function to read the expression matrix in "h5" format under the 10X Visium path, and call the rhdf5::h5write() function to import it into the sptmatrix instance in the mapping file; then call the png::readPNG() function to read the tissue images at various resolutions under the 10X Visium path, call the rjson::fromJSON() function to read the scaling factor under the 10X Visium path, call the Seurat::Read10X_Image() function to read the spatial coordinates of the barcodes under the 10X Visium path, and finally call the rhdf5::h5write() function to import the above information into the sptimages instance in the mapping file.
[0074] Alternatively, when the user needs to import Slide-seq spatial transcriptome data, first read the expression matrix in "count" format under the Slide-seq path in the R language environment, and call the rhdf5::h5write() function to import it into the sptmatrix instance in the mapping file; then read the index list in "idx" format under the Slide-seq path, and call the rhdf5::h5write() function to import it into the sptimages instance in the mapping file.
[0075] Furthermore, on the R language platform, the operation of converting between the mapping file and the data structure in R language is specifically as follows:
[0076] Import the data in the mapping file using the "rhdf5" package. Among them, the "rhdf5" package is an R language toolkit that contains a series of functions for reading and writing h5 files. Importing the data in the mapping file is specifically as follows:
[0077] S1, Use the rhdf5::h5read() function to import the mapping file into the R language environment;
[0078] S2, Read the column-stored sparse matrix of the sptmatrix instance in the mapping file, including data, indices, indptr, and shape. Finally, use the Matrix::sparseMatrix() function in R language to generate a spatial expression matrix or a single-cell expression matrix;
[0079] S3, Read the barcode items (Barcodes) of the sptmatrix instance in the mapping file, including information such as barcodes, annotations, cell types, etc., and the feature items (features) of the sptmatrix instance, including information such as gene names, gene numbers, and whether they are marker genes;
[0080] S4, Read the histological image information of the sptimages instance in the mapping file, including histological images at various resolutions, spatial coordinates of barcodes, and scale factors.
[0081] Furthermore, on the R language platform, the operation of converting between the mapping file and the data structure in R language also includes:
[0082] When the user uses the Seurat spatial transcriptome analysis tool, convert the content in the mapping file into a Seurat object, or save the analysis results of Seurat in the mapping file.
[0083] Specifically, convert the content in the mapping file into a Seurat object, including:
[0084] First, execute steps S1 to S4; second, call the Seurat::CreateSeuratObject() function to create a Seurat object, import the expression matrix in step S2 into the "@assay$spatial@data" member of the Seurat object, import the barcode items in step S3 into the "@meta.data" member of the Seurat object, and import the feature items in step S3 into the "@assay$spatial@meta.features" member of the Seurat object; finally, import the histological images at various resolutions in step S4 into the "@images$slice1@image" member of the Seurat object, import the spatial coordinates of the barcodes in step S4 into the "@images$slice1@coordinates" member of the Seurat object, and import the scale factor in step S4 into the "@images$slice1@scale.factors" member of the Seurat object.
[0085] Save the analysis results of Seurat in the mapping file, including:
[0086] When finding the instance of the spatial expression matrix corresponding to the Seurat object in the mapping file, append the "@meta.data" member of the Seurat object under the barcode item of this instance; append the "@assay$spatial@meta.features" member of the Seurat object under the feature item of the mapping file. Here, append means using the rhdf5::h5write() function of the "rhdf5" package to add the additional data generated by the analysis tool to an already initialized mapping file, including but not limited to: creating an additional array under the barcode item to save annotation information (idents) and cell type proportion information (CellTypes); creating an additional array under the feature item to save highly variable gene information (HVGs) and marker gene information (marker genes).
[0087] Saving the analysis results of Seurat in the mapping file also includes:
[0088] When the spatial expression matrix instance corresponding to the Seurat object is not found in the mapping file, or when there are changes in the dimensions, expression values, etc. of the data in the Seurat object, a new spatial expression matrix instance is created in the mapping file, and the modified "@assay$spatial@data" member of the Seurat object is imported into this instance, including data, indices, indptr, and shape; the "@meta.data" member of the Seurat object is imported into the Barcodes; and the "@assay$spatial@meta.features" member of the Seurat object is imported into the Features.
[0089] Further, on the R language platform, to implement the mutual conversion operation between the mapping file and the data structure in R language, it further includes:
[0090] When the user uses analysis tools such as BayesSpace, scater, SPCS, etc. for spatial transcriptome or single-cell transcriptome analysis, the content in the mapping file is converted into a SingleCellExperiment object, or the analysis results in the SingleCellExperiment are saved in the mapping file.
[0091] Specifically, converting the content in the mapping file into a SingleCellExperiment object includes:
[0092] First, perform steps S1 to S4 to generate a spatial expression matrix;
[0093] Then call the SingleCellExperiment() function to create a SingleCellExperiment object, import the expression matrix in step S2 into the "@assays$counts" member of the SingleCellExperiment object, import the barcode items in step S3 into the "@colData" member of the SingleCellExperiment object, and import the feature items in step S3 into the "@rowData" member of the SingleCellExperiment object; import the spatial coordinates of the barcodes in step S4 into the "@colData" member of the SingleCellExperiment object, and import the scaling factor in step S4 into the "@meta.data" member of the SingleCellExperiment object.
[0094] Converting the content in the mapping file into a SingleCellExperiment object further includes:
[0095] First, execute the steps of S1 to S3 to generate a single-cell expression matrix;
[0096] Then, call the SingleCellExperiment() function to create a SingleCellExperiment object, import the expression matrix in step S2 into the "@assays$counts" member of the SingleCellExperiment object, import the barcode items in step S3 into the "@colData" member of the SingleCellExperiment object, and import the feature items in step S3 into the "@rowData" member of the SingleCellExperiment object.
[0097] Save the analysis results in the SingleCellExperiment in a mapping file, including:
[0098] Save the analysis results of spatial transcriptomics, or save the analysis results of single-cell transcriptomics.
[0099] When finding the corresponding spatial expression matrix instance or single-cell expression matrix instance of the SingleCellExperiment object in the mapping file, append the "@colData" member of the SingleCellExperiment object under the barcode item of the expression matrix; append the "@rowData" member of the SingleCellExperiment object under the feature item of the expression matrix; append the "@meta.data" member of the SingleCellExperiment object in other unstructured information. When the corresponding expression matrix of the SingleCellExperiment object is not found in the mapping file, or when there are changes in the dimensions, expression values, etc. of the data in the SingleCellExperiment object, create a new expression matrix in the mapping file, and import the modified "@assays$counts" member of the SingleCellExperiment object into the expression matrix, including data, indices, indptr, and shape; import the "@colData" member of the SingleCellExperiment object into the barcode item (Barcodes); import the "@rowData" member of the SingleCellExperiment object into the feature item (Features).
[0100] Furthermore, on the R language platform, to implement the mutual conversion operation between the mapping file and the data structure in the R language, it further includes:
[0101] When users perform spatial transcriptome analysis using analysis tools such as SPOTlight, they can convert the content in the mapping file into a SpatialExperiment object, or save the analysis results in the SpatialExperiment in the mapping file.
[0102] Specifically, converting the content in the mapping file into a SpatialExperiment object includes:
[0103] First, execute steps S1 to S4 to generate an instance of the spatial expression matrix; then call the SpatialExperiment() function to create a SpatialExperiment object, import the expression matrix in step S2 into the "@assays$counts" member of the SpatialExperiment object, import the barcode items in step S3 into the "@colData" member of the SpatialExperiment object, and import the feature items in step S3 into the "@rowData" member of the SpatialExperiment object; then call the SpatialImage() function to create a spatial image container SpatialImage, import the spatial tissue images at various resolutions in step S4 into the SpatialImage container, and import this container into the "@imageData" member of the SpatialExperiment object; finally, import the spatial coordinates of the barcodes in step S4 into the "@spatialData" member of the SpatialExperiment object, and import the scale factor in step S4 into the "@meta.data" member of the SpatialExperiment object.
[0104] Saving the analysis results in the SpatialExperiment in the mapping file includes:
[0105] When the spatial expression matrix instance corresponding to the SpatialExperiment object is found in the mapping file, append the "@colData" member of the SpatialExperiment object to the barcode item of the expression matrix; append the "@rowData" member of the SpatialExperiment object to the feature item of the expression matrix; append the "@meta.data" member of the SpatialExperiment object to other unstructured information. When the expression matrix corresponding to the SpatialExperiment object is not found in the mapping file, or when there are changes in the dimensions, expression values, etc. of the data in the SpatialExperiment object, create a new expression matrix in the mapping file, and import the modified "@assays$counts" member of the SpatialExperiment object into this expression matrix, including data, indices, indptr, and shape; import the "@colData" member of the SpatialExperiment object into the barcode item (Barcodes); import the "@rowData" member of the SpatialExperiment object into the feature item (Features).
[0106] On the R language platform, the operation of converting between the mapping file and the data structure in R language also includes:
[0107] When users perform spatial transcriptome analysis using analysis tools such as RCTD and spacexr, the content in the mapping file can be converted into a SpatialRNA object.
[0108] Specifically, converting the content in the mapping file into a SpatialRNA object includes:
[0109] First, execute steps S1 to S4 to generate a spatial expression matrix instance;
[0110] Then call the SeuratObject::CreateAssayObject() function to create an Assay object, import the expression matrix in step S2 into the "$counts" member of the Assay object, and respectively use the barcodes in the barcode item in step S3 as the column names of the expression matrix, and use the gene names in the feature item in step S3 as the row names of the expression matrix; finally, call the SpatialRNA() function to create a SpatialRNA object, import the Assay object created in the previous step into the "$counts" member, and import the spatial coordinates of the barcodes in step S4 into the "$coord" member of the SpatialRNA object.
[0111] On the R language platform, the operation of realizing the mutual conversion between the mapping file and the data structure in the R language further includes:
[0112] When the user performs spatial transcriptome analysis using the SPARK analysis tool, the content in the mapping file can be converted into a SPARK object, and the analysis results of SPARK can also be saved in the mapping file.
[0113] Specifically, converting the content in the mapping file into a SPARK object includes:
[0114] First, perform steps S1 to S4 to generate a spatial expression matrix;
[0115] Then call the SPARK::CreateSPARKObject() function, import the expression matrix in step S2 into the "$counts" member of the SPARK object, use the barcodes of the barcode items in step S3 as the column names of the expression matrix respectively, use the gene names of the feature items in step S3 as the row names of the expression matrix; finally, import the spatial coordinates of the barcodes in step S4 into the "$location" member of the SPARK object.
[0116] Saving the analysis results of SPARK in the mapping file includes:
[0117] When finding the instance of the spatial expression matrix corresponding to the SPARK object in the mapping file, append the "$stats" and "$res_mtest" members of the SpatialExperiment object under the feature item of the expression matrix. When the expression matrix corresponding to the SpatialExperiment object is not found in the mapping file, or when there are changes in the dimensions, expression values, etc. of the data in the SpatialExperiment object, create a new expression matrix in the mapping file, and import the modified "$counts" member of the SPARK object into the expression matrix, including data, indices, indptr, and shape.
[0118] On the R language platform, the operation of realizing the mutual conversion between the mapping file and the data structure in the R language further includes:
[0119] When the user performs spatial transcriptome analysis using the Giotto analysis tool, the content in the mapping file can be converted into a Giotto object, and the analysis results of Giotto can also be saved in the mapping file.
[0120] Specifically, converting the content in the mapping file into a Giotto object includes:
[0121] First, execute steps S1 to S4 to generate a spatial expression matrix;
[0122] Then, call the CreateGiottoObject() function to create a Giotto object, import the expression matrix in step S2 into the "$raw_exprs" member of the Giotto object, and use the barcodes of the barcode items in step S3 as the column names of the expression matrix and the gene names of the feature items in step S3 as the row names of the expression matrix respectively; finally, import the spatial coordinates of the barcodes in step S4 into the "$spatial_locs" member of the Giotto object.
[0123] Save the analysis results of Giotto in a mapping file, including:
[0124] When finding the instance of the spatial expression matrix corresponding to the Giotto object in the mapping file, append the "@cell_metadata" member of the Giotto object under the barcode item of the expression matrix; append the "@gene_metadata" member of the Giotto object under the feature item of the expression matrix; append the "@meta.data" member of the Giotto object in other unstructured information. When the expression matrix corresponding to the SpatialExperiment object is not found in the mapping file, or when there are changes in the dimensions, expression values, etc. of the data in the Giotto object, create a new expression matrix in the mapping file and import the modified "$raw_exprs" member of the Giotto object into this expression matrix, including data, indices, indptr, and shape; import the "@colData" member of the Giotto object into the barcode items (Barcodes); import the "@rowData" member of the Giotto object into the feature items (Features).
[0125] Furthermore, on the Python platform, implement the mutual conversion operation between the mapping file and the data structure in Python, specifically including:
[0126] Use the "h5py" library to import the data in the mapping file. Among them, the "h5py" package is a third-party library in the Python language that contains a series of functions for reading and writing h5 files. Import the data in the mapping file, specifically:
[0127] S11, use the h5py.File() function to import the mapping file into the Python language environment;
[0128] S12. Read the column - stored sparse matrix of the sptmatrix instance in the mapping file, including data, indices, indptr, and shape. Finally, use the scipy.csc_matrix() function in the Python language to generate a spatial expression matrix or a single - cell expression matrix;
[0129] S13. Read the barcode items (Barcodes) of the sptmatrix instance in the mapping file, including information such as barcodes, annotations, cell types, etc., and the feature items (Features) of the sptmatrix instance in the mapping file, including information such as gene names, gene numbers, and whether they are marker genes;
[0130] S14. Read the histological image information of the sptimages instance in the mapping file, including histological images at various resolutions, spatial coordinates of barcodes, and scale factors.
[0131] Furthermore, on the Python platform, the operation of mutual conversion between the mapping file and the data structure in Python is also included:
[0132] When users perform spatial transcriptome analysis using analysis tools such as scanpy, SpaGCN, and SpatialDE, the content in the mapping file can be converted into an AnnData object, and the analysis results in the AnnData object can also be saved in the mapping file.
[0133] Specifically, as Figure 4 shown, converting the content in the mapping file into an AnnData object includes:
[0134] First, execute the steps of S11 to S14, then call the AnnData() function to create an AnnData object, import the spatial expression matrix in step S12 into the ".X" member of the AnnData object, import the barcode items in step S13 into the ".obs" member of the AnnData object, import the feature items in step S13 into the ".var" member of the AnnData object, import the histological images at various resolutions in step S14 into the ".uns['spatial']['images']" member of the AnnData object, import the spatial coordinates of barcodes in step S14 into the ".obsm['spatial']" member of the AnnData object, and import the scale factors in step S14 into the ".uns['spatial']['scalefactors']" member of the AnnData object.
[0135] Save the analysis results in the AnnData object to a mapping file, including:
[0136] When finding the instance of the spatial expression matrix corresponding to the AnnData object in the mapping file, append the ".obs" member of the AnnData object under the barcode item of this instance; append the ".obs" member of the AnnData object under the feature item of the mapping file. Here, appending means using the create_dataset() function of the "h5py" library to add the additional data generated by the analysis tool to an already initialized mapping file, including but not limited to: creating an additional array under the barcode item to save annotation information (idents), cell type proportion information (CellTypes); creating an additional array under the feature item to save highly variable gene information (HVGs), marker gene information (marker genes).
[0137] Saving the analysis results in the AnnData object to the mapping file also includes:
[0138] When the instance of the spatial expression matrix corresponding to the AnnData object is not found in the mapping file, or when there are changes in the dimensions, expression values, etc. of the data in the AnnData object, create a new instance of the spatial expression matrix in the mapping file, and import the modified ".X" member of the AnnData object into this instance, including data, indices, indptr, and shape; import the ".obs" member of the AnnData object into the barcode item (Barcodes); import the ".var" member of the AnnData object into the feature item (Features).
[0139] Alternatively, when the user performs spatial transcriptome analysis using the stLearn analysis tool, the content in the mapping file can be converted into an stData object, or the analysis results of the stData can be saved in the mapping file.
[0140] Specifically, converting the content in the mapping file into an stData object includes:
[0141] The steps performed, except for importing the spatial coordinates of the barcodes in step S14 into the ".obs" member of the AnnData object, are the same as those of the Load_spt_to_AnnData() function.
[0142] Saving the analysis results of the stData in the mapping file includes:
[0143] The steps performed are the same as those of the Save_spt_from_AnnData() function.
[0144] Alternatively, when the user needs to view the content in the mapping file on the Python platform, the content in the mapping file can be converted into a dictionary format in Python.
[0145] Converting the content in the mapping file into a dictionary format in Python specifically includes:
[0146] Execute steps S11 to S14, and save the expression matrix, barcode items, feature items, and tissue image information of the above four steps in the dictionary object in the form of key-value pairs of {name: data}.
[0147] Figure 3 A storage example of the sequencing data of the axial section of the mouse brain tissue obtained by the 10X Visium sequencing technology is given. Among them, tissue section 1... tissue section n represent the tissue image data of multiple sections, belonging to the sptimages type. Raw matrix 1... raw matrix n represent the raw gene expression matrices of multiple sections, and preprocessing matrix 1... preprocessing matrix n represent the spatial transcriptome expression matrices obtained after a series of preprocessing methods on the corresponding expression matrices; reference matrix 1... reference matrix n represent the single-cell transcriptome expression matrices obtained by the single-cell sequencing technology scRNA-seq, and the above expression matrices all belong to the expression matrix type.
[0148] A method for exchanging spatial transcriptome data based on the mapping between R and Python. The solution provides a well-organized file storage format: the "spt" format. For the intermediate results obtained by running on one platform (such as the R language) in the workflow, the interface on "sptr" can be called to store them in the "spt" format file, and then the interface on "sptpy" can be called under another platform (such as Python) to read the intermediate results and continue running, and vice versa. At the same time, due to the variety of sequencing technologies for spatial transcriptomics, such as 10X Visium, Slide-Seq, etc., the present disclosure also provides interfaces for reading and saving these data, so that the method developed based on the "spt" file format can be compatible with multiple sequencing technologies, which helps to improve the efficiency and performance of developing spatial transcriptome tools.
[0149] The described solution can achieve the conversion of different data structures on the Python platform and the R platform, enabling the intermediate results generated on one platform to be read and continued to run on the other platform; it can be compatible with multiple read and write interfaces for original sequencing data, that is, it can read "spt" files in the same way as reading original "h5" format files; it provides a unified file format for the original data generated by various sequencing technologies, which can save storage space while improving the development efficiency of downstream analysis tools; in addition, the data structure of the described solution can also serve as the main data read and write hub for the spatial transcriptome data analysis workflow, improving the simplicity, usability, and scalability of the workflow.
[0150] As Figure 2 shown, this technology creates a mapping file in the "spt" format. In this file, there are tissue image classes and expression matrix classes. When the user provides the paths of spatial transcriptome or single-cell transcriptome data, this module uses computer instructions to instantiate the tissue image classes and expression matrix classes, import the tissue image data and expression matrix data, and save them in the mapping file.
[0151] This technology provides an R language toolkit "sptr" implemented by computer instructions as the conversion interface between the mapping file and the R language data structure. When the user needs to create R language spatial transcriptome objects, including but not limited to: Seurat, SingleCellExperiment, SpatialExperiment, functions such as Load_spt_to_Seurat(), Load_spt_to_SCE(), Load_spt_to_SE() in this module can be called to achieve data reading. When the user needs to save the intermediate analysis results in R language, functions such as Save_spt_from_Seurat(), Save_spt_from_SCE(), Save_spt_from_SE() in this module can be called to achieve data storage.
[0152] The present technology provides a Python language toolkit "sptpy" implemented by computer instructions, which serves as an interface for converting mapping files and Python language data structures. When a user needs to create a spatial transcriptomics object in the Python language, including but not limited to: AnnData, stData, Dictionary, functions such as Load_spt_to_AnnData(), Load_spt_to_stData(), Load_spt_to_Dict() in this module can be called to implement data reading. When a user needs to save the intermediate analysis results in the Python language, functions such as Save_spt_from_AnnData() and Save_spt_from_stData() in this module can be called to implement data storage.
[0153] Example Two
[0154] This embodiment provides a spatial transcriptomics data conversion system across language platforms;
[0155] The spatial transcriptomics data conversion system across language platforms includes:
[0156] A reading and storage module, which is configured to: read and store the spatial transcriptomics data, single-cell transcriptomics reference data, or intermediate results generated by spatial transcriptomics analysis tools on the first language platform using a mapping file;
[0157] An operation module, which is configured to: read the stored results on the second language platform and continue to operate;
[0158] Wherein, the reading and storage using the mapping file specifically includes: initializing the mapping file;
[0159] When the first language platform is the R language platform; then on the R language platform, implement the mutual conversion operation between the mapping file and the data structure in the R language;
[0160] When the first language platform is the Python language platform; then on the Python platform, implement the mutual conversion operation between the mapping file and the data structure in Python.
[0161] It should be noted here that the above reading and storage module and operation module correspond to steps S101 to S102 in Embodiment One. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment One above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.
[0162] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for converting spatial transcriptome data across language platforms, characterized in that Including: Read and store the spatial transcriptome data of the first language platform, the single-cell transcriptome reference data, or the intermediate results generated by the spatial transcriptome analysis tool using a mapping file; The basic structure of the mapping file consists of two classes: the spatial image class and the expression matrix class; Among them, the spatial image class is used to store spatial information; The expression matrix class is used to store the expression matrix and the intermediate results generated by the spatial transcriptome analysis tool; On the second language platform, read the stored results and continue to run; Among them, the reading and storage using the mapping file specifically includes: initializing the mapping file; When the first language platform is the R language platform, then on the R language platform, implement the mutual conversion operation between the mapping file and the data structure in R; the implementation of the mutual conversion operation between the mapping file and the data structure in R on the R language platform specifically includes: Use the "rhdf5" package to import the data in the mapping file; the import of the data in the mapping file is specifically: S1, Use the rhdf5::h5read() function to import the mapping file into the R language environment; S2, Read the column-stored sparse matrix of the sptmatrix instance in the mapping file, including data, index, pointer, and dimension; finally, use the Matrix::sparseMatrix() function in R to generate a spatial expression matrix or a single-cell expression matrix; S3, Read the barcode items of the sptmatrix instance in the mapping file, including barcodes, annotations, cell type information, and the feature items of the sptmatrix instance, including gene names, gene numbers, and whether they are marker gene information; S4, Read the histological image information of the sptimages instance in the mapping file, including histological images at various resolutions, the spatial coordinates of the barcodes, and the scale factors; When the first language platform is the Python language platform, then on the Python platform, implement the mutual conversion operation between the mapping file and the data structure in Python; the implementation of the mutual conversion operation between the mapping file and the data structure in Python on the Python platform is specifically: Use the "h5py" library to import the data in the mapping file; S11, Use the h5py.File() function to import the mapping file into the Python language environment; S12, Read the column-stored sparse matrix of the sptmatrix instance in the mapping file, including data, index, pointer, and dimension; finally, use the scipy.csc_matrix() function in Python to generate a spatial expression matrix or a single-cell expression matrix; S13, Read the barcode items of the sptmatrix instance in the mapping file, including barcodes, annotations, cell type information, and the feature items of the sptmatrix instance in the mapping file, including gene names, gene numbers, and whether they are marker gene information; S14, Read the histological image information of the sptimages instance in the mapping file, including histological images at various resolutions, the spatial coordinates of the barcodes, and the scale factors.
2. The method for converting spatial transcriptome data of a cross-language platform according to claim 1, characterized in that, The initialization mapping file specifically includes: Using the rhdf5 package to create the structure of the mapping file and import the spatial transcriptome data into the mapping file; Among them, the use of the rhdf5 package to create the structure of the mapping file specifically includes: Calling the rhdf5::h5createGroup() function to create the basic structure of the mapping file for storing the expression matrix and tissue images.
3. The method for converting spatial transcriptome data of a cross-language platform according to claim 2, characterized in that, The basic structure of the mapping file, where the spatial image class is used to store spatial information, that is, histological images at different resolutions, the specific coordinates of each sampling point in the tissue section image, and the scaling factor between the original high-resolution image and the low-resolution image; Among them, the expression matrix is a sparse matrix in column storage format; the sparse matrix includes: data, used to store the expression values of matrix elements; index, used to store the row numbers of elements; pointer, used to represent the starting position of each row of elements; dimension, used to record the dimension of the sparse matrix; name, used to record the name of the sptmatrix instance; the intermediate results with the same dimension as the barcode and feature are respectively appended under the barcode item and the feature item, and other data are appended in the unstructured information.
4. The spatial transcriptome data conversion method for a cross-lingual platform according to claim 1, characterized in that, On the R language platform, the implementation of the mutual conversion operation between the mapping file and the data structure in R language also includes: When the user uses the Seurat spatial transcriptome analysis tool, convert the content in the mapping file into a Seurat object, or save the analysis results of Seurat in the mapping file; When the user uses the BayesSpace, scater, SPCS analysis tools for spatial transcriptome or single-cell transcriptome analysis, convert the content in the mapping file into a SingleCellExperiment object, or save the analysis results in the SingleCellExperiment in the mapping file.
5. The method for converting spatial transcriptome data of a cross - language platform according to claim 1, wherein, On the R language platform, the implementation of the mutual conversion operation between the mapping file and the data structure in R language also includes: When the user uses the SPOTlight analysis tool for spatial transcriptome analysis, convert the content in the mapping file into a SpatialExperiment object, or save the analysis results in the SpatialExperiment in the mapping file; When the user uses the RCTD, spacexr analysis tools for spatial transcriptome analysis, convert the content in the mapping file into a SpatialRNA object.
6. The method for converting spatial transcriptome data of a cross - language platform according to claim 1, characterized in that, On the R language platform, the implementation of the mutual conversion operation between the mapping file and the data structure in R language also includes: When the user uses the SPARK analysis tool for spatial transcriptome analysis, convert the content in the mapping file into a SPARK object, or save the analysis results of SPARK in the mapping file; When the user uses the Giotto analysis tool for spatial transcriptome analysis, convert the content in the mapping file into a Giotto object, or save the analysis results of Giotto in the mapping file.
7. The method for converting spatial transcriptome data of a cross-language platform according to claim 1, wherein On the Python platform, implementing the conversion operations between the mapped file and the data structures in Python also includes: When the user performs spatial transcriptome analysis using analysis tools such as scanpy, SpaGCN, and SpatialDE, converting the content in the mapped file into an AnnData object, or saving the analysis results in the AnnData object in the mapped file; When the user performs spatial transcriptome analysis using the stLearn analysis tool, converting the content in the mapped file into an stData object, or saving the analysis results of stData in the mapped file; When the user needs to view the content in the mapped file on the Python platform, converting the content in the mapped file into a dictionary format in Python.
8. A spatial transcriptome data conversion system for cross - language platforms, characterized in that, Including: A read and storage module, which is configured to: read and store the spatial transcriptome data, single-cell transcriptome reference data, or intermediate results generated by spatial transcriptome analysis tools on the first language platform using a mapped file; the basic structure of the mapped file consists of two classes: a spatial image class and an expression matrix class; Among them, the spatial image class is used to store spatial information; The expression matrix class is used to store the expression matrix and the intermediate results generated by spatial transcriptome analysis tools; A running module, which is configured to: on the second language platform, read the stored results and continue running; Among them, the reading and storing using the mapped file specifically includes: initializing the mapped file; When the first language platform is the R language platform; then on the R language platform, implementing the conversion operations between the mapped file and the data structures in the R language; the implementation of the conversion operations between the mapped file and the data structures in the R language on the R language platform specifically includes: Importing the data in the mapped file using the "rhdf5" package; the import of the data in the mapped file is specifically: S1, using the rhdf5::h5read() function to import the mapped file into the R language environment; S2, reading the column-stored sparse matrix of the sptmatrix instance in the mapped file, including data, index, pointer, and dimension; finally, using the Matrix::sparseMatrix() function in the R language to generate a spatial expression matrix or a single-cell expression matrix; S3, reading the barcode items of the sptmatrix instance in the mapped file, including barcodes, annotations, cell type information, and the feature items of the sptmatrix instance, including gene names, gene numbers, and information on whether it is a marker gene; S4, reading the histological image information of the sptimages instance in the mapped file, including histological images at various resolutions, spatial coordinates of barcodes, and scale factors; When the first language platform is the Python language platform; then on the Python platform, implementing the conversion operations between the mapped file and the data structures in Python; the implementation of the conversion operations between the mapped file and the data structures in Python on the Python platform is specifically: Importing the data in the mapped file using the "h5py" library; S11. Use the h5py.File() function to import the mapping file into the Python language environment; S12. Read the column-stored sparse matrix of the sptmatrix instance in the mapping file, including data, indices, pointers, and dimensions; finally, use the scipy.csc_matrix() function in the Python language to generate a spatial expression matrix or a single-cell expression matrix; S13. Read the barcode items of the sptmatrix instance in the mapping file, including barcodes, annotations, and cell type information, as well as the feature items of the sptmatrix instance in the mapping file, including gene names, gene numbers, and information on whether they are marker genes; S14. Read the histological image information of the sptimages instance in the mapping file, including histological images at various resolutions, the spatial coordinates of barcodes, and the scale factors.
Citation Information
Patent Citations
Transcriptome data automatic analysis method based on gene chip
CN112289380A
System, method and software for robust transcriptomic data analysis
US20170277826A1