Spatial omics data processing method and system and computing equipment

By converting spatial omics text data into gene name files, information files and binary files and compressed storage, the problems of data redundancy and low storage efficiency in the prior art are solved, and the effect of saving storage space and improving analysis efficiency is achieved.

CN120048363APending Publication Date: 2025-05-27SHENZHEN HUADA SANJIAN QIFA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311593859.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The storage format of existing spatial omics data results in data redundancy, consumes a lot of storage space, and affects the storage and transmission efficiency of data.

Method used

By converting spatial omics text data into gene name files, information files and binary files and performing compression storage, redundant information is reduced and storage space is saved.

Benefits of technology

It effectively saves storage space, improves data reading and analysis efficiency, and facilitates later spatial omics analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048363A_ABST
    Figure CN120048363A_ABST
Patent Text Reader

Abstract

The invention discloses a spatial omics data processing method and system and computing equipment. The spatial omics data processing method comprises the steps of obtaining spatial omics text data; obtaining a gene name file and an information file according to the spatial omics text data, and obtaining a binary file according to the gene name file and the spatial omics text data; and compressing and storing the gene name file, the information file and the binary file. According to the embodiment of the invention, the storage space can be effectively saved, and reading and analysis of space omics in the later period are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electronic products, and in particular to a method, system and computing device for processing spatial omics data. Background Art

[0002] The original spatial group data provides simple spatial group information, including coordinate information, gene name, gene expression level, etc. For the reasonable use of spatial group data, it is necessary to cooperate with ssDNA staining images and add additional information such as cell staining segmentation, boundary division, and functional area segmentation. The common spatial omics data storage format is spatial omics text data (i.e., RNA long matrix) with gene + coordinates as the main and other information as the auxiliary. Such spatial omics text data will generate a large amount of redundant data, and the data storage consumption is large, usually in the range of tens of GB to hundreds of GB, which affects data storage and transmission. Moreover, the spatial omics text data is read once when used, and a lot of redundant information is not used during analysis, which wastes storage space and reading time, is not convenient for subsequent spatial omics analysis, and affects data storage and transmission. In addition, such a long matrix needs to read the spatial omics text data once in spatial omics analysis, and a lot of redundant information is not used during analysis, which wastes storage space. Summary of the invention

[0003] In view of the above defects or deficiencies in the prior art, it is desirable to provide a method, system and computing device for processing spatial omics data, which can effectively save storage space and facilitate the subsequent reading and analysis of spatial omics.

[0004] In a first aspect, an embodiment of the present application provides a method for processing spatial omics data, comprising:

[0005] Obtain spatial omics text data;

[0006] Obtaining a gene name file and an information file according to the spatial omics text data, and obtaining a binary file according to the gene name file and the spatial omics text data;

[0007] The gene name file, information file and binary file are compressed and stored.

[0008] In some examples, obtaining spatial omics text data includes:

[0009] Obtaining spatial omics raw data, and generating spatial omics text data according to the spatial omics raw data, including:

[0010] Extract coordinate information, gene expression and gene information from raw spatial omics data;

[0011] Obtain additional information, where the additional information includes ssDNA staining information;

[0012] Generate the spatial omics text data according to the coordinate information, gene expression levels, gene information, and the additional information.

[0013] In some examples, obtaining the gene name file and the information file according to the spatial omics text data includes:

[0014] Read the gene names from the spatial omics text data, and obtain the gene name file according to the gene names and the subscripts of the gene names.

[0015] Extract the file information of the spatial omics text data from the spatial omics text data, and form the information file according to the file information of the spatial omics text data, where the information file includes at least the number of rows, the number of columns, and the header information of the spatial omics text data.

[0016] In some examples, obtaining the binary file according to the gene name file and the spatial omics text data includes:

[0017] Read the gene names in the spatial omics text data line by line, and convert the read gene names into corresponding subscripts according to the subscripts corresponding to the gene names in the gene name file.

[0018] Form the binary file according to the converted subscripts and the gene coordinates in the spatial omics text data.

[0019] In some examples, obtaining the binary file according to the gene name file and the spatial omics text data includes:

[0020] Divide the spatial omics text data into multiple spatial omics text data blocks;

[0021] Perform parallel processing on multiple spatial omics text data blocks, where the processing process for each spatial omics text data block is: read the gene names in the spatial omics text data block line by line, convert the read gene names into corresponding subscripts according to the subscripts corresponding to the gene names in the gene name file, and form a binary file block according to the converted subscripts and the gene coordinates in the spatial omics text data block;

[0022] Form the binary file according to the binary file blocks obtained from each spatial omics text data block.

[0023] In some examples, it further includes:

[0024] Read the compressed file;

[0025] Unzip the compressed file to obtain the gene name file, binary file, and information file;

[0026] Convert the binary file into the spatial omics text data according to the gene name file and the information file.

[0027] In some examples

[0028] The converting the binary file into the spatial omics text data according to the gene name file and the information file includes:

[0029] Read the information file to obtain the header information and the table information;

[0030] According to the number of elements in each row of the information file, sequentially read the corresponding bytes in the binary file;

[0031] Write the read table information and bytes into the spatial omics text data; or,

[0032] Read the information file to obtain the header information and the table information;

[0033] Obtain multiple row number sub-intervals according to the information file;

[0034] Perform parallel processing on the binary file based on multiple threads, with one thread corresponding to at least one row number sub-interval. Among them, the processing process of each thread includes: according to the row number sub-interval corresponding to the thread, sequentially read the corresponding bytes in the binary file, and write the read table information and bytes into the spatial omics text data block;

[0035] Obtain the spatial omics text data according to each spatial omics text data block.

[0036] In some examples, after reading the spatial omics text data, it further includes: converting the spatial omics text data into a data dictionary of a predetermined data structure.

[0037] In a second aspect, an embodiment of the present application provides a processing system for spatial omics data, including:

[0038] An acquisition module for obtaining spatial omics text data;

[0039] A partitioning module for obtaining a gene name file and an information file according to the spatial omics text data, and obtaining a binary file according to the gene name file and the spatial omics text data;

[0040] A compression storage module for compressing and storing the gene name file, the information file, and the binary file.

[0041] In a third aspect, an embodiment of the present application provides a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for processing spatial omics data as described in the first aspect above.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to implement the method for processing spatial omics data as described in the first aspect above.

[0043] In a fifth aspect, an embodiment of the present application provides a computer program product, on which a computer program is stored, and the computer program is used to implement the method for processing spatial omics data as described in the first aspect above.

[0044] The method, system, and computing device for processing spatial omics data provided by the embodiments of the present application obtain a gene name file, a binary file, and an information file from the spatial omics text data, and then compress and store the gene name file, the binary file, and the information file. Thus, the storage space can be effectively saved, which is convenient for the later reading and analysis of spatial omics. That is: the spatial omics information is read from the spatial omics original data bam file to generate spatial omics text data (such as an RNA long matrix). When the data is not in use, the spatial omics text data is converted into a data format mainly based on coordinate information and supplemented by other information for storage. Thus, the amount of data stored is reduced, and the storage space is saved. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more obvious:

[0046] Figure 1 It is a flowchart of the method for processing spatial omics data according to an embodiment of the present application;

[0047] Figure 2 It is a schematic diagram of the spatial omics text data of spatial omics data;

[0048] Figure 3 It is an architecture diagram of the method for processing spatial omics data according to an embodiment of the present application;

[0049] Figure 4 It is a schematic diagram of the structure of data compression storage;

[0050] Figure 5 It is a schematic diagram of the implementation process in the specific application of the method for processing spatial omics data;

[0051] Figure 6 It is a schematic diagram of a data recovery process;

[0052] Figure 7 Schematic diagram of the data dictionary storage method after data recovery;

[0053] Figure 8 Schematic diagram of the input data format;

[0054] Figure 9 Diagram showing the result after a data compression command;

[0055] Figure 10 Schematic diagram of the input format for a data recovery;

[0056] Figure 11 Diagram showing the result of a data recovery;

[0057] Figure 12 Schematic diagram of a data compression input;

[0058] Figure 13 Schematic diagram showing the result of a data compression;

[0059] Figure 14 Schematic diagram of an input for a data recovery;

[0060] Figure 15 Diagram showing the result of another data recovery;

[0061] Figure 16 Schematic diagram of the structure of the spatial omics data processing system according to an embodiment of the present application;

[0062] Figure 17 Schematic diagram of the structure of the computing device according to an embodiment of the present application. Detailed implementation manners

[0063] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant disclosure, rather than limiting the disclosure. Additionally, it should be noted that for the sake of description, only the parts related to the disclosure are shown in the drawings.

[0064] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0065] The spatial omics data processing method, system and computing device according to embodiments of the present application are described below in conjunction with the accompanying drawings.

[0066] The implementation environment of the application example can obtain spatial omics text data from personal computing devices such as computers and mobile terminals; obtain a gene name file and an information file according to the spatial omics text data, and obtain a binary file according to the gene name file and the spatial omics text data; compress and store the gene name file, the information file, and the binary file.

[0067] Alternatively, it can also be implemented by a server. For example, a personal computing device sends a request to the server, and the server obtains spatial omics text data; obtains a gene name file and an information file according to the spatial omics text data, and obtains a binary file according to the gene name file and the spatial omics text data; compresses and stores the gene name file, the information file, and the binary file, and finally returns the stored file to the personal computing device.

[0068] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0069] Figure 1 is a flowchart of a method for processing spatial omics data according to an embodiment of the present application. As Figure 1 shown, according to a method for processing spatial omics data according to an embodiment of the present application, the following steps are included:

[0070] S101: Obtain spatial omics text data.

[0071] In a specific example, the spatial omics text data is, for example, an RNA long matrix.

[0072] In a specific example, the acquisition of spatial omics text data can be achieved in the following manner:

[0073] Obtain spatial omics raw data, and generate spatial omics text data according to the spatial omics raw data. Specifically: extract coordinate information, gene expression levels, and gene information from the spatial omics raw data; obtain additional information, where the additional information includes ssDNA staining information; generate the spatial omics text data according to the coordinate information, gene expression levels, gene information, and the additional information.

[0074] Taking the single-cell spatial omics raw data as an example, for instance, obtaining the single-cell spatial omics raw data bam file (i.e., the spatial omics raw data), and then, based on the single-cell spatial omics raw data bam file, obtaining the spatial omics text data (i.e., the RNA long matrix). Specifically, extracting the effective information from the single-cell spatial omics raw data bam file, and the effective information includes but is not limited to coordinate information, gene expression levels, gene names, etc., and combining with additional information such as ssDNA staining information like cell boundary division information, functional area division information, etc., to form a long matrix with RNA as the unit, that is, the spatial omics text data. As Figure 2 shown, it is a schematic diagram of a long matrix with RNA as the unit generated from single-cell spatial omics raw data.

[0075] As Figure 3 shown, the single-cell spatial omics raw data bam file is converted into text data, that is, the spatial omics text data.

[0076] S102: Obtaining a gene name file and an information file based on the spatial omics text data, and obtaining a binary file based on the gene name file and the spatial omics text data.

[0077] Among them, in the embodiments of the present application, in order to make the spatial omics text data have higher sparsity, that is: the data storage is large and there is more redundant information, which is not conducive to spatial omics analysis. Therefore, the embodiments of the present application perform sparsification processing on it, and the specific method is to divide the file and compress and store it, so as to reduce its storage. For example: dividing the spatial omics text data into three different files, namely the gene name file, the binary file, and the information file.

[0078] In a specific example, obtaining a gene name file and an information file based on the spatial omics text data includes: reading the gene names from the spatial omics text data, and obtaining the gene name file according to the gene names and the subscripts of the gene names; extracting the file information of the spatial omics text data from the spatial omics text data, and forming the information file according to the file information of the spatial omics text data, where the information file at least includes the number of rows, the number of columns, and the header information of the spatial omics text data, etc.

[0079] Furthermore, obtaining a binary file based on the gene name file and the spatial omics text data includes two methods, one is the single-thread method, and the other is the multi-thread method, that is: according to different requirements, converting the spatial omics text data into a binary file includes the single-thread method and the multi-thread method, where the single-thread method includes:

[0080] Obtaining a binary file based on the gene name file and the spatial omics text data includes:

[0081] Read the gene names in the spatial omics text data line by line, and convert the read gene names into corresponding subscripts according to the subscripts corresponding to the gene names in the gene name file; form the binary file according to the converted subscripts and the gene coordinates in the spatial omics text data. Specifically, read the spatial omics text data line by line, where the gene names are converted into the subscripts in their corresponding gene name files, with the coordinate information as the main and other information such as the gene expression levels at the current coordinates as the supplement, and convert them into binary and write them into the binary file.

[0082] The multi-threaded method includes: obtaining a binary file according to the gene name file and the spatial omics text data, including: dividing the spatial omics text data into multiple spatial omics text data blocks; performing parallel processing on the multiple spatial omics text data blocks, where the processing process for each spatial omics text data block is: read the gene names in the spatial omics text data block line by line, and convert the read gene names into corresponding subscripts according to the subscripts corresponding to the gene names in the gene name file, and form a binary file block according to the converted subscripts and the gene coordinates in the spatial omics text data block; form the binary file according to the binary file blocks obtained from each spatial omics text data block. Specifically, number each thread, divide the spatial omics text data into several parts according to the number of input threads, each thread is responsible for a small part of the spatial omics text data. For each thread, a single thread reads a small part of the spatial omics text data line by line, with the coordinate information as the main and the gene names converted into the subscripts in their gene name files, count the file information of the spatial omics text data, such as the number of rows and columns, etc., and write this information into the long matrix information text successively, convert the coordinate information and gene expression information into binary and store them in the binary file generated by the current thread, name and save it to the specified path according to the current thread label. Read the binary files generated by each thread and store them in the final binary file in sequence according to their numbers, and delete the binary files generated by each thread under the specified path.

[0083] As Figure 3 shown, divide the text data, i.e., the spatial omics text data, into a byte file (information file), a gene name file (gene name file), and a binary file.

[0084] S103: Compress and store the gene name file, information file, and binary file.

[0085] Among them, the gene name file, binary file, and information file are associated and can be stored in one file, and then this folder is compressed and stored uniformly. For example: save the gene name file, binary file, and information file to a specified folder and save it in zip format to a specified path.

[0086] According to the method for processing spatial omics data according to the embodiments of the present application, a gene name file, a binary file, and an information file are obtained from the spatial omics text data, and then the gene name file, information file, and binary file are compressed and stored. Thus, the storage space can be effectively saved, which is convenient for later reading and analysis of spatial omics. That is: read the spatial group information from the spatial omics raw data bam file to generate spatial omics text data (such as RNA long matrix), and when the data is not in use, convert the spatial omics text data into a data format mainly based on coordinate information and supplemented by other information for storage. Thus, the amount of data stored is reduced, and the storage space is saved.

[0087] In an embodiment of the present application, the method for processing spatial omics data further includes: reading the compressed file; decompressing the compressed file to obtain the gene name file, binary file, and information file; and converting the binary file into the spatial omics text data according to the gene name file and information file.

[0088] Specifically, converting the binary file into the spatial omics text data according to the gene name file and information file includes: reading the information file to obtain the header information and table information; sequentially reading the corresponding bytes in the binary file according to the number of elements in each row of the information file; writing the read table information and bytes into the spatial omics text data; or, reading the information file to obtain the header information and table information; obtaining multiple row number sub-intervals according to the information file; performing parallel processing on the binary file based on multiple threads, where one thread corresponds to at least one row number sub-interval, and the processing process of each thread includes: sequentially reading the corresponding bytes in the binary file according to the row number sub-interval corresponding to the thread, and writing the read table information and bytes into the spatial omics text data block; and obtaining the spatial omics text data according to each spatial omics text data block. Finally, after reading the spatial omics text data, convert the spatial omics text data into a data dictionary of a predetermined data structure.

[0089] Combined with Figure 3 As shown, the decompressed file can be converted into a long matrix according to different reading methods. For example:

[0090] Single-threaded mode:

[0091] Read the information file to obtain the long matrix header information, number of rows, and number of columns, and write the header information into the spatial omics text data. According to the number of columns, sequentially read the corresponding number of bytes from the binary file and convert them into strings, then write them into the spatial omics text data.

[0092] Multi-threaded approach:

[0093] Read the information file to obtain the long matrix header information, number of rows, and gene expression information at each coordinate, and write the header information into the spatial omics text data. Assign a number to each thread, divide the number of rows into several parts according to the input number of threads, and each thread is responsible for a small part of the number of rows. For each thread, the single thread reads the binary byte file according to the rows it is responsible for, mainly based on the coordinate information, converts the gene number into the gene name according to the gene name file, reads the corresponding number of bytes according to the gene expression information at the current coordinate, converts them into strings, stores them in the RNA file generated by the current thread, names and saves them to the specified path according to the current thread number. Read the RNA files generated by each thread and store them sequentially into the final spatial omics text data according to their numbers, and delete the spatial omics text data generated by each thread under the specified path. Check whether the newly generated spatial omics text data is consistent with the original long matrix data according to the number of rows, columns, and header information in the information file. If they are consistent, proceed to the next step; if not, report an error and terminate the program.

[0094] Combine Figure 3 As shown, read the newly generated spatial omics text data, which can be converted into STOCdata for subsequent spatial omics analysis. Read the header information, number of rows, number of columns, total number of strings, etc. and store them as header information in dictionary form. Read the coordinate information, gene names, gene expression levels, etc. and store them in pandas format as the main data information MDATA. Read the region information, boundary information, etc. and store them in pandas format as other data information ODATA. Read the gene names and store them in pandas format as gene name information geneInfo. In this way, when the data is used, read the data and convert it into the specified data structure STOCdata for downstream spatial omics analysis. This data structure can reduce data redundancy during spatial omics analysis, that is, only provide the necessary information, thereby improving the data utilization rate and reading efficiency.

[0095] In the specific application of the spatial omics data processing method in the embodiments of the present application, its main purpose of storage is to read the spatial omics text data from the bam file, convert the long matrix into a data format mainly based on coordinate information and supplemented by other information, including gene list information files, common information binary files, and information files, and store them in zip format, such as Figure 4As shown, it is the structure of data compression storage.

[0096] The implementation process in the specific application of the method for processing spatial omics data in the embodiments of this application is as Figure 5 shown. In the data compression part, according to the command input by the user, the scheme to be adopted is determined. According to different schemes, the bam file is read to generate a long matrix, and the long matrix is converted into a binary file, a gene name file, and a file information file, and then these three files are stored in the specified path in zip format.

[0097] When selecting Scheme 1, a single-threaded method is used to process the data without considering duplicate coordinates. The data input format is as follows: script, bam file path, number of threads used, and method number.

[0098] The specific execution process is as follows:

[0099] Print the system name and give a brief introduction, including the system name, version information, and a brief introduction;

[0100] Print the current time, and the system starts running;

[0101] Read the bam file to generate a long matrix mainly based on RNA, including information such as x, y, gene, intron, and exon;

[0102] For the preprocessing of the long matrix, obtain information such as the gene list information, the number of rows and columns of the long matrix, etc.;

[0103] Print that the data is processed using Scheme 1;

[0104] Read the long matrix, convert the gene name into its label geneID, and convert the information such as x, y, geneID, and intron into binary and save it in a specific binary long matrix;

[0105] Generate a gene information long matrix and a file information long matrix according to the long matrix information;

[0106] Save the binary long matrix, the gene information long matrix, and the file information long matrix in zip format to the specified path;

[0107] Print the current time, and the system running section ends;

[0108] Display the total time used for the current file compression process.

[0109] When using a multi-threaded method, consider processing the data in a multi-threaded manner and consider duplicate coordinates. The input command is as follows: script, bam file path, number of threads used, and scheme number.

[0110] The specific execution process is as follows:

[0111] Print the name of the printing system and give a brief introduction, including the system name, version information, and a brief introduction;

[0112] Print the current time and the start of system operation;

[0113] Read the bam file to generate a long matrix mainly based on RNA, including information such as x, y, gene, intron, exon, etc.;

[0114] Preprocess the long matrix to obtain information such as the gene list information, the number of rows and columns of the long matrix;

[0115] Print that the data is processed using Scheme II;

[0116] Read the long matrix, and according to the number of input threads, split the long matrix. Each thread is responsible for a part of the long matrix; in the form of a thread pool, convert the gene name into its label geneID, and convert information such as x, y, geneID, intron, etc. into binary and save it to a specific binary long matrix; according to the thread number, integrate the binary long matrix processed by each thread into the final binary long matrix;

[0117] Generate a gene information long matrix and a file information long matrix according to the long matrix information;

[0118] Save the binary long matrix, the gene information long matrix, and the file information long matrix in zip format to the specified path;

[0119] Print the current time and the end of system operation;

[0120] Display the total time used in the current file compression process.

[0121] The main purpose of the data recovery process is to read the compressed file, decompress it into a gene name file, a binary file, an information file, restore it to a sparse matrix, and generate a new data structure. The data recovery process is as Figure 6 shown, and the storage method of the data dictionary after data recovery is as Figure 7 shown. Among them, combined with Figure 6 shown, the specific data recovery process is as follows: In the data compression recovery part, according to the command input by the user, determine the scheme to be adopted, read the zip file according to different schemes, decompress the zip file to generate a gene name file, a binary file, an information file, generate a long matrix, and generate dictionary data.

[0122] When selecting Scheme I, a single-threaded method is used to process the data without considering duplicate coordinates. The data input format is as follows: script, bam file path, number of threads used, and method number.

[0123] The specific execution process of the single-threaded method is as follows:

[0124] Print the system name and provide a brief introduction, including the system name, version information, and a brief introduction;

[0125] Print the current time and start the system operation;

[0126] Read the zip file and decompress it to generate a binary long matrix, a long matrix of gene names, and a long matrix of file information;

[0127] Read the long matrix of gene names and the long matrix of file information;

[0128] Read the binary file and convert it into a long matrix mainly based on RNA signals;

[0129] Convert the long matrix of gene names and the long matrix mainly based on RNA signals into dictionary data for subsequent analysis;

[0130] Print the current time and end the system operation session;

[0131] Display the total time used for the current file compression process.

[0132] In the multi-threaded mode, consider using a multi-threaded approach to process data, taking into account duplicate coordinates. The input commands are as follows: script, bam file path, number of threads used, and mode number.

[0133] The specific execution process is as follows:

[0134] Print the system name and provide a brief introduction, including the system name, version information, and a brief introduction;

[0135] Print the current time and start the system operation;

[0136] Read the zip file and decompress it to generate a binary long matrix, a long matrix of gene names, and a long matrix of file information;

[0137] Read the long matrix of gene names and the long matrix of file information;

[0138] Print that data is processed using Method 2;

[0139] Read the long matrix. According to the input number of threads, divide the long matrix. Each thread is responsible for a part of the long matrix. Adopt the thread pool method to read the binary file and convert it into a long matrix mainly based on RNA signals. According to the thread number, integrate the long matrices mainly based on RNA signals converted by each thread into the final spatial omics text data;

[0140] Convert the long matrix of gene names and the long matrix mainly based on RNA signals into dictionary data for subsequent analysis;

[0141] Print the current time and end the system operation session;

[0142] Displays the total time taken for the current file compression process.

[0143] As Figure 8 shown, the input data format is shown:

[0144] Adopt Solution 1 to process the data, use a single thread, and do not consider duplicate coordinates. Dataset size: 81.24M, one million rows.

[0145] During the data compression process, the input command is like python txtToByte.py. / data / aa.bam 1 1, including the environment, script, bam file path, number of threads used, and method number. The result display after executing the data compression command is as Figure 9 shown.

[0146] During the data recovery process, as Figure 10 shown, the input command is for example python byteToTxt.py. / aa.zip 11, including the environment, script, zip file path, number of threads used, and method number. After execution: the input recovery result is as Figure 11 shown.

[0147] Use multiple threads and consider duplicate coordinates. Dataset size: 825.67M, ten million rows. During the data compression process, the input command is as Figure 12 shown:

[0148] python txtToByte.py. / data / LYP01.bam 4 2

[0149] Script, bam file path, number of threads used, solution number

[0150] The data compression result after the input command is as Figure 13 shown.

[0151] During the data recovery process, as Figure 14 shown, the input command:

[0152] python byteToTxt.py. / LYP01.zip 42

[0153] Script, zip file path, number of threads used, solution number

[0154] The display result after the input data is recovered is as Figure 15 shown.

[0155] According to the processing method of spatial omics data in the embodiments of the present application, the consumption of storage resources for spatial omics data can be effectively reduced, and moreover, it has the advantages of fast reading speed of spatial omics data and being convenient for subsequent analysis.

[0156] On the other hand, as Figure 16 shown, an embodiment of the present application provides a processing system for spatial omics data, including: an acquisition module 1610, a partitioning module 1620, and a compression storage module 1630, where:

[0157] The acquisition module 1610 is configured to obtain spatial omics text data;

[0158] The partitioning module 1620 is configured to obtain a gene name file and an information file according to the spatial omics text data, and obtain a binary file according to the gene name file and the spatial omics text data;

[0159] The compression storage module 1630 is configured to compress and store the gene name file, the information file, and the binary file.

[0160] According to the processing system for spatial omics data in the embodiment of the present application, the consumption of storage resources for spatial omics data can be effectively reduced, and it has the advantages of fast reading speed of spatial omics data and being convenient for subsequent analysis.

[0161] It should be noted that the specific implementation manner of the processing system for spatial omics data in the embodiment of the present application is similar to the specific implementation manner of the processing method for spatial omics data in the embodiment of the present application. For details, please refer to the description in the method part, and details will not be repeated here.

[0162] Figure 17 It is a schematic structural diagram of a computing device according to an embodiment of the present application.

[0163] As Figure 17 shown, the computing device 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 602 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computing device 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0164] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 610 as required so that a computer program read therefrom is installed into the storage section 608 as required.

[0165] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a machine-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by a central processing unit (CPU) 601, the above-described functions defined in the computing device of the present application are performed.

[0166] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor computing device, system, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction-executing computing device, system, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction-executing computing device, system, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0167] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of a processing receiving device, method, and computer program product according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in a block can occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and the combination of blocks in a block diagram and / or flowchart, can be implemented by a dedicated hardware-based computing device that executes a specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0168] The units or modules involved in the embodiments of the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor, and when the processor executes the program, it implements a method for processing spatial omics data:

[0169] Obtain spatial omics text data;

[0170] Obtain a gene name file and an information file according to the spatial omics text data, and obtain a binary file according to the gene name file and the spatial omics text data;

[0171] Compress and store the gene name file, the information file, and the binary file.

[0172] As another aspect, the present application also provides a computer-readable storage medium, which can be included in the computing device described in the above embodiments; or can exist alone without being assembled into the computing device. The above computer-readable storage medium stores one or more programs, and when the foregoing programs are executed by one or more processors, they implement the method for processing spatial omics data described in the present application:

[0173] Obtain spatial omics text data;

[0174] Obtain a gene name file and an information file according to the spatial omics text data, and obtain a binary file according to the gene name file and the spatial omics text data;

[0175] Compress and store the gene name file, the information file, and the binary file.

[0176] As another aspect, the present application also provides a computer program product, which can be included in the computing device described in the above embodiments; or can exist alone without being assembled into the computing device. The above computer program product stores one or more programs, and when the foregoing programs are executed by one or more processors, they implement the method for processing spatial omics data described in the present application:

[0177] Obtain spatial omics text data;

[0178] Obtain a gene name file and an information file according to the spatial omics text data, and obtain a binary file according to the gene name file and the spatial omics text data;

[0179] Compress and store the gene name file, the information file, and the binary file.

[0180] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the foregoing disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.

Claims

1. A method for processing spatial omics data, characterized in that, it includes: obtaining spatial omics text data; obtaining a gene name file and an information file according to the spatial omics text data, and obtaining a binary file according to the gene name file and the spatial omics text data; compressing and storing the gene name file, the information file and the binary file.

2. The method for processing spatial omics data according to claim 1, characterized in that, the obtaining of the spatial omics text data includes: obtaining spatial omics raw data, and generating spatial omics text data according to the spatial omics raw data, including: extracting coordinate information, gene expression levels and gene information from the spatial omics raw data; obtaining additional information, wherein the additional information includes ssDNA staining information; generating the spatial omics text data according to the coordinate information, gene expression levels, gene information and the additional information.

3. The method for processing spatial omics data according to claim 1 or 2, characterized in that, the obtaining of the gene name file and the information file according to the spatial omics text data includes: reading gene names from the spatial omics text data, and obtaining the gene name file according to the gene names and the subscripts of the gene names; extracting the file information of the spatial omics text data from the spatial omics text data, and forming the information file according to the file information of the spatial omics text data, wherein the information file at least includes the number of rows, the number of columns and the header information of the spatial omics text data.

4. The method for processing spatial omics data according to claim 3, characterized in that, the obtaining of the binary file according to the gene name file and the spatial omics text data includes: reading the gene names in the spatial omics text data line by line, and converting the read gene names into corresponding subscripts according to the subscripts corresponding to the gene names in the gene name file; forming the binary file according to the converted subscripts and the gene coordinates in the spatial omics text data.

5. The method for processing spatial omics data according to claim 3, characterized in that, the obtaining of the binary file according to the gene name file and the spatial omics text data includes: dividing the spatial omics text data into multiple spatial omics text data blocks; performing parallel processing on the multiple spatial omics text data blocks, wherein the processing process of each spatial omics text data block is: reading the gene names in the spatial omics text data block line by line, converting the read gene names into corresponding subscripts according to the subscripts corresponding to the gene names in the gene name file, and forming a binary file block according to the converted subscripts and the gene coordinates in the spatial omics text data block; forming the binary file according to the binary file blocks obtained from each spatial omics text data block.

6. The method for processing spatial omics data according to claim 1, characterized in that, it further includes: reading the compressed file; decompressing the compressed file to obtain the gene name file, the binary file and the information file; Convert the binary file into the spatial omics text data according to the gene name file and the information file.

7. The method for processing spatial omics data according to claim 6, wherein, the converting the binary file into the spatial omics text data according to the gene name file and the information file includes: Read the information file to obtain the header information and the table information; Read the corresponding bytes in the binary file sequentially according to the number of elements in each row of the information file; Write the read table information and bytes into the spatial omics text data; or, Read the information file to obtain the header information and the table information; Obtain multiple row number sub-intervals according to the information file; Perform parallel processing on the binary file based on multi-threading, where one thread corresponds to at least one row number sub-interval. Among them, the processing process of each thread includes: sequentially read the corresponding bytes in the binary file according to the row number sub-interval corresponding to the thread, and write the read table information and bytes into the spatial omics text data block; Obtain the spatial omics text data according to each spatial omics text data block.

8. The method for processing spatial omics data according to claim 6 or 7, wherein, After reading the spatial omics text data, it further includes: converting the spatial omics text data into a data dictionary of a predetermined data structure.

9. A processing system for spatial omics data, wherein, it includes: An acquisition module for obtaining spatial omics text data; A division module for obtaining a gene name file and an information file according to the spatial omics text data, and obtaining a binary file according to the gene name file and the spatial omics text data; A compression storage module for compressing and storing the gene name file, the information file and the binary file.

10. A computing device, wherein, the computing device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is used to implement the method for processing spatial omics data according to any one of claims 1-8 when executing the program.