A construction method and system of a disease-specific omics database

By constructing a special disease omics database, combining the mapping of gene expression data with biological pathways and immune cell levels, the problem that a single-dimensional database in the existing technology is difficult to meet the multi-dimensional disease research, and the integration and in-depth analysis of multi-dimensional data is achieved.

CN118588169BActive Publication Date: 2025-07-25QINGDAO KELI BIOPHARMACEUTICAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410636989.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-22
Publication Date
2025-07-25
Estimated Expiration
2044-05-22

AI Technical Summary

Technical Problem

The existing omics databases mainly focus on single-dimensional data, which is difficult to meet the data presentation and research needs of complex diseases in multiple dimensions, and cannot effectively reveal the connections between diseases at other levels such as genes and cells.

Method used

A special disease omics database was constructed, and by obtaining the gene expression data of the disease sample set, mapping it to the biological pathway and immune cell levels, establishing the association relationship between genes, pathways, and immune cells, and storing it. The enrichment difference score was used for GSVA, GSEA and other methods, and combining algorithms such as cibersort to calculate the relative abundance of cell types, and a multi-dimensional database was constructed.

Benefits of technology

The data integration of diseases in multiple dimensions has been achieved, providing a more comprehensive means of disease research, and helping to deeply understand the mechanism of the disease and judge the efficacy of drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118588169B_ABST
    Figure CN118588169B_ABST
Patent Text Reader

Abstract

This application relates to the field of bioinformatics, and specifically relates to a method and system for constructing a disease-specific omics database. The method includes obtaining gene expression data and gene IDs of a disease sample set; mapping the gene expression data to the biological pathway level to obtain pathway data and pathway IDs; mapping the gene expression data to the immune cell level to obtain immune cell data and immune cell IDs; associating the pathway IDs, the immune cell IDs with the gene IDs to obtain an association relationship; and storing the association relationship, gene expression data, pathway data, and immune cell data to obtain a disease-specific omics database. The database constructed by this method contains data on genes, pathways, and immune cells, and can obtain data on a certain disease in multiple dimensions simultaneously, having good clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics, and specifically relates to a method and system for constructing a disease-specific omics database. Background Art

[0002] With the progress of technology, especially the popularization of large-scale sequencing technology, genomic data is generated faster and in larger quantities. This requires omics databases to provide more powerful storage and processing capabilities to meet the needs of scientists. In addition, the rapid development of single-cell sequencing technology has also enabled the collection and analysis of more genomic data, providing researchers with opportunities to study more complex biological phenomena, such as cell differentiation and tumor heterogeneity. Secondly, there are a variety of omics databases, each with its own characteristics. For example, genotype and phenotype databases (such as dbGaP) are used to archive, select, and publish research information on the interaction between genotypes and phenotypes; The Human Protein Atlas database provides information on target proteins in immunohistochemistry or tumor immunity experiments; while the PhosphoSitePlus database focuses on summarizing protein modification sites, helping to study the role of PTMs in normal and pathological cells / tissues. However, these databases focus on single-dimensional data, and the causes of human diseases are often intricate. Single-dimensional databases are difficult to meet the data presentation of diseases in multiple dimensions and are not conducive to studying the connections between single diseases at other levels such as genes and cells. Summary of the Invention

[0003] In view of the above problems, the present invention proposes a method for constructing a disease-specific omics database, including:

[0004] Obtaining gene expression data and gene IDs of a disease sample set;

[0005] Mapping the gene expression data to the biological pathway level to obtain pathway data and pathway IDs;

[0006] Mapping the gene expression data to the immune cell level to obtain immune cell data and immune cell IDs;

[0007] Associating the pathway IDs, the immune cell IDs with the gene IDs to obtain an association relationship;

[0008] Storing the association relationship, gene expression data, pathway data, and immune cell data to obtain a disease-specific omics database.

[0009] Furthermore, the mapping to the biological pathway level is achieved by converting gene expression data into a gene set, performing an enrichment difference score on the genes in the gene set, and then sorting the scores to obtain high-expression level pathways;

[0010] Optionally, the conversion method includes one or more of the following: GSVA, GSEA, SSGSEA, ZSCORE, PLAGE;

[0011] Optionally, the gene expression data is from one or more of the following: GO dataset, KEGG dataset.

[0012] The mapping to the immune cell level is as follows: perform linear support vector regression analysis on the gene expression data, calculate the relative abundance of each cell type in the sample, and compare and calculate the relative abundance of the cell type with the data in the gene expression library unique to the cell type to obtain the estimated proportion of each cell type;

[0013] Optionally, the method for mapping to immune cells includes one or more of the following: cibersort, ssGSEA, Cibersoft, TIMER, EPIC.

[0014] Optionally, the gene expression library unique to the cell type is the LM22 dataset.

[0015] The method further includes data preprocessing, and the preprocessing includes one or more of the following: decentralization, data grouping, duplicate removal, ID matching, normalization.

[0016] Optionally, for duplicate removal, calculate the median gene expression based on the genes in the gene expression data, sort the medians, and then remove duplicate gene data based on the sorted medians.

[0017] Optionally, the normalization is performed based on z-score. Among them, calculate the gene length in the expression data through the GenomicFeatures function, and the gene length is matched one by one with the genes; for genes for which the gene length cannot be matched, calculate the mean value of the genes, and then use the value obtained by dividing the mean value by the gene median as the gene length.

[0018] The method further includes annotation file collation. Obtain disease pathway information, disease gene information, and disease immune cell information from the database, and the disease pathway information, disease gene information, and disease immune cell information constitute the annotation file;

[0019] Optionally, the database for extracting annotation data of the gene information includes one or more of the following: GTF file;

[0020] Optionally, the database for extracting annotation data of the pathway information includes one or more of the following: msigdb database, GO database;

[0021] Optionally, the database for the immunocyte information extraction annotation data includes one or more of the following: cibersort.

[0022] The method further includes retrieval file production. The retrieval file is associated with the annotation file, and the disease gene information, disease pathway information, and disease immunocyte information are combined into one file to obtain the retrieval file.

[0023] The object of the present invention is to provide a construction system for a disease-specific omics database, including:

[0024] A memory and a processor. The memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, the above-mentioned construction method of the disease-specific omics database is implemented.

[0025] The object of the present invention is to provide a construction system for a disease-specific omics database, including:

[0026] An acquisition unit: acquiring gene expression data of a disease sample set;

[0027] A mapping unit: mapping the gene expression data to the biological pathway level to obtain pathway data and pathway IDs; mapping the gene expression data to the immunocyte level to obtain immunocyte data and immunocyte IDs;

[0028] An association unit: associating the pathway IDs, the immunocyte IDs, and gene IDs to obtain an association relationship;

[0029] A storage unit: storing the association relationship, gene expression data, pathway data, and immunocyte data to obtain a disease-specific omics database.

[0030] The object of the present invention is to provide a disease-specific omics database, including: the disease-specific omics database is obtained by executing the above-mentioned construction method of the disease-specific omics database;

[0031] A storage unit: storing the data of the disease-specific omics database;

[0032] A query unit: querying data from the disease-specific omics database through a query instruction;

[0033] An output unit: outputting data;

[0034] Optionally, the query data includes one or more of the following: gene information of a disease, pathway information, and immunocyte information;

[0035] Optionally, the query instruction includes: retrieving by gene type or retrieving by pathway type or retrieving by immunocyte type.

[0036] The object of the present invention is to provide a method for obtaining disease-specific information based on a disease-specific database, including: the method obtains disease-specific information from the above-mentioned disease-specific omics database, wherein a query instruction is input to a query unit in the disease-specific omics database, and the disease-specific omics database calls data based on the query instruction and outputs through an output unit.

[0037] Optionally, the disease-specific omics database further includes a server and a client. The server, the client, and the disease-specific omics database are communicatively connected to each other in pairs. The client logs in to access the server and inputs a search instruction. The server accesses the disease-specific omics database based on the search instruction. The disease-specific omics database obtains the search content through arithmetic processing and returns it to the client;

[0038] Optionally, data analysis is performed on the disease-specific information obtained through the query to obtain disease-specific analysis data;

[0039] Optionally, data visualization is performed on the disease-specific analysis data to obtain disease-specific analysis visualization data.

[0040] Optionally, data visualization is performed on the disease-specific information obtained through the query to obtain disease-specific visualization data.

[0041] Furthermore, the retrieval of the disease-specific database includes one or more of the following: when obtaining data from the disease-specific omics database, retrieving according to gene type to obtain corresponding gene data, the pathways corresponding to the genes, and immune cell data; retrieving according to pathway type to obtain corresponding pathway data, the genes corresponding to the pathways, and immune cell data; retrieving according to immune cell type to obtain corresponding immune cell data, the genes corresponding to the immune cells, and pathway data;

[0042] Optionally, the method for obtaining data from the disease-specific database further includes feature data retrieval, and data satisfying the features are obtained through feature screening in N dimensions, where N is a natural number greater than or equal to 1.

[0043] Advantages of the present invention:

[0044] A disease-specific omics database of a certain disease is constructed using data at the gene level, pathway level, and cell level. The disease-specific database is beneficial for studying the relationships of data at different levels of the disease, can obtain disease data in multiple dimensions simultaneously, and helps in the in-depth study of the disease. Description of the Drawings

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0046] Figure 1 Schematic flow diagram of the construction method of the disease-specific omics database provided by the embodiments of the present invention;

[0047] Figure 2 Schematic diagram of the construction system of the disease-specific omics database provided by the embodiments of the present invention;

[0048] Figure 3 Schematic diagram of the primary retrieval of the disease-specific omics database provided by the embodiments of the present invention;

[0049] Figure 4 Schematic diagram of the secondary retrieval of the disease-specific omics database provided by the embodiments of the present invention;

[0050] Figure 5 Schematic diagram of the data retrieval of the disease-specific omics database provided by the embodiments of the present invention;

[0051] Figure 6 Schematic diagram of the data visualization of the disease-specific omics database provided by the embodiments of the present invention;

[0052] Figure 7 Schematic diagram of the feature query of the disease-specific omics database provided by the embodiments of the present invention;

[0053] Figure 8 Schematic diagram of the data analysis of the disease-specific omics database provided by the embodiments of the present invention. Detailed implementation manners

[0054] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention.

[0055] In some of the processes described in the specification, claims, and the above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0056] Figure 1 The schematic diagram of the method for constructing a disease omics database provided by the embodiment of the present invention specifically includes:

[0057] S101: Obtain the gene expression data and gene IDs of a disease sample set;

[0058] In one embodiment, the gene expression data reflects the abundance of the gene transcript mRNA directly or indirectly measured in cells. These data can be used to analyze which genes have changed in expression, what the correlations are between genes, and how the activities of genes are affected under different conditions. They have important applications in aspects such as medical clinical diagnosis, judgment of drug efficacy, and revealing the mechanism of disease occurrence.

[0059] In one embodiment, Gene ID (also known as Entrez ID): is a number provided by the gene-specific database of the National Center for Biotechnology Information (NCBI) in the United States, and it is an integer. Gene ID is unique and stable and is currently the most authoritative gene ID internationally. One Entrez ID may correspond to multiple Gene symbols. For example, AKT also has other names such as PKB, RAC, etc.

[0060] In one embodiment, the method includes data preprocessing, and the preprocessing includes one or several of the following: decentralization, data grouping, duplicate removal, ID matching, and standardization.

[0061] Optionally, for the duplicate removal, the median of gene expression is calculated based on the genes in the gene expression data, the medians are sorted, and then the duplicate gene data is removed based on the sorted medians.

[0062] Optionally, the standardization is based on z-score. Among them, the gene lengths in the expression data are calculated through the GenomicFeatures function, and the gene lengths are in one-to-one correspondence with the genes; for the genes for which the gene lengths cannot be obtained by matching, the mean value of the genes is calculated, and then the value obtained by dividing the mean value by the gene median is used as the gene length.

[0063] In a specific embodiment, the expression data sorting module, as shown in Table 1, is used to sort the expression matrix data;

[0064] a) Uniformly summarize and number the project grouping information;

[0065] b) Remove the problems that may be caused by coding issues;

[0066] c) Implement the correspondence between gene IDs and gene symbols in the expression matrix;

[0067] d) Standardize the expression matrix;

[0068] Table 1 Expression Data Sorting Module

[0069]

[0070] Specifically, the expression matrix data is sorted to obtain a unified expression data format for facilitating the next-step analysis, and the gene expression data downloaded from NCBI, TCGA, etc. is sorted.

[0071] a) Sort the sample grouping information, convert the grouping provided in the official document to lowercase, and uniformly summarize the similar information of different projects; connect multiple words with _.

[0072] b) For project information, character encoding problems are likely to occur when directly copying from the web page. After reading into R and sorting, it is necessary to write the file in utf-8 to check whether there are warning messages;

[0073] c) For the expression matrix, some matrices will become character matrices when read into R and need to be converted to numbers; reserve the first two columns for gene IDs, and those not provided in the original file can be pre-filled with 0. NA cannot be written into the h5 file (written according to filling 0 in the R code); remove duplicates according to the gene expression median provided by the gene to facilitate handling the one-to-many situation that occurs during gene conversion.

[0074] d) When there are expression matrices of different tissues / sequencing types in a GSE project, the project name in the h5 is split, GSE***__, two underscores, the first underscore is followed by the sequencing type, and the second underscore is followed by the tissue type; otherwise, according to GSE***__, that is, only use two underscores.

[0075] e) Gene id conversion: Uniformly name the gene formats of each project, establish a one-to-one correspondence with the gene names (gene symbols), and write them into the first two columns of the expression matrix, as shown in Table 2. The uniformly named gene format for circRNA is "chromosome_start position_end position_plus or minus strand" (e.g.

[0076] chr12_938227_939110_+), and the uniformly named gene format for miRNA is miR-**;

[0077] Table 2 Gene id conversion

[0078]

[0079] f) Standardization of the expression matrix: Convert the count data in the expression matrix into tpm / fpkm data, and then standardize the expression data into tpm data.

[0080] unitrans calculates the z-score; count to tpm / fpkm; fpkm to tpm;

[0081] Use the v43 annotation file, and the r package GenomicFeatures calculates the gene length (May 2023, v43 is the latest version).

[0082] Some genes in the lnc / m expression matrix cannot be matched to the genes in the gtf file, and there is no gene length;

[0083] Supplementary method: Calculate the mean / median = gene length, and then calculate the expression matrix.

[0084] S102: Map the gene expression data to the biological pathway level to obtain pathway data and pathway IDs;

[0085] In one embodiment, gene set variation analysis (GSVA) is a special type of gene set enrichment method that achieves a path-centered analysis of molecular data through a conceptually simple but powerful change to the functional unit being analyzed - from genes to gene sets. Simply put, it changes the analysis object from genes to gene sets and performs differential analysis at the gene set (pathway) level.

[0086] In one embodiment, GO (Gene Ontology) is a database established by the Gene Ontology Consortium, aiming to establish a semantic vocabulary standard applicable to various species, which defines and describes the functions of genes and proteins and can be updated as research progresses. GO is one of the multiple biological ontology languages and provides a system definition method with a three-layer structure for describing the functions of gene products. They divide the functions of genes into three parts, namely: Cellular Component (CC), Molecular Function (MF), and Biological Process (BP).

[0087] In one embodiment, KEGG (Kyoto Encyclopedia of Genes and Genomes) was established in 1995 by the Kanehisa Laboratory at the Bioinformatics Center of Kyoto University in Japan. It is a database that integrates genomic, chemical, and systematic functional information. KEGG correlates the gene catalog obtained from the genomes that have been completely sequenced with the systematic functions at the higher levels of cells, species, and ecosystems.

[0088] In one embodiment, GSEA (Gene Set Enrichment Analysis) uses predefined gene sets to rank genes according to their differential expression levels in two types of samples, and then tests whether the predefined gene sets are enriched at the top or bottom of this ranking list. Gene set enrichment analysis detects the expression changes of gene sets rather than individual genes, so it can include these subtle expression changes and is expected to obtain more ideal results.

[0089] In one embodiment, mapping to the biological pathway level is achieved by converting gene expression data into gene sets, performing enrichment difference scoring on the genes in the gene sets, and then ranking the scores to obtain high-expression level pathways;

[0090] Optionally, the conversion methods include one or more of the following: GSVA, GSEA, SSGSEA, ZSCORE, PLAGE;

[0091] Optionally, the gene expression data comes from one or more of the following: GO dataset, KEGG dataset.

[0092] In a specific embodiment, the expression matrix analysis module analyzes the organized expression data to obtain the biological pathways mapped by genes and the expression analysis results at the immune cell level;

[0093] a) At the level of biological pathways, GSVA was used to score the KEGG and GO gene sets;

[0094] The KEGG and GO data sets were downloaded from the GSEA database:

[0095] c2.cp.kegg.v2022.1.Hs.symbols.gmt","c5.go.v2022.1.Hs.symbols.gmt;

[0096] Description of the GSVA algorithm:

[0097] Normalize the expression data: First, the input gene expression data was normalized to ensure comparability between different samples.

[0098] Gene ranking: For each sample, all genes were ranked according to their expression values. This step was for calculating the enrichment score of the gene set based on the gene expression level subsequently.

[0099] Kernel smoothing transformation: Using the kernel smoothing method (such as Gaussian kernel), according to the gene ranking and expression values, calculate the enrichment score of each gene set in each sample. This step evaluates the expression activity of the gene set as a whole in a specific sample by considering the expression ranking and relative position of each gene in the gene set.

[0100] Calculate the enrichment score: Through kernel smoothing transformation, an enrichment score (ES) was calculated for each gene set in each sample. This score reflects the degree of deviation of the gene set's expression level in the sample from the background expression level, which can be positive (upregulated expression) or negative (downregulated expression).

[0101] S103: Map the gene expression data to the immune cell level to obtain immune cell data and immune cell IDs;

[0102] In one embodiment, the CIBERSORT algorithm is a tool for cell type analysis based on gene expression data. It uses the principle of linear support vector regression to deconvolve the expression matrix of immune cell subtypes to estimate the abundance of immune cells. This algorithm requires preparing two main files:

[0103] A reference data set containing the gene expression profiles of known cell types, namely the LM22 file, which covers 22 common immune infiltrating cells, including immune cells with different cell types and functional states. An expression matrix file, whose row names are gene names and column names are sample information.

[0104] The advantage of the CIBERSORT algorithm is that it can decompose the cell type abundances from bulk samples without physically separating the cells, and it also does not require the use of antibodies or live substances to study the detailed description of tissue composition.

[0105] In one embodiment, mapping to the immune cell level is as follows: performing linear support vector regression analysis on gene expression data, calculating the relative abundances of each cell type in the sample, and comparing and calculating the abundances of the cell types with the data in the gene expression library specific to the cell types to obtain the estimated proportion of each cell type;

[0106] Optionally, the method of mapping to immune cells includes one or more of the following: cibersort, ssGSEA, Cibersoft, TIMER, EPIC.

[0107] Optionally, the gene expression library specific to the cell type is the LM22 dataset.

[0108] In a specific embodiment, at the immune cell level, using LM22 as the reference dataset, cibersort is used to calculate the civersort score;

[0109] Description of the Cibersort algorithm:

[0110] Input data processing: Input the gene expression data of the mixed cell sample. These data usually need to be pre-normalized for comparison with the reference gene expression signature library.

[0111] Prepare the reference gene expression signature: Construct a reference library containing the gene expression patterns specific to different cell types. This library is crucial for the operation of CIBERSORT because the algorithm estimates the proportions of each cell type in the sample by comparing with these reference signatures.

[0112] Apply linear support vector regression (SVR): CIBERSORT uses linear support vector regression to analyze the gene expression data in the mixed sample. By minimizing the difference between the actual gene expression values and the values predicted by the reference gene expression signatures, the algorithm can estimate the relative abundances of each cell type in the mixed sample.

[0113] Estimation of cell components (score): The algorithm outputs the estimated proportions of each cell type in each sample, which are based on the results of the optimization process, taking into account the comparison between the expression patterns of the reference signatures and the actual sample data.

[0114] S104: Associate the pathway ID, the immune cell ID with the gene ID to obtain an association relationship;

[0115] In a specific embodiment, gene expression data is mapped to the biological pathway level and immune cell level to obtain pathway data and pathway IDs, immune cell data and immune cell IDs, as shown in Table 3;

[0116] Table 3

[0117]

[0118] The pathway IDs and immune cell IDs corresponding to the gene IDs are associated in a data table, and other IDs corresponding to the gene IDs are obtained by querying the ID data table. Then, relevant data content is obtained based on the other IDs;

[0119] In one embodiment, three-dimensional data content of genes, pathways, and immune cells is obtained by querying an ID data table in a database.

[0120] S105: Store the association relationship, gene expression data, pathway data, and immune cell data to obtain a disease omics database.

[0121] In a specific embodiment, the results are written into an H5 file:

[0122] Use the h5write function of the rhdf5 package in R language to save all the above data. The h5write function will generate an h5 format file locally. The H5 file has the following advantages:

[0123] Hierarchical data organization: HDF5 supports complex data organization forms. Data can be stored hierarchically in a single file, similar to the directory and file structure in a file system.

[0124] Efficient data access: Through optimized I / O operations and support for partial read and write of data sets, HDF5 can efficiently process large-scale data sets.

[0125] Cross-platform compatibility: HDF5 files have good compatibility between different operating systems, including Windows, Linux, and MacOS.

[0126] Support for multiple data types: HDF5 can store different types of data, including images, tables, text, and multi-dimensional arrays, etc.

[0127] Scalability and flexibility: Users can customize the metadata of the data, making data organization and management more flexible.

[0128] In one embodiment, the method further includes annotation file collation. By obtaining disease pathway information, disease gene information, and disease immune cell information from the database, the disease pathway information, disease gene information, and disease immune cell information form an annotation file;

[0129] Optionally, the database for extracting and annotating gene information includes one or more of the following: GTF files;

[0130] Optionally, the database for extracting and annotating pathway information includes one or more of the following: msigdb database, GO database;

[0131] Optionally, the database for extracting and annotating immune cell information includes one or more of the following: cibersort.

[0132] In one embodiment, the method further includes retrieving file production. The retrieved file is associated with the annotation file, and the disease gene information, disease pathway information, and disease immune cell information are combined into one file to obtain the retrieved file.

[0133] In a specific embodiment, the annotation file sorting module includes a gene annotation unit, a pathway information annotation unit, and an immune cell annotation unit, which sort the annotation files at the gene, pathway, and immune cell levels and write them into an H5 file.

[0134] Specifically, the downloaded annotation files are sorted. For the gene annotation files, the GTF files in gencode are used. For the pathway information, the annotation files in the mgigdb and GO databases are used. For the immune cells, the additional materials provided in the cibersort literature are used.

[0135] For gene annotation, the GTF file (v43) downloaded from the gencode official website is used. The "NA" in the annotation file is replaced with "-", and the sorted information is written into the h5 file.

[0136] For pathway information, first, the human kegg and go gene set data are downloaded from the GSEA MSigDB database. Then, the rvest package is used to crawl the information of the msigdb and GO databases, and the obtained information is integrated, mainly including the Standard name, Systematic name, Brief description, Full description or abstract, Collection, Source publication, Exact source, Related gene sets, External links, Filteredbysimilarity, Source species, Contributed by, Source platformoridentifier namespace, Dataset references, etc. of each gene set.

[0137] The immune cell information uses the gene expression feature data of 22 immune cells provided by the cibersort literature, including 7 T cell types, naive and memory B cells, plasma cells, NK cells, and myeloid subsets. The sorted data contains information such as "Cell_Type_Description", "Reference_PMID", "Cell_Separation_Method", "Markers_used", "Purity", etc.

[0138] All the sorted data above is written into an H5 file.

[0139] In a specific embodiment, the retrieval file module is used to create a retrieval file for facilitating the retrieval of annotation information; an index file is established for the obtained annotation file to facilitate data retrieval, and this step is implemented simultaneously during the file sorting process. The project data is written into the H5 file. When the project to which the file belongs does not exist, h5createGroup is used to create the project and then write the data.

[0140] Furthermore, obtain the overall information file of the project, which includes 10 items such as project, category, strategy, disease, tissue, platform, source, condition, samples, expr. Among them, the information of project, category, condition, samples, expr comes from the H5 file, and the other 5 columns of information need to be filled in manually.

[0141] Figure 2 The schematic diagram of the construction system of the disease-specific omics database provided by the embodiment of the present invention specifically includes:

[0142] A memory and a processor, where the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, the above-mentioned method for constructing the disease-specific omics database is implemented.

[0143] A disease-specific omics database, including: the disease-specific omics database is obtained by executing the above-mentioned method for constructing the disease-specific omics database;

[0144] A storage unit: storing the data of the disease-specific omics database;

[0145] A query unit: querying data from the disease-specific omics database through a query instruction;

[0146] An output unit: outputting data;

[0147] Optionally, the query data includes one or more of the following: gene information of a disease, pathway information, immune cell information;

[0148] Optionally, the query instruction includes: retrieving by gene type, or by pathway type, or by immune cell type.

[0149] A method for obtaining disease-specific information based on a disease-specific database, including: the method obtains disease-specific information by obtaining disease-specific information from the above-mentioned disease-specific omics database, wherein the query instruction is input to the query unit in the disease-specific omics database, and the disease-specific omics database calls data based on the query instruction and outputs through the output unit.

[0150] Optionally, the disease-specific omics database further includes a server and a client. The server, the client, and the disease-specific omics database are communicatively connected to each other. The client logs in to access the server and inputs a search instruction. The server accesses the disease-specific omics database based on the search instruction. The disease-specific omics database obtains the search content through arithmetic processing and returns it to the client;

[0151] Optionally, data analysis is performed on the disease-specific information obtained through the query to obtain disease-specific analysis data;

[0152] Optionally, data visualization is performed on the disease-specific analysis data to obtain disease-specific analysis visualization data.

[0153] Optionally, data visualization is performed on the disease-specific information obtained through the query to obtain disease-specific visualization data.

[0154] In one embodiment, the retrieval of the disease-specific database includes one or more of the following: when obtaining data from the disease-specific omics database, retrieving by gene type to obtain corresponding gene data and the pathways and immune cell data corresponding to the genes, retrieving by pathway type to obtain corresponding pathway data and the genes and immune cell data corresponding to the pathways, and retrieving by immune cell type to obtain corresponding immune cell data and the genes and pathways data corresponding to the immune cells;

[0155] Optionally, the manner of obtaining data from the disease-specific database further includes feature data retrieval, and data satisfying the features is obtained through feature screening in N dimensions, where N is a natural number greater than or equal to 1.

[0156] In a specific embodiment, the disease-specific (stroke) omics database is a database of various omics data of stroke diseases (such as genomics, transcriptomics, epigenomics, etc.). It provides rich data resources and analyzes the data on a single item or between multiple items based on three dimensions: gene expression characteristics, functional pathway expression characteristics, and immune cell expression characteristics, which can be used for disease research and bioinformatics analysis.

[0157] In a specific embodiment, the software installation of the disease-specific omics database: The minimum version of the operating system is centos8. The running environment requires the installation of the docker engine, and the Nginx service and the software core service are deployed in a containerized manner. The Nginx container needs to modify the default configuration file and change the network mode. The software core depends on the running environments of python and R language, configures the startup mode, and mounts the docker core file of the host machine. The software analysis module image needs to be loaded in tar format, and the running environment of the R language is internally dependent in the analysis module image.

[0158] In a specific embodiment, the retrieval of the disease-specific omics database:

[0159] i. Retrieval item settings;

[0160] Select the retrieval bar, and select the corresponding retrieval keyword in the pop-up drop-down box. The retrieval bar is divided into two retrieval items. The first-level retrieval item selects the retrieval type of the data, and the classification includes as Figure 3 shown. After selecting the first-level retrieval item, the content of the second-level retrieval item pops up the corresponding directory according to the first-level retrieval item. For example, if the first-level retrieval item selects Gene, the second-level retrieval item directory is as Figure 4 shown. Multiple selections are allowed for the first-level retrieval item. When multiple options are selected for the first-level retrieval item, the drop-down directory of the second-level retrieval item is the union of the directories of multiple first-level retrieval items.

[0161] ii. Data retrieval;

[0162] After the retrieval items are selected, click the search button to jump to the detailed data page. Click on each index of the data, and the data can be viewed by sorting according to the index, as Figure 5 shown.

[0163] iii. Data feature visualization;

[0164] You can click on the feature information visualization and feature classification visualization above the data table to visually view the distribution of the data, as Figure 6 shown.

[0165] In a specific embodiment, the function navigation module of the database provides four functions. Hover the mouse over the corresponding label, and the for details button is displayed on the label. Clicking it can jump to the corresponding function page:

[0166] i. Software Introduction. Click to view the basic introduction of the project;

[0167] ii. Data Browse. Click to jump to the data browse page. You can select features through six dimensions.

[0168] After selection, click Search to view the data containing the corresponding features, as shown in Figure 7.

[0169] As shown in Figure 7;

[0170] iii. Data Analysis. Provide multiple data analysis tools, including differential expression analysis, box plot, heatmap, roc, similar and other tools. After selecting the analysis tool you want to use on the left, on the parameter page on the right, select the data and analysis parameters for the analysis, and then click Analysis. After waiting for a while, the analysis results will be displayed, as shown in Figure 8 Figure.

[0171] The verification results of this verification embodiment show that assigning fixed weights to the indications can improve the performance of this method compared to the default settings. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk, or optical disc, etc.

[0172] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned medium storage can be read-only memory, magnetic disk, or optical disc, etc.

[0173] The above has introduced in detail a computer device provided by the present invention. For those of ordinary skill in the art, according to the idea of the embodiments of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A construction method of a disease-specific omics database, characterized in that, The method includes: Obtaining gene expression data and gene IDs of a disease sample set; Mapping the gene expression data to the biological pathway level to obtain pathway data and pathway IDs; the mapping to the biological pathway level is achieved by converting the gene expression data into a gene set, performing an enrichment difference score on the genes in the gene set, and then sorting the scores to obtain high-expression level pathways; Mapping the gene expression data to the immune cell level to obtain immune cell data and immune cell IDs; the mapping to the immune cell level is as follows: performing linear support vector regression analysis on the gene expression data, calculating the relative abundance of each cell type in the sample, and comparing and calculating the relative abundance of the cell type with the data in the gene expression library unique to the cell type to obtain the estimated proportion of each cell type; Associating the pathway IDs, the immune cell IDs with the gene IDs to obtain an association relationship; associating the pathway IDs and immune cell IDs corresponding to the gene IDs in a data table, querying the other IDs corresponding to the gene IDs by retrieving the ID data table, and then querying the relevant data content based on the other IDs; Storing the association relationship, gene expression data, pathway data, and immune cell data to obtain a disease omics database; querying the three-dimensional data content of genes, pathways, and immune cells through the ID data table in the database; The conversion methods are GSVA and GSEA; The gene expression data is the GO dataset and the KEGG dataset; The method for mapping to immune cells is: cibersort; The gene expression library unique to the cell type is the LM22 dataset.

2. The construction method of the disease-specific omics database according to claim 1, wherein The method also includes data preprocessing, and the preprocessing includes one or more of the following: decentralization, data grouping, duplicate removal, ID matching, and standardization.

3. The construction method of the disease-specific omics database according to claim 2, characterized in that, The duplicate removal is to calculate the gene expression median according to the genes in the gene expression data, sort the medians, and then remove the duplicate gene data based on the sorted medians.

4. The construction method of the disease-specific omics database according to claim 2, wherein The standardization is performed based on z-score. Among them, the gene length in the expression data is calculated through the GenomicFeatures function, and the gene length is matched with the gene one by one; for genes that cannot match to obtain the gene length, the mean value is calculated by calculating the gene mean, and then the value obtained by dividing the mean value by the gene median is used as the gene length.

5. The construction method of the disease-specific omics database according to claim 1, characterized in that The method also includes annotation file arrangement. By obtaining disease pathway information, disease gene information, and disease immune cell information from the database, the disease pathway information, disease gene information, and disease immune cell information form an annotation file.

6. The construction method of the disease-specific omics database according to claim 5, wherein, The database for extracting annotation data of gene information is the GTF file.

7. The construction method of the disease-specific omics database according to claim 5, wherein The databases for extracting annotation data of pathway information include the msigdb database and the GO database.

8. The construction method of the disease-specific omics database according to claim 5, characterized in that, The database for extracting annotation data of immune cell information is cibersort.

9. The construction method of the disease-specific omics database according to claim 5, wherein, The method also includes retrieval file production. The retrieval file is associated with the annotation file, and the disease gene information, disease pathway information, and disease immune cell information are combined into one file to obtain the retrieval file.

10. A construction system for a disease-specific omics database, which has a computer program thereon, characterized in that, Including: A memory and a processor, where the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, the method for constructing the disease-specific omics database described in claims 1-9 is implemented.

11. A construction system for a disease-specific omics database, characterized in that, Including: The disease-specific omics database is obtained by executing the method for constructing the disease-specific omics database described in claims 1-9; A storage unit: storing the data of the disease-specific omics database; A query unit: querying data from the disease-specific omics database through a query instruction; An output unit: outputting data.

12. The construction system of the disease-specific omics database according to claim 11, wherein, The queried data includes one or more of the following: gene information of a disease, pathway information, immune cell information.

13. The construction system of the disease-specific omics database according to claim 11, characterized in that, The query instruction includes: retrieving by gene type or by pathway type or by immune cell type.

14. A method for obtaining disease-specific information based on a disease-specific database, characterized in that, Including: The method obtains disease-specific information from the disease-specific omics database construction system described in any one of claims 11-13. Among them, the query instruction is input to the query unit in the disease-specific omics database, and the disease-specific omics database calls data based on the query instruction and outputs it through the output unit.

15. The method for obtaining disease-specific information based on a disease-specific database according to claim 14, wherein The disease-specific omics database further includes a server and a client. The server, the client, and the disease-specific omics database are communicatively connected to each other. The client logs in to access the server and inputs a search instruction. The server accesses the disease-specific omics database based on the search instruction, and the disease-specific omics database obtains the search content through arithmetic processing and returns it to the client.

16. The method for obtaining specialized disease information based on a specialized disease database according to claim 14, wherein Performing data analysis on the disease-specific information obtained by query to obtain disease-specific analysis data.

17. The method for obtaining specialized disease information based on a specialized disease database according to claim 16, wherein Performing data visualization on the disease-specific analysis data to obtain disease-specific analysis visualization data.

18. The method for obtaining specialized disease information based on a specialized disease database according to claim 14, wherein Performing data visualization on the disease-specific information obtained by query to obtain disease-specific visualization data.

19. The method for obtaining specialized disease information based on a specialized disease database according to claim 14, wherein When obtaining the data of the disease-specific omics database, retrieving by gene type to obtain the corresponding gene data and the pathways and immune cell data corresponding to the gene, retrieving by pathway type to obtain the corresponding pathway data and the genes and immune cell data corresponding to the pathway, and retrieving by immune cell type to obtain the corresponding immune cell data and the genes and pathways data corresponding to the immune cell.

20. The method for obtaining specialized disease information based on a specialized disease database according to claim 19, wherein The way of obtaining the data of the disease-specific omics database further includes feature data retrieval, and obtaining the data that meets the features through feature screening in N dimensions, where N is a natural number greater than or equal to 1.

Citation Information

Patent Citations

  • Alzheimer's disease data processing method and system

    CN115472219A

  • Classification system and method based on tumor immune subtypes

    CN117854596A