A gene expression pattern discovery system and method based on information entropy
By acquiring and preprocessing cancer samples, mapping them to visible light and building a gene co-expression network, using information entropy to predict cancer-related lncRNAs, the problem of high cost in traditional methods is solved and efficient cancer association analysis is achieved.
Patent Information
- Application Number
- CN202311204385.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2043-09-18
AI Technical Summary
The prior art is difficult to efficiently discover the association between lncRNA and cancer through computational methods. Traditional biological experiments are costly and limited in scope, making it difficult to discover the expression patterns of special genes.
By obtaining disease and normal samples of multiple cancers, pretreatment and mapping, gene expression values are converted to visible light, a gene co-expression network is constructed, and lncRNAs related to cancer are predicted using information entropy.
Discrete and coarse grain of continuous gene expression data were achieved, and the gene expression patterns associated with cancer were discovered, which reduced research time and cost, and provided a new cancer research strategy.
Smart Images

Figure CN117198401B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of gene recognition, and in particular relates to a gene expression pattern discovery system and method based on information entropy. Background Art
[0002] In recent years, the continuous development of high-throughput sequencing has elevated our understanding of lncRNAs to a new level. Research has shown that lncRNAs are widely involved in physiological processes such as metabolism and immunity. Their dysregulated expression contributes to the occurrence and progression of various diseases, and they play a crucial role in life. Current research on lncRNA functions relies primarily on biological experiments, but this approach is costly and time-consuming, which has limited our understanding of lncRNAs. In this interdisciplinary context, computational bioinformatics has emerged, addressing the long timelines and high research costs of traditional biological experiments and providing significant support for in-depth research on lncRNAs. Therefore, computational methods are crucial for studying the relationship between lncRNAs and diseases, and are of great significance for understanding the pathogenesis of diseases, as well as for their diagnosis and treatment. However, the significant time and cost associated with traditional biological experiments has limited our understanding of lncRNAs. Therefore, computational methods for studying the relationship between lncRNAs and cancer can narrow the scope of biological experiments, focusing researchers' attention on lncRNAs that are more likely to be involved in cancer development, thereby reducing the time and cost of studying cancer-related lncRNAs.
[0003] Currently, there are two main approaches to studying the link between lncRNAs and cancer: the first is prediction through machine learning, which primarily uses the biological characteristics of lncRNAs and diseases to train classifiers to predict potential lncRNA-disease associations; the second is to predict lncRNA associations by analyzing the topological characteristics of biological networks. However, most of these methods simply use lncRNA-disease association data, PPI networks, gene-disease association data, lncRNA-protein-coding gene association data, and disease similarity data to predict lncRNA-disease associations, making it difficult to identify the expression patterns of specific genes. Summary of the Invention
[0004] To address the aforementioned issues in related technologies, the present invention provides a system and method for discovering gene expression patterns based on information entropy. The technical issues to be addressed by the present invention are achieved through the following technical solutions:
[0005] The present invention provides a gene expression pattern discovery system based on information entropy, comprising:
[0006] An acquisition module is used to acquire disease samples of multiple different cancers and normal samples of z1 types of cancer among the multiple different cancers; each sample contains expression values of multiple different genes and annotations of each gene; the multiple different genes contain multiple lncRNAs; z1 is a positive integer;
[0007] A preprocessing module, used for preprocessing the disease sample and the normal sample respectively;
[0008] a mapping module, configured to perform batch processing, outlier processing, mapping, and conversion on the pre-processed disease samples and the pre-processed normal samples to map the expression values onto visible light, thereby obtaining a disease gene expression spectrum for each cancer and a normal gene expression spectrum for each of the z1 types of cancer;
[0009] A construction module is used to obtain prior expression data of different human tissues and construct a gene co-expression network based on the prior expression data;
[0010] A prediction module is used to predict lncRNA associated with each cancer based on the disease gene expression spectrum, the normal gene expression spectrum and the gene co-expression network.
[0011] The present invention also provides a method for discovering gene expression patterns based on information entropy, comprising:
[0012] Obtain disease samples of multiple different cancers and normal samples of z1 types of cancer among the multiple different cancers; each sample contains expression values of multiple different genes and annotations of the different genes; the multiple different genes contain multiple lncRNAs; z1 is a positive integer;
[0013] Preprocessing the disease sample and the normal sample respectively;
[0014] performing batch processing, outlier processing, mapping, and conversion on the pre-processed disease samples and the pre-processed normal samples to map the expression values onto visible light, thereby obtaining a disease gene expression spectrum for each cancer and a normal gene expression spectrum for each of the z1 types of cancer;
[0015] Acquiring prior expression data of different human tissues, and constructing a gene co-expression network based on the prior expression data;
[0016] Based on the disease gene expression spectrum, the normal gene expression spectrum and the gene co-expression network, lncRNA associated with each cancer is predicted.
[0017] The present invention has the following beneficial technical effects:
[0018] 1) This paper relies on the characteristics of high-performance computer computing and proposes a new analysis model for lncRNA-disease association by converting gene expression data into visible light.
[0019] 2) The present invention maps continuous gene expression values onto visible light, achieving discretization and coarse-graining of continuous gene expression data, and constructing a gene expression spectrum. Based on the calculation and analysis of genes in the constructed gene expression spectrum, relevant gene expression patterns, including cancer-associated genes, were discovered, providing a method for studying the occurrence of cancer from the perspective of biological processes.
[0020] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A structural diagram of a gene expression pattern discovery system based on information entropy provided by an embodiment of the present invention;
[0022] Figure 2 A flowchart of a gene expression pattern discovery method based on information entropy provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.
[0024] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0025] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.
[0026] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0027] This invention aims to screen for cancer-associated lncRNAs (long noncoding RNAs). Unlike cancer diagnosis, this invention aims to discover the association between specific lncRNAs and cancer development and progression. By identifying these lncRNAs, we can better understand the molecular mechanisms of cancer and further provide new targets and strategies for cancer treatment. This invention does not involve any technology or method related to cancer diagnosis. Instead, it focuses on screening for cancer-associated lncRNAs in the hope of contributing to advances in cancer research.
[0028] Figure 1 This is a structural diagram of a gene expression pattern discovery system based on information entropy provided by an embodiment of the present invention, such as Figure 1 As shown, the system includes: an acquisition module 10, a preprocessing module 20, a mapping module 30, a construction module 40 and a prediction module 50.
[0029] The acquisition module 10 is used to obtain disease samples of multiple different cancers and normal samples of z1 types of cancer among the multiple different cancers; each sample contains expression values of multiple different genes and annotations of each gene; the multiple different genes contain multiple lncRNAs; z1 is a positive integer.
[0030] Specifically, the acquisition module 10 is used to obtain original disease samples of multiple different cancers, and when z1 types of cancer among the multiple different cancers have original normal samples, obtain the original normal samples of each cancer in the z1 types of cancer; download the gene annotation file from the gene annotation database, and annotate the genes in the original disease samples and the original normal samples respectively to obtain disease samples and normal samples; each sample contains the expression value of each gene in n genes; each cancer has m1 disease samples, and the value of m1 corresponding to different cancers is different; each cancer in the z1 types of cancer has m2 normal samples, and the value of m2 corresponding to different cancers is different. For example, you can download disease RNA-seq data and clinical data for 33 cancers from the TCGA database. When 10 cancers have corresponding normal RNA-seq data, download the normal RNA-seq data for these 10 cancers. Each cancer corresponds to a set of disease RNA-seq data, and each set of disease RNA-seq data corresponds to m1 disease samples, each disease sample contains clinical information, and the clinical information of a sample includes the clinical data corresponding to the sample (i.e., the patient information corresponding to the sample) and the batch information of the sample, or only the batch information of the sample; and a set of normal RNA-seq data corresponds to m2 normal samples, each disease sample or normal sample contains a total of 60,483 genes. Download the gene annotation file from the Gencode database, and annotate the genes in each sample according to the gene annotation file.
[0031] It should be noted that samples corresponding to the same clinical data but belonging to different batches of information are obtained by analyzing the same human tissue of the same patient in different analytical environments (for example, using different experimental methods and / or different experimental times, different laboratory environments, etc.).
[0032] The preprocessing module 20 is used to preprocess disease samples and normal samples respectively.
[0033] Specifically, the preprocessing module 20 is used to construct a disease sample matrix of n rows and m1 columns for each cancer based on the disease samples of each cancer, the expression values of the genes contained in each sample, and the annotations of the genes, and when the cancer has normal samples, construct a normal sample matrix of n rows and m2 columns for the cancer; delete the genes with expression values of 0 in A% of the samples in the disease sample matrix of the cancer, and delete the genes repeated on the Y chromosome in the disease sample matrix of the cancer according to the annotations to obtain the disease sample matrix of the cancer after the value is deleted; when the cancer has a normal sample matrix, delete the genes with expression values of 0 in A% of the samples in the normal sample matrix of the cancer Genes with expression values of 0 are deleted, and genes repeated on the Y chromosome in the normal sample matrix of the cancer are deleted according to annotations to obtain a normal sample matrix after value deletion of the cancer; A is a positive integer greater than 80; when the data format of the disease sample matrix after value deletion and the normal sample matrix after value deletion of the cancer is Counts data, the data format of the disease sample matrix after value deletion and the normal sample matrix after value deletion of the cancer is standardized to FPKM data; the FPKM data is logarithmized to make the matrix conform to the normal distribution, to obtain a pre-processed disease matrix and a pre-processed normal matrix of the cancer. For example, for each of the 33 types of cancer mentioned above, m1 disease samples of the cancer are represented by a matrix, and then a matrix of 60483 rows and m1 columns for the cancer can be obtained. Then, the genes that are not expressed in 90% (for example, A=90) of the samples in the matrix of 60483 rows and m1 columns are deleted, and the genes that are repeated on the Y chromosome in the Gencode gene annotation file in the matrix of 60483 rows and m1 columns are deleted. When the data format of the matrix after deleting some genes is Counts data, the data format of the matrix is standardized to FPKM data to remove the technical deviation of the sequencing data and eliminate the influence of sequencing depth and gene length on the data. Then, the FPKM data is logarithmized, and the expression value of the FPKM data is changed to Log2(FPKM+1) so that the matrix conforms to the normal distribution. In this way, the preprocessed disease matrix of the matrix is obtained. For each of the 33 cancer types mentioned above, when the cancer has m2 normal samples, the m2 normal samples of the cancer can be represented by a matrix, and a 60483-row and m2-column matrix can be obtained for the cancer. Then, the matrix is preprocessed using the same processing principle as above to obtain the preprocessed normal matrix for the cancer. Here, the expression of FPKM data is: Where N represents the number of genes in the matrix, C g Indicates the number of reads aligned to gene g, L g represents the sum of the lengths of all exons of gene g.
[0034] The mapping module 30 is used to perform batch processing, outlier processing, mapping and conversion on the pre-processed disease samples and pre-processed normal samples to map the expression values to visible light, thereby obtaining the disease gene expression spectrum of each cancer and the normal gene expression spectrum of each cancer in z1 types of cancer.
[0035] Specifically, the preprocessed disease samples for each cancer form a preprocessed disease matrix with n1 rows and m1 columns, and the preprocessed normal samples for each of the z1 cancer types form a preprocessed normal matrix with n2 rows and m1 columns. n, m1, and m2 are all positive integers, and n1 and n2 are less than or equal to n. Mapping module 30 is specifically configured to perform batch processing, outlier processing, and expression value mapping and conversion on the preprocessed disease matrix for each cancer to obtain a disease gene expression spectrum for that cancer. If the cancer has a preprocessed normal matrix, batch processing, outlier processing, and expression value mapping and conversion are performed on the preprocessed normal matrix to obtain a normal gene expression spectrum for that cancer.
[0036] Specifically, among the m1 samples of each of the z2 types of cancer among the above-mentioned multiple cancers, there are samples from different batches but belonging to the same human tissue; z2 is a positive integer; based on this, the pre-processed disease matrix of each cancer is batch processed, outlier processed, and expression value mapped and converted to obtain a disease gene expression spectrum of the cancer. The principle is: in the pre-processed disease matrix of each of the z2 types of cancer, the samples that do not contain clinical data are deleted to obtain the effective disease matrix of each of the z2 types of cancer; the ComBat method is used to batch process the effective disease matrix of each of the z2 types of cancer to obtain the corrected disease matrix of each of the z2 types of cancer; according to the corrected disease matrix of each of the z2 types of cancer, and the disease matrix of multiple different cancers except z2 types, For each cancer type other than z2, the pre-processed disease matrix is used to determine positive and negative outliers for each cancer type, as well as positive and negative outliers for each cancer type other than z2, in the multiple different cancers. Based on the corresponding positive and negative outliers, expression values outside the interval defined by the positive and negative outliers are deleted from the corrected disease matrix for each cancer type and the pre-processed disease matrix for each cancer type other than z2, thereby obtaining a reasonable disease matrix for each cancer type, as well as a reasonable disease matrix for each cancer type other than z2. Pre-set mapping and transformation formulas are used to map the reasonable disease matrix for each cancer type to multiple visible light sources, thereby obtaining a disease gene expression spectrum for each cancer type. It should be noted that when the batch information for each sample is long, the samples can be renumbered according to the order of the batch information to facilitate subsequent calculations. However, when the batch information for each sample is short, renumbering is not required.
[0037] For example, when the disease samples of 25 of the 33 cancers mentioned above come from different batches, the samples that do not contain clinical data in the pre-processed disease matrix of each of the 25 cancers are deleted to obtain the effective disease matrix of each of the 25 cancers. Then, the ComBat method is used to batch process the effective disease matrix of each of the 25 cancers to obtain the corrected disease matrix of each of the 25 cancers. Then, the expression values in the corrected disease matrices of the 25 cancers are observed, and the pre-processed disease matrices of the 8 cancers other than the 25 cancers among the 33 cancers are observed, and the negative values of the corrected disease matrices of the 25 cancers are respectively calculated. The cluster value is set to 0.0001, the positive outlier value is set to 0.9999, and the negative outlier value of the pre-processed disease matrix of the 8 cancers is set to 0.0001, and the positive outlier value is set to 0.9995; then, the expression values outside the interval [0.0001, 0.9999] in the corrected disease matrix of each of the 25 cancers are deleted to obtain the reasonable disease matrix of each of the 25 cancers, and the expression values outside the interval [0, 0.9995] in the pre-processed disease matrix of each of the 8 cancers are deleted to obtain the reasonable disease matrix of each of the 8 cancers; finally, for a reasonable disease matrix of each cancer, the formula WL= By mapping the reasonable disease matrix to 7 kinds of visible light and using 1 to 7 to encode these 7 kinds of visible light, a color matrix of the cancer is obtained. The color matrix is a disease gene expression spectrum of the cancer. In this case, E represents each expression value in the reasonable disease matrix, and E max Indicates the maximum expression value in the reasonable disease matrix where E is located, E min represents the minimum expression value in the reasonable disease matrix where E is located, WL represents the mapped wavelength of E, 780 and 380 represent the maximum and minimum wavelengths of visible light, respectively. The wavelength range of red light is 620nm-780nm, the wavelength range of orange light is 590nm-620nm, the wavelength range of yellow light is 560nm-590nm, the wavelength range of green light is 490nm-560nm, the wavelength range of blue light is 450nm-490nm, the wavelength range of indigo light is 420nm-450nm, and the wavelength range of violet light is 380nm-420nm.
[0038] It should be noted that the principle of performing batch processing, outlier processing, and expression value mapping and conversion on the preprocessed disease matrix of each cancer to obtain a disease gene expression spectrum of the cancer is the same as the principle of performing batch processing, outlier processing, and expression value mapping and conversion on the preprocessed normal matrix of each cancer to obtain a normal gene expression spectrum of the cancer.
[0039] The construction module 40 is used to obtain prior expression data of different human tissues and construct a gene co-expression network based on the prior expression data.
[0040] Specifically, the construction module 40 is used to construct a matrix with q rows and p columns; preprocess the matrix with q rows and p columns to obtain a preprocessed matrix; calculate the Pearson correlation coefficient of the expression values of all genes in the preprocessed matrix; for all genes in the preprocessed matrix, when the Pearson correlation coefficient between the expression values of two different genes is greater than or equal to a preset threshold, it indicates that there is a correlation between the two different genes, and the two different genes are used as two nodes in the gene co-expression network, and the two nodes are connected by a line to obtain an edge in the gene co-expression network. After traversing all genes in the preprocessed matrix, a gene co-expression network containing multiple nodes and multiple edges is obtained. For example, 11,688 samples of different human tissues can be obtained from the GTEx database, each containing a total of 56,202 genes, so a matrix with 56,202 rows and 11,688 columns can be constructed. Then, genes with expression values of 0 or constant values in at least 80% of the samples in the matrix with 56,202 rows and 11,688 columns are deleted, resulting in a matrix with 34,923 rows and 11,688 columns. Then, the formula is used Calculate the Pearson correlation coefficient between each two different genes in the matrix of 34923 rows and 11688 columns, where X represents one of the two different genes, σ X represents the standard deviation of the row where X is located, Y represents the other gene in two different genes, σ Y Indicates the standard deviation of the row where Y is located, P X,Y represents the Pearson correlation coefficient between X and Y; when the threshold is 0.74, the difference in clustering coefficient between the gene co-expression network and the random network is the largest, so 0.74 is selected as the threshold. When the P value between the expression values of two different genes in the matrix of 34923 rows and 11688 columns is X,Y When it is greater than or equal to 0.74, it indicates that there is an association between the two different genes. The two different genes are regarded as two nodes in the gene co-expression network, and the two nodes are connected by a line to obtain an edge in the gene co-expression network. After traversing all genes in the matrix of 34923 rows and 11688 columns, a gene co-expression network consisting of 20419 gene nodes and 5383167 gene edges is obtained, and 2932 lncRNAs are involved in these 20419 genes.
[0041] The prediction module 50 is used to predict the lncRNA associated with each cancer based on the disease gene expression spectrum, the normal gene expression spectrum and the gene co-expression network.
[0042] Specifically, the prediction module 50 is used to calculate the information entropy of each gene in the disease gene expression spectrum for each cancer, and when the cancer has a normal gene expression spectrum, calculate the information entropy of each gene in the normal gene expression spectrum; when the cancer has a disease gene expression spectrum and a normal gene expression spectrum, according to the prior characteristics of the gene distribution of the cancer, a plurality of genes are screened from the disease gene expression spectrum and the normal gene expression spectrum respectively, and the selected genes are intersected, and the genes belonging to the intersection are removed from the plurality of genes screened from the disease gene expression spectrum to obtain the remaining genes of the cancer; or, when the cancer does not have a normal gene expression spectrum, the prediction module 50 is used to calculate the information entropy of each gene in the disease gene expression spectrum; When analyzing the expression spectrum, multiple genes selected from the disease gene expression spectrum are used as the remaining genes for the cancer; a subnetwork containing the remaining genes for the cancer is searched from the gene co-expression network; the subnetwork is mined to obtain multiple gene blocks within the subnetwork; the central node of each gene block is the hub gene of the gene block; the biological function of each gene block for the cancer is analyzed, and lncRNAs associated with the hub gene of the gene block are searched within the gene block; the correlation between the biological function of the gene block and the cancer is determined based on prior knowledge, and the correlation between the lncRNAs associated with the hub gene of the gene block and the cancer is predicted based on the correlation. Specifically, the prediction module 50 is configured to, for each gene block for each cancer, identify the lncRNA associated with the hub gene of the gene block as the lncRNA associated with the cancer when the biological function of the gene block is related to the cancer.
[0043] For example, taking breast cancer as an example to illustrate the specific functions of the prediction module, the distribution of all genes in breast cancer is similar to a power law distribution, showing a "top-heavy" characteristic. The information entropy of most genes is less than 1. Therefore, according to the 80 / 20 rule, the genes with the top 20% information entropy are screened from the disease gene expression spectrum of breast cancer as a group of genes with large information entropy. When breast cancer has a normal gene expression spectrum, the genes with the top 20% information entropy are screened from the normal gene expression spectrum of breast cancer as another group of genes with large information entropy. Afterwards, these two groups of genes with large information entropy are combined. The intersection was taken, and the genes belonging to the intersection were removed from the top 20% of the information entropy genes screened from the disease gene expression spectrum of breast cancer to obtain the remaining genes of breast cancer (it should be noted that when breast cancer does not have a normal gene expression spectrum, the genes ranked in the top 20% of the information entropy from the disease gene expression spectrum of breast cancer are used as the remaining genes of breast cancer). There are a total of 2784 genes remaining in breast cancer. Then, from the gene co-expression network composed of 20419 gene nodes and 5383167 gene edges, the 2784 genes were found. The subnetwork of 4 genes was constructed and found to be composed of 1253 gene nodes and 13958 gene edges, and these 1253 gene nodes contained 77 lncRNAs. Then, the CNM community detection algorithm based on modularity was used to mine the subnetwork, and multiple gene blocks of breast cancer were obtained. These multiple gene blocks were the gene blocks with the most compact structure in the subnetwork. The central node of each gene block was called the hub gene of the gene block. GO function enrichment analysis and KEGG pathway enrichment analysis were performed on each gene block of breast cancer to obtain the biological function of the gene block of breast cancer. The method can determine the function of a gene block (for example, whether it is involved in metabolism, etc.), and search for lncRNAs with genetic links to the hub gene of the gene block from the gene block. Then, based on existing databases and existing literature records, determine whether the biological function of the gene block is related to breast cancer. When the biological function of the gene block is related to breast cancer, the lncRNAs with genetic links to the hub gene of the gene block are lncRNAs related to breast cancer. When the biological function of the gene block is not related to breast cancer, the lncRNAs with genetic links to the hub gene of the gene block are lncRNAs not related to breast cancer.
[0044] The present invention also provides a method for discovering gene expression patterns based on information entropy, which is used to execute the content of the method corresponding to the above-mentioned method for discovering gene expression patterns based on information entropy. Figure 2 As shown, the gene expression pattern discovery method based on information entropy includes:
[0045] S101. Obtain disease samples of multiple different cancers and normal samples of z1 types of cancer among the multiple different cancers; each sample contains expression values of multiple different genes and annotations of different genes; the multiple different genes contain multiple lncRNAs; z1 is a positive integer.
[0046] S102. Preprocess the disease samples and normal samples respectively.
[0047] S103. Batch processing, outlier processing, mapping, and conversion are performed on the pretreated disease samples and the pretreated normal samples to map the expression values onto visible light, thereby obtaining a disease gene expression spectrum for each cancer and a normal gene expression spectrum for each of the z1 cancer types.
[0048] S104. Obtain prior expression data of different human tissues, and construct a gene co-expression network based on the prior expression data.
[0049] S105. Predict lncRNAs associated with each cancer based on disease gene expression spectra, normal gene expression spectra, and gene co-expression networks.
[0050] This invention provides a method for discovering gene expression patterns based on information entropy. Leveraging the power of high-performance computing, this method discretizes and coarse-grains continuous gene expression data by mapping continuous gene expression values onto seven visible light spectrums through scaling, thereby constructing gene expression spectra. By calculating and analyzing the information entropy of genes in gene expression spectra, relevant gene expression patterns, including those associated with cancer, were discovered. This invention proposes a novel analytical model for lncRNA-disease associations, reducing the time and cost required to study cancer-related lncRNAs.
[0051] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A gene expression pattern discovery system based on information entropy, characterized in that: include: An acquisition module is configured to acquire disease samples of multiple different cancers and normal samples of z1 types of cancer among the multiple different cancers; each sample comprises expression values of multiple different genes and annotations of each gene; the multiple different genes comprise multiple lncRNAs; z1 is a positive integer; each sample comprises expression values of each gene among n genes; each cancer has m1 disease samples; each of the z1 types of cancer has m2 normal samples; a preprocessing module, configured to preprocess the disease samples and the normal samples respectively; the preprocessed disease samples of each cancer are a preprocessed disease matrix with n1 rows and m1 columns, and the preprocessed normal samples of each of the z1 cancer types are a preprocessed normal matrix with n2 rows and m1 columns; n, m1 and m2 are all positive integers, n1 and n2 are less than or equal to n; A mapping module is configured to perform batch processing and outlier processing on the preprocessed disease matrix of each cancer to obtain a reasonable disease matrix for the cancer, and map the reasonable disease matrix of the cancer to multiple visible lights using a mapping and conversion formula to obtain a disease gene expression spectrum for the cancer; when the cancer has the preprocessed normal matrix, batch processing and outlier processing are performed on the preprocessed normal matrix to obtain a reasonable normal matrix for the cancer, and map the reasonable normal matrix of the cancer to multiple visible lights using a mapping and conversion formula to obtain a normal gene expression spectrum for the cancer; wherein the mapping and conversion formula is expressed as follows: ,in, represents each expression value in each reasonable disease matrix / each reasonable normal matrix, express The maximum expression value in the reasonable disease matrix / reasonable normal matrix, express The minimum expression value in the reasonable disease matrix / reasonable normal matrix, express After mapping, 780 and 380 represent the maximum and minimum wavelengths of visible light, respectively; A construction module is used to obtain prior expression data of different human tissues and construct a gene co-expression network based on the prior expression data; A prediction module is used to calculate the information entropy of each gene in the disease gene expression spectrum of each cancer, and when the cancer has the normal gene expression spectrum, calculate the information entropy of each gene in the normal gene expression spectrum; based on the information entropy of each gene in the disease gene expression spectrum, the information entropy of each gene in the normal gene expression spectrum and the gene co-expression network, predict the lncRNA associated with each cancer.
2. The gene expression pattern discovery system based on information entropy according to claim 1, characterized in that: Among the m1 samples of each of z2 types of cancer among the plurality of different cancers, there are samples from different batches but belonging to the same human tissue; z2 is a positive integer; and the mapping module is further configured to: Deleting samples that do not contain clinical data from the pre-processed disease matrix for each of the z2 types of cancer to obtain a valid disease matrix for each of the z2 types of cancer; Using the ComBat method to batch process the effective disease matrix of each of the z2 cancers to obtain a corrected disease matrix for each of the z2 cancers; determining, based on the corrected disease matrix for each of the z2 cancers and the pre-treated disease matrix for each of the plurality of different cancers excluding the z2 cancers, positive and negative outliers for each of the plurality of different cancers excluding the z2 cancers; According to the corresponding positive and negative outliers, the expression values outside the interval formed by the positive and negative outliers in the corrected disease matrix of each cancer in the z2 types of cancer and the pre-treated disease matrix of each cancer in the multiple different cancers except the z2 types of cancer are deleted accordingly to obtain a reasonable disease matrix for each cancer in the z2 types of cancer and a reasonable disease matrix for each cancer in the multiple different cancers except the z2 types of cancer.
3. The gene expression pattern discovery system based on information entropy according to claim 1, characterized in that: The prior expression data of different human tissues include: P different samples, and each of the P different samples contains q genes of human tissues; the construction module is further used to: Construct a matrix with q rows and p columns; Preprocessing the matrix of q rows and p columns to obtain a preprocessing matrix; Calculate the Pearson correlation coefficient of the expression values of all genes in the preprocessing matrix; For all genes in the preprocessing matrix, when the Pearson correlation coefficient between the expression values of two different genes is greater than or equal to a preset threshold, it indicates that there is a correlation between the two different genes. The two different genes are used as two nodes in the gene co-expression network, and the two nodes are connected by a line to obtain an edge in the gene co-expression network. After traversing all genes in the preprocessing matrix, a gene co-expression network containing multiple nodes and multiple edges is obtained.
4. The gene expression pattern discovery system based on information entropy according to claim 1, characterized in that: The prediction module is further used to: When the cancer has the disease gene expression spectrum and the normal gene expression spectrum, a plurality of genes are screened from the disease gene expression spectrum and the normal gene expression spectrum according to a priori characteristics of gene distribution of the cancer, and an intersection is taken for the selected genes, and genes belonging to the intersection are removed from the plurality of genes screened from the disease gene expression spectrum to obtain remaining genes for the cancer; Alternatively, when the cancer does not have the normal gene expression spectrum, the plurality of genes screened from the disease gene expression spectrum are used as the remaining genes of the cancer; Searching for a sub-network containing the remaining genes of the cancer from the gene co-expression network; Mining the sub-network to obtain multiple gene blocks of the sub-network; the central node of each gene block is the hub gene of the gene block; Analyze the biological function of each gene block in the cancer and search for lncRNAs associated with the hub genes in the gene block; The correlation between the biological function of the gene block and the cancer is determined based on prior knowledge, and the correlation between the lncRNA associated with the hub gene of the gene block and the cancer is predicted based on the correlation.
5. The gene expression pattern discovery system based on information entropy according to claim 4, characterized in that: The prediction module is further used to: For each gene block of each cancer, when the biological function of the gene block is related to the cancer, the lncRNA associated with the hub gene of the gene block is regarded as the lncRNA associated with the cancer.
6. The gene expression pattern discovery system based on information entropy according to claim 1, characterized in that: The acquisition module is further used to: Obtaining original disease samples of multiple different cancers, and when z1 types of cancer among the multiple different cancers have original normal samples, obtaining original normal samples of each of the z1 types of cancer; The gene annotation file is downloaded from the gene annotation database, and the genes in the original disease sample and the original normal sample are annotated respectively to obtain the disease sample and the normal sample.
7. The gene expression pattern discovery system based on information entropy according to claim 1, characterized in that: The preprocessing module is further used for: Based on the disease samples of each cancer, the expression values of the genes contained in each sample, and the gene annotations, a disease sample matrix with n rows and m1 columns for the cancer is constructed. If the cancer has normal samples, a normal sample matrix with n rows and m2 columns for the cancer is constructed. Deleting genes whose expression values are 0 in A% of samples from the disease sample matrix of the cancer, and deleting genes that are repeated on chromosome Y in the disease sample matrix of the cancer according to the annotation, to obtain a disease sample matrix after value deletion of the cancer; when the cancer has a normal sample matrix, deleting genes whose expression values are 0 in A% of samples from the normal sample matrix of the cancer, and deleting genes that are repeated on chromosome Y in the normal sample matrix of the cancer according to the annotation, to obtain a normal sample matrix after value deletion of the cancer; A is a positive integer greater than 80; When the data format of the cancer's disease sample matrix after value deletion and the normal sample matrix after value deletion is Counts data, the data format of the cancer's disease sample matrix after value deletion and the normal sample matrix after value deletion is standardized to FPKM data; The FPKM data are logarithmized to make the matrix conform to normal distribution, thereby obtaining a pre-treatment disease matrix of the cancer and a pre-treatment normal matrix of the cancer.
8. A gene expression pattern discovery method based on information entropy, applied to the gene expression pattern discovery system based on information entropy according to any one of claims 1 to 7, characterized in that: include: Obtain disease samples of multiple different cancers and normal samples of z1 types of cancer among the multiple different cancers; each sample contains expression values of multiple different genes and annotations of the different genes; the multiple different genes contain multiple lncRNAs; z1 is a positive integer; Preprocessing the disease sample and the normal sample respectively; performing batch processing, outlier processing, mapping, and conversion on the pre-processed disease samples and the pre-processed normal samples to map the expression values onto visible light, thereby obtaining a disease gene expression spectrum for each cancer and a normal gene expression spectrum for each of the z1 types of cancer; Acquiring prior expression data of different human tissues, and constructing a gene co-expression network based on the prior expression data; Based on the disease gene expression spectrum, the normal gene expression spectrum and the gene co-expression network, lncRNA associated with each cancer is predicted.
Citation Information
Patent Citations
Hybrid network gene screening method based on gene expression data
CN106055922A
Spectral light splitting node determination method and related equipment
CN115683336A