Method and system for mining potential proto-oncogenes based on chromatin three-dimensional structure
Through a method based on the three-dimensional structure of chromatin, CTCF binding site predictor and deep learning model are used to combine gene expression data to screen out damaged insulating regions and differentially expressed genes, solving the problems of high time and money costs in the existing technology, and achieving efficient and accurate screening of potential proto-oncogenes.
Patent Information
- Application Number
- CN202311017259.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-11
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-08-11
AI Technical Summary
The prior art consumes a lot of time and money when mining multiple gene mutations, and the gene differential expression analysis method is inefficient, making it difficult to accurately screen out potential proto-oncogenes.
Through a method based on the three-dimensional structure of chromatin, CTCF binding site predictor and deep learning model were used to combine gene expression data to screen out damaged insulating regions and differentially expressed genes, perform survival analysis and gene set enrichment, and select genes related to adverse prognosis.
On the basis of saving time and economic costs, the accuracy and reliability of potential oncogene mining are improved, and the functionally altered genes are accurately identified, reducing human and material consumption.
Smart Images

Figure CN116935962B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computational biology technology, and in particular to a system and method for mining potential proto-oncogenes based on the three-dimensional structure of chromatin. Background Art
[0002] In human eukaryotic cells, chromosomes are highly folded and organized into dynamic three-dimensional (3D) structures. This complex and precise folding of chromosomes within the cell nucleus ensures correct gene expression and replication. In 3D genomics, insulating regions are defined as chromatin loops formed by CTCF-CTCF homodimers, co-bound to cohesin, and containing at least one gene.
[0003] Mutations in CTCF binding sites frequently occur in cancer, and some variants in these sites can weaken CTCF binding, leading to the loss of insulating regions. This can then lead to local gene dysregulation due to inappropriate enhancer-promoter interactions. Mutated insulating regions have been found to be significantly enriched for known oncogenes. These silenced oncogenes are often located within insulating regions, and the loss of insulating regions can activate these genes, causing cancer. Disruption of insulating regions is a driver of cancer.
[0004] By comparing gene expression differences between cancerous and normal tissues, differential gene expression analysis can identify differentially expressed genes that may play a key role in the development and progression of cancer. Differential gene expression analysis can reveal differences in gene expression between cancerous and normal tissues, which may be closely related to the pathophysiology of cancer. Bioinformatics analysis of these differentially expressed genes can reveal important signaling pathways, regulatory factors, and target genes involved in cancer development and progression.
[0005] Deep learning is a learning method based on multi-layer neural networks that can automatically learn representations and features from data and perform high-level abstraction and analysis. Deep learning is an effective method for sequence prediction. In sequence prediction tasks, it can effectively capture the complexity and temporal relationships of sequence data, improving prediction accuracy and generalization capabilities.
[0006] Current technological advances have made it possible to identify insulating regions in the human genome. We can identify insulating regions using CTCF ChIA-PET data and RAD21 (a cohesin-forming protein) ChIP-seq data. We can also identify CTCF binding sites from CTCF ChIP-seq data and construct and train a CTCF binding site predictor.
[0007] Currently, only a few mutations that disrupt CTCF binding site activity have been reported through biological experiments, leaving a vast number of mutations to be investigated. This approach, when faced with tens of thousands of mutations, often consumes significant time and money.
[0008] Currently, proto-oncogene discovery is typically performed through differential gene expression analysis. This method relies on identifying genes based on whether their expression differs significantly between normal and cancerous samples. However, this method typically yields hundreds or even thousands of genes, requiring significant human and material resources to analyze each gene individually before ultimately identifying the proto-oncogene. Summary of the Invention
[0009] The present invention solves the problem of consuming a large amount of time and money costs in the prior art when mining multiple gene mutations by providing a method and system for mining potential proto-oncogenes based on the three-dimensional structure of chromatin. It achieves the goal of improving the accuracy and reliability of mining potential proto-oncogenes while saving time and economic expenses.
[0010] In a first aspect, an embodiment of the present invention provides a method for mining potential proto-oncogenes based on the three-dimensional structure of chromatin, the method comprising:
[0011] Obtaining multiple mutation-insulating regions based on chromatin data and cancer mutation data, and inputting the multiple mutation-insulating regions into a trained binding site predictor to obtain prediction results;
[0012] using the prediction result to determine whether the sudden insulating region is a sudden insulating region destroyed by the mutation, thereby obtaining a destroyed insulating region;
[0013] Obtaining multiple differentially expressed gene sets according to the cancer gene expression data, and taking the intersection of the multiple differentially expressed gene sets to obtain a final differentially expressed gene set;
[0014] Taking the intersection of the final differentially expressed gene set and the genes in the destroyed insulation region to obtain intersection genes;
[0015] Performing survival analysis on the intersection genes to obtain analysis results, and screening the intersection genes according to the analysis results to obtain genes associated with poor prognosis;
[0016] Gene set enrichment is used to screen the genes associated with poor prognosis to obtain potential oncogenes.
[0017] In combination with the first aspect, in one possible implementation, the chromatin data includes chromatin three-dimensional interaction data and genomics data, wherein the chromatin three-dimensional interaction data is the number of long-range interactions of CTCF in the genome; and the genomics data is the binding site distribution data of CTCF protein in the genome.
[0018] In combination with the first aspect, in one possible implementation, the cancer mutation data includes somatic mutation data of multiple cancer types; and the cancer gene expression data includes gene expression profiles of cancer types.
[0019] In conjunction with the first aspect, in a possible implementation, obtaining multiple mutation-insulating regions based on chromatin data and cancer mutation data specifically includes:
[0020] Extracting the first to fifth columns of the CTCF CHIA-PET data in the chromatin data to obtain the start and end positions of two genomic regions of the interaction mediated by CTCF on each chromosome;
[0021] Extracting the first, second, and third columns of the RAD21 CHIP-seq data from the chromatin data to obtain the starting and ending positions of the genome bound by RAD21 on each chromosome;
[0022] The start and end positions of the two genomic regions interacting via CTCF were combined with the start and end positions of the genome bound by RAD21 to generate multiple mutation-insulating regions.
[0023] In combination with the first aspect, in a possible implementation, the binding site predictor includes: an embedding layer, a one-dimensional convolutional layer, a maximum pooling layer, a bidirectional gated recurrent neural network layer and a fully connected layer, and the layers in the binding site predictor are connected in sequence.
[0024] In conjunction with the first aspect, in a possible implementation, before inputting the multiple mutant insulating regions into the trained binding site predictor, the method further includes:
[0025] Obtaining the first, second, and third columns of CTCF CHIP-seq data of the K562 cell line, the HepG2 cell line, and the GM12878 cell line, respectively, to obtain a plurality of screening data, and performing an intersection operation on the plurality of screening data to obtain conserved CTCF CHIP-seq data of the cell lines;
[0026] Analyze the tenth column of the conserved CTCF CHIP-seq data of the cell line to obtain the position with the strongest interaction in each data;
[0027] intercepting a portion of data centered on the position with the strongest interaction as positive sample data for the binding site predictor;
[0028] By using the R package gkmSVM to control the sequence length, GC content and repetitive sequence score characteristics of the positive sample data, negative sample data corresponding to each positive sample data is obtained;
[0029] The binding site predictor is trained using the positive sample data and the negative sample data to obtain the trained binding site predictor.
[0030] In conjunction with the first aspect, in one possible implementation, inputting the multiple mutant insulating regions into a trained binding site predictor to obtain prediction results specifically includes:
[0031] Obtaining a mutant insulating region with a CTCF binding site from the multiple mutant insulating regions, to obtain multiple mutant insulating regions with the CTCF binding site;
[0032] The plurality of mutant insulating regions with CTCF binding sites are respectively input into the binding site predictor for prediction to obtain a plurality of prediction results.
[0033] With reference to the first aspect, in one possible implementation, determining whether the sudden insulating region is a sudden insulating region destroyed by a mutation using the prediction result to obtain the destroyed insulating region specifically includes:
[0034] Obtaining the prediction result, and determining whether the prediction result is a combination;
[0035] If the prediction result is binding, determining the mutant insulating region with the CTCF binding site as an insulating region that will not be destroyed;
[0036] If the prediction result is no binding, the mutant insulating region with the CTCF binding site is determined to be a destroyed insulating region.
[0037] In conjunction with the first aspect, in one possible implementation, obtaining multiple differentially expressed gene sets based on cancer gene expression data specifically includes:
[0038] The cancer gene expression data are analyzed using the limma gene differential expression analysis package, the DESeq2 gene differential expression analysis package, or the edgeR gene differential expression analysis package to obtain multiple differentially expressed gene sets.
[0039] In a second aspect, the present invention provides a system for mining potential proto-oncogenes based on the three-dimensional structure of chromatin, the system comprising:
[0040] A prediction module is used to obtain multiple mutation-insulating regions based on chromatin data and cancer mutation data, and input the multiple mutation-insulating regions into a trained binding site predictor to obtain prediction results;
[0041] a region acquisition module, configured to use the prediction result to determine whether the sudden insulation region is a sudden insulation region destroyed by the mutation, and obtain the destroyed insulation region;
[0042] A differentially expressed gene acquisition module is used to obtain multiple differentially expressed gene sets based on cancer gene expression data, and to obtain a final differentially expressed gene set by taking the intersection of the multiple differentially expressed gene sets;
[0043] An intersection gene acquisition module, configured to obtain intersection genes by intersecting the final differentially expressed gene set and the genes in the destroyed insulation region;
[0044] A gene screening module, configured to perform survival analysis on the intersection genes to obtain analysis results, and screen the intersection genes based on the analysis results to obtain genes associated with poor prognosis;
[0045] The proto-oncogene acquisition module is used to screen the genes associated with poor prognosis using gene set enrichment to obtain potential proto-oncogenes.
[0046] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:
[0047] The present invention adopts a method for mining potential proto-oncogenes based on the three-dimensional structure of chromatin, which includes: obtaining multiple mutation insulation regions based on chromatin data and cancer mutation data, inputting the multiple mutation insulation regions into a trained binding site predictor to obtain prediction results, and relying on the characteristics of high-performance computer computing, greatly saving time costs and economic expenses; using the prediction results to determine whether the mutation insulation region is a mutation insulation region destroyed by mutation, and obtaining the destroyed insulation region; obtaining multiple differentially expressed gene sets based on cancer gene expression data, and taking the intersection of the multiple differentially expressed gene sets to obtain a final differentially expressed gene set; taking the intersection of the final differentially expressed gene set and the genes in the destroyed insulation region to obtain intersection genes; performing survival analysis on the intersection genes to obtain The method combines the analysis results of the isolation region with the analysis of differential gene expression to more accurately determine which genes have undergone functional changes after the isolation region is destroyed and may promote the occurrence of cancer. It can also exclude some genes without obvious differential expression and focus on genes that show significant differences in cancer tissues, thereby improving the accuracy and reliability of mining potential oncogenes. It uses gene set enrichment to screen genes related to poor prognosis and obtain potential oncogenes, which effectively solves the problem of consuming a lot of time and money costs when mining multiple gene mutations in the existing technology, and improves the accuracy and reliability of mining potential oncogenes while saving time and economic expenses. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments of the present invention or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 A flowchart of the steps of the method for mining potential proto-oncogenes based on the three-dimensional structure of chromatin provided in an embodiment of the present invention;
[0050] Figure 2A An insulating area in a normal state provided by an embodiment of the present invention;
[0051] Figure 2B The insulating region destroyed by the mutation site provided in the embodiment of the present invention;
[0052] Figure 3 A schematic diagram of a CTCF binding site predictor provided in an embodiment of the present invention;
[0053] Figure 4 Schematic diagram of a potential proto-oncogene mining system based on chromatin three-dimensional structure provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.
[0055] The embodiment of the present invention provides a method for mining potential proto-oncogenes based on the three-dimensional structure of chromatin. Figure 1 As shown, the method includes the following steps S101 to S106.
[0056] Before inputting the plurality of mutant insulating regions into the trained binding site predictor, that is, before S101, the following steps are included.
[0057] (1) Obtain the first, second, and third columns of CTCF CHIP-seq data for K562, HepG2, and GM12878 cell lines, respectively, to obtain multiple screening data, and perform an intersection operation on the multiple screening data to obtain cell line conserved CTCF CHIP-seq data. The data extracted in this step can obtain chromatin loops mediated by CTCF. The insulating region is maintained by a pair of CTCF and cohesin, and CTCF and cohesin are co-bound. Cohesin is a tetrameric protein, one of whose components is RAD21. Therefore, when two genomic regions of a CTCF-mediated interaction overlap the RAD21 binding site at the same time, we define this CTCF-mediated chromatin loop as an insulating region.
[0058] (2) Analyze the tenth column of the cell line conserved CTCF CHIP-seq data to obtain the position with the strongest interaction in each data.
[0059] (3) The data centered on the position with the strongest interaction are intercepted as positive sample data for the binding site predictor.
[0060] (4) By using the R package gkmSVM to control the sequence length, GC content, and repetitive sequence score characteristics of the positive sample data, the negative sample data corresponding to each positive sample data was obtained.
[0061] (5) Use positive sample data and negative sample data to train the binding site predictor to obtain a trained binding site predictor.
[0062] In the present invention, chromatin data includes three-dimensional chromatin interaction data and genomic data. The three-dimensional chromatin interaction data is the number of long-range interactions of CTCF in the genome; the genomic data is the distribution of CTCF protein binding sites in the genome. Cancer mutation data includes somatic mutation data for multiple cancer types; and cancer gene expression data includes gene expression profiles for each cancer type. In a specific embodiment provided by the present invention, three-dimensional chromatin interaction data for a human B-cell lymphoma cell line, specifically CTCF-mediated CHIA-PET data, was downloaded from the Encode database website. This data contains 92,808 interactions between target chromatin fragments. Subsequently, genomic data for the human chronic myeloid leukemia cell line K562, the human liver cancer cell line HepG2, and the human B-cell lymphoma cell line GM12878 were downloaded from the database. Specifically, CTCF CHIP-seq data for these three cell lines showed the number of CTCF binding sites in the genome to be 56,889, 60,229, and 43,631, respectively. Cancer mutation data, specifically somatic mutation data for liver cancer, were downloaded from the ICGC database, which includes 1706 samples. Cancer gene expression data, specifically gene expression profiles for liver cancer, were downloaded from the TCGA database, which includes 50 normal samples and 374 cancer samples.
[0063] S101, obtaining multiple mutation insulation regions based on chromatin data and cancer mutation data, inputting the multiple mutation insulation regions into a trained binding site predictor to obtain prediction results.
[0064] In step S101 , multiple mutation-insulating regions are obtained based on chromatin data and cancer mutation data, which specifically includes the following steps S1011 to S1012 .
[0065] S1011, extracting the first to fifth columns of the CTCF CHIA-PET data in the chromatin data to obtain the start and end positions of the two genomic regions of the interaction mediated by CTCF on each chromosome.
[0066] It's worth noting that the data extracted in step S1011 can reveal CTCF-mediated chromatin loops. Insulating regions are maintained by a pair of CTCF and cohesin, which co-bind. Cohesin is a tetrameric protein, one of whose components is RAD21. Therefore, when two genomic regions in a CTCF-mediated interaction overlap a RAD21 binding site, this CTCF-mediated chromatin loop is defined as an insulating region.
[0067] S1012 specifically includes three steps: (1) extracting the first, second, and third columns of the RAD21 CHIP-seq data in the chromatin data to obtain the starting and ending positions of the genome bound by RAD21 on each chromosome, and extracting the GM12878 cell line in the same way as the steps to obtain the CTCF binding sites on each chromosome; (2) intersecting the CTCF binding site data and the cancer point mutation data to obtain the CTCF binding sites containing mutations; (3) intersecting the data obtained in the step again with the data of the insulating region obtained in the previous step to finally obtain the insulating region containing mutations in the CTCF binding site.
[0068] S1013 combines the start and end positions of the two genomic regions interacting via CTCF with the start and end positions of the genomic regions bound by RAD21 to generate multiple mutation-insulating regions. Insertion and deletion mutations and point mutations can occur in humans, but this invention focuses solely on point mutations.
[0069] The binding site predictor in step S101 includes: an embedding layer, a one-dimensional convolution layer, a maximum pooling layer, a bidirectional gated recurrent neural network layer, and a fully connected layer. The layers in the binding site predictor are connected in sequence.
[0070] In step S101, multiple mutant insulating regions are input into a trained binding site predictor to obtain prediction results, which specifically includes the following steps.
[0071] (1) Obtain multiple mutant insulating regions with CTCF binding sites in the mutant insulating regions to obtain multiple mutant insulating regions with CTCF binding sites.
[0072] (2) Multiple mutant insulating regions with CTCF binding sites were input into the binding site predictor for prediction, and multiple prediction results were obtained.
[0073] In step S101, potential oncogenes are mined based on two key features: whether the insulating region containing the gene is disrupted and whether the gene is differentially expressed. Combining analysis of insulating region disruption and differential gene expression more accurately identifies genes that undergo functional changes after insulating region disruption and may contribute to cancer development. This comprehensive analysis eliminates genes with no significant differential expression and focuses on genes that exhibit significant differences in cancer tissue, improving the accuracy and reliability of potential oncogene mining.
[0074] S102, using the prediction result, determining whether the sudden insulating region is a sudden insulating region destroyed by the mutation, and obtaining the destroyed insulating region.
[0075] Step S1012 specifically includes:
[0076] (1) Obtain the prediction result and determine whether the prediction result is combined.
[0077] (2) If the prediction result is binding, the mutant insulating region with the CTCF binding site is determined to be an insulating region that will not be destroyed.
[0078] (3) If the prediction result is no binding, the mutant insulating region with the CTCF binding site is determined to be the destroyed insulating region.
[0079] S103: Obtaining multiple differentially expressed gene sets based on the cancer gene expression data, and intersecting the multiple differentially expressed gene sets to obtain a final differentially expressed gene set. This specifically includes analyzing the cancer gene expression data using the limma gene differential expression analysis package, the DESeq2 gene differential expression analysis package, or the edgeR gene differential expression analysis package to obtain the multiple differentially expressed gene sets.
[0080] S104, taking the intersection of the final differentially expressed gene set and the genes in the destroyed insulation region to obtain intersection genes.
[0081] S105: Performing a survival analysis on the intersection genes to obtain analysis results, and screening the intersection genes based on the analysis results to obtain genes associated with poor prognosis. Specifically, this includes performing a survival analysis on the intersection genes using a KM Plot database to obtain analysis results, and screening the intersection genes based on the analysis results to obtain genes with high expression levels associated with poor prognosis in cancer patients.
[0082] S106. Genes associated with poor prognosis are screened using gene set enrichment to identify potential oncogenes. Specifically, gene set enrichment analysis is performed on the genes obtained above to further identify genes whose high expression activates pathways associated with cancer development. Ultimately, these genes are identified as potential oncogenes.
[0083] It is worth noting that step S106 is a gene enrichment analysis based on single genes. Specifically, cancer samples are grouped according to the mean expression level of the target gene, with those above the mean being grouped into the high expression group and those below the mean being grouped into the low expression group. Then, the HALL MARK database is selected to screen out genes associated with high expression and activation of cancer.
[0084] In a specific embodiment provided by the present invention, Figure 2A He Ru Figure 2B A schematic diagram of the insulating area performing the regulation function provided by an embodiment of the present invention.
[0085] like Figure 2AThe image shows an insulating region in its normal state, demonstrating that CTCF binds to the DNA sequence, forming a homodimer that brings the two distal DNA ends closer together to form a chromatin loop. Cohesin is then loaded into the chromatin loop and occupies these CTCF sites, ultimately forming an insulating region. When the insulating region is in its normal state, enhancers outside the insulating region cannot act on genes within it, causing the genes to be silenced or unexpressed.
[0086] like Figure 2B The image shows the insulating region disrupted by the mutation. The mutation opens the insulating region, allowing enhancers outside the insulating region to activate genes within the original insulating region, resulting in high gene expression.
[0087] Figure 3 Schematic diagram of the CTCF binding site predictor provided by an embodiment of the present invention. Considering that CTCF can bind to either the sense or antisense strand of DNA, the sense and antisense strands of the CTCF binding site are fed into the model in parallel, sequentially undergoing an embedding layer, a one-dimensional convolution and max pooling layer, a bidirectional gated recurrent neural network layer, and a fully connected layer, ultimately outputting a predicted result of binding or non-binding. The embedding layer uses dna2vec for embedding, a one-dimensional convolution kernel to capture important features in the sequence, a recurrent neural network to capture long-term dependencies in the sequence, and a fully connected layer as the output layer.
[0088] The present invention relies on the characteristics of high-performance computer computing, and efficiently predicts the interaction between transcription factors and genes through data mining technology. While considering the regulation of linear neighbor genes of the genome by transcription factors, the influence of the three-dimensional structure of chromatin on regulation is also taken into account, and the regulatory relationship formed by spatial neighbors is integrated to obtain a more comprehensive transcription factor and gene regulatory relationship, thereby constructing a more complete cancer type-specific gene regulatory network. The present invention can realize the fusion of chromatin state and structural information with gene expression function information, follow the basic law that structure determines function, and avoid both the problem of insufficient information of single type data and the problem of information redundancy of too many types of data. Compared with traditional biological experimental methods, it greatly saves time cost and economic expenditure.
[0089] The present invention provides a system 400 for mining potential proto-oncogenes based on the three-dimensional structure of chromatin, such as Figure 4 As shown, it includes: a prediction module 401, a region acquisition module 402, a differentially expressed gene acquisition module 403, an intersection gene acquisition module 404, a gene screening module 405 and a proto-oncogene acquisition module 406.
[0090] The prediction module 401 is used to obtain multiple mutation-insulating regions based on chromatin data and cancer mutation data, and input the multiple mutation-insulating regions into a trained binding site predictor to obtain prediction results.
[0091] The region acquisition module 402 is configured to use the prediction result to determine whether the sudden insulation region is a sudden insulation region destroyed by the mutation, and obtain the destroyed insulation region.
[0092] The differentially expressed gene acquisition module 403 is used to obtain multiple differentially expressed gene sets based on the cancer gene expression data, and to obtain a final differentially expressed gene set by taking the intersection of the multiple differentially expressed gene sets.
[0093] The intersection gene acquisition module 404 is used to obtain the intersection of the final differentially expressed gene set and the genes in the destroyed isolation region to obtain the intersection genes.
[0094] The gene screening module 405 is used to perform survival analysis on the intersection genes to obtain analysis results, and screen the intersection genes according to the analysis results to obtain genes associated with poor prognosis.
[0095] The proto-oncogene acquisition module 406 is used to screen genes associated with poor prognosis using gene set enrichment to obtain potential proto-oncogenes.
[0096] Some modules in the system described herein may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0097] Through the description of the above implementation methods, it can be seen that those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, or can be embodied through the implementation process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.
[0098] The various embodiments in this specification are described in a progressive manner. References to the same or similar parts between the various embodiments are sufficient. Each embodiment focuses on the differences from other embodiments. All or part of the present invention can be used in a variety of general or specialized computer system environments or configurations. For example, personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.
[0099] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the present invention.
Claims
1. A method for mining potential proto-oncogenes based on chromatin three-dimensional structure, characterized in that: include: A plurality of mutation-insulating regions are obtained based on chromatin data and cancer mutation data, and the plurality of mutation-insulating regions are input into a trained binding site predictor to obtain prediction results; wherein the binding site predictor comprises: an embedding layer, a one-dimensional convolutional layer, a maximum pooling layer, a bidirectional gated recurrent neural network layer, and a fully connected layer, and each layer in the binding site predictor is sequentially connected; Before inputting the multiple mutation insulating regions into the trained binding site predictor, the method further includes: respectively obtaining the first column, the second column and the third column of CTCF CHIP-seq data of the K562 cell line, the HepG2 cell line and the GM12878 cell line to obtain multiple screening data, and performing an intersection operation on the multiple screening data to obtain cell line conserved CTCF CHIP-seq data; analyzing the tenth column of the cell line conserved CTCF CHIP-seq data to obtain the position with the strongest interaction in each data; intercepting part of the data centered on the position with the strongest interaction as the positive sample data of the binding site predictor; controlling the sequence length, GC content and repeat sequence score features of the positive sample data by using the R package gkmSVM to obtain negative sample data corresponding to each positive sample data; using the positive sample data and the negative sample data to train the binding site predictor to obtain the trained binding site predictor; using the prediction result to determine whether the sudden insulating region is a sudden insulating region destroyed by the mutation, thereby obtaining a destroyed insulating region; Obtaining multiple differentially expressed gene sets according to the cancer gene expression data, and taking the intersection of the multiple differentially expressed gene sets to obtain a final differentially expressed gene set; Taking the intersection of the final differentially expressed gene set and the genes in the destroyed insulation region to obtain intersection genes; Performing survival analysis on the intersection genes to obtain analysis results, and screening the intersection genes according to the analysis results to obtain genes associated with poor prognosis; Gene set enrichment is used to screen the genes associated with poor prognosis to obtain potential oncogenes.
2. The system for mining potential proto-oncogenes based on chromatin three-dimensional structure according to claim 1, characterized in that: The chromatin data includes chromatin three-dimensional interaction data and genomics data, wherein the chromatin three-dimensional interaction data is the number of long-range interactions of CTCF in the genome; The genomic data is the binding site distribution data of the CTCF protein in the genome.
3. The system for mining potential proto-oncogenes based on chromatin three-dimensional structure according to claim 1, characterized in that: The cancer mutation data includes somatic mutation data of multiple cancer types; and the cancer gene expression data includes gene expression profiles of cancer types.
4. The system for mining potential proto-oncogenes based on chromatin three-dimensional structure according to claim 1, characterized in that: The method of obtaining multiple mutation insulation regions based on chromatin data and cancer mutation data specifically includes: Extracting the first to fifth columns of the CTCF CHIA-PET data in the chromatin data to obtain the start and end positions of two genomic regions of the interaction mediated by CTCF on each chromosome; Extracting the first, second, and third columns of the RAD21 CHIP-seq data from the chromatin data to obtain the starting and ending positions of the genome bound by RAD21 on each chromosome; The start and end positions of the two genomic regions interacting via CTCF were combined with the start and end positions of the genome bound by RAD21 to generate multiple mutation-insulating regions.
5. The system for mining potential proto-oncogenes based on chromatin three-dimensional structure according to claim 1, characterized in that: Inputting the plurality of mutant insulating regions into a trained binding site predictor to obtain prediction results specifically includes: Obtaining a mutant insulating region with a CTCF binding site from the multiple mutant insulating regions, to obtain multiple mutant insulating regions with the CTCF binding site; The plurality of mutant insulating regions with CTCF binding sites are respectively input into the binding site predictor for prediction to obtain a plurality of prediction results.
6. The system for mining potential proto-oncogenes based on chromatin three-dimensional structure according to claim 5, characterized in that: The step of using the prediction result to determine whether the sudden insulating region is a sudden insulating region destroyed by a mutation, and obtaining the destroyed insulating region, specifically includes: Obtaining the prediction result, and determining whether the prediction result is a combination; If the prediction result is binding, determining the mutant insulating region with the CTCF binding site as an insulating region that will not be destroyed; If the prediction result is no binding, the mutant insulating region with the CTCF binding site is determined to be a destroyed insulating region.
7. The system for mining potential proto-oncogenes based on chromatin three-dimensional structure according to claim 1, characterized in that: The method of obtaining multiple differentially expressed gene sets based on cancer gene expression data specifically includes: The cancer gene expression data are analyzed using the limma gene differential expression analysis package, the DESeq2 gene differential expression analysis package, or the edgeR gene differential expression analysis package to obtain multiple differentially expressed gene sets.
Citation Information
Patent Citations
Cell specific genome G-quadruplex prediction method
CN113160877A
Optimizing device
JP2000029858A