A computer prediction method and system for protein-DNA-RNA triplet and its function

By acquiring and analyzing cellular DNA-RNA and protein interaction data through computer prediction methods, the difficulty of detecting endogenous protein-DNA-RNA interactions in cells was solved, the role of DRBP in diseases was revealed, and the understanding of gene regulation research was improved.

CN118116465BActive Publication Date: 2025-09-16SUN YAT SEN UNIVERSITY SHENZHEN +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410321905.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2025-09-16
Estimated Expiration
2044-03-20

AI Technical Summary

Technical Problem

Existing technologies lack effective methods to detect endogenous protein-DNA-RNA interactions in cells, especially the role of DRBP in the occurrence and development of complex diseases is unknown.

Method used

A computer prediction method for protein-DNA-RNA triplets is provided. By obtaining DNA-RNA interaction data, protein-DNA interaction data and protein-RNA interaction data of target cells, interaction scores and confidence level scores are calculated using formulas to predict protein-DNA-RNA triplets and their functions.

Benefits of technology

It has enabled the exploration of the functions and effects of the interactions between DNA, RNA and proteins in cells, improved the understanding of gene expression regulation, and provided new insights into the research of basic molecular biology and gene regulation in disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118116465B_ABST
    Figure CN118116465B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a computer prediction method and system for protein-DNA-RNA triplets and their functions, which belongs to the field of biotechnology. The solution obtains target cells from a target cell line; obtains DNA-RNA interaction data of the target cells; obtains protein-DNA interaction data of the target cells; obtains protein-RNA interaction data of the target cells; processes the DNA-RNA interaction data of the target cells to obtain the DNA-RNA interaction score, protein-DNA-RNA interaction score, and protein-DNA-RNA triplet prediction results of the target cells; and obtains functional information of downstream target genes that regulate the interaction of protein-DNA-RNA triplet based on the prediction results. The present invention can explore the functions and effects of DNA, RNA and protein interactions in cells, and provide a new type of interactome data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of biotechnology, and in particular to a computer prediction method and system for protein-DNA-RNA triplets and their functions. Background Art

[0002] Nucleic acid-binding proteins play a significant role in the development and progression of many diseases. For decades, DNA-binding proteins and RNA-binding proteins have been studied separately as two distinct proteins, but increasing evidence suggests that these proteins may include dual DNA- and RNA-binding proteins (DRBPs), which can bind to both DNA and RNA simultaneously, fulfilling complex gene regulatory functions. However, there are no effective methods to detect endogenous protein-DNA-RNA interactions at the cellular level, and the role of DRBPs in the development and progression of complex diseases remains poorly understood. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to provide a computer prediction method and system for protein-DNA-RNA triplets and their functions.

[0004] The technical solution adopted by the present invention is:

[0005] In one aspect, an embodiment of the present invention provides a computer prediction method for protein-DNA-RNA triplets and their functions, comprising the following steps:

[0006] Obtaining target cells from a target cell line; the target cell line is a group of cultured and screened cells;

[0007] Acquiring DNA-RNA interaction data of the target cell;

[0008] Obtaining protein-DNA interaction data of the target cell; the protein-DNA interaction data includes protein-DNA binding motifs and protein-DNA interaction experimental data of the target cell line;

[0009] Obtaining protein-RNA interaction data of the target cell; the protein-RNA interaction data includes protein-RNA binding motifs and protein-RNA interaction experimental data of the target cell line;

[0010] Processing the DNA-RNA interaction data of the target cell to obtain a DNA-RNA interaction score of the target cell;

[0011] Obtaining a protein-DNA-RNA interaction score according to the DNA-RNA interaction score, the protein-DNA interaction data, and the protein-RNA interaction data;

[0012] According to the protein-DNA-RNA interaction score, a prediction result of the protein-DNA-RNA triplet is obtained;

[0013] According to the prediction results of the protein-DNA-RNA triplet, functional information of the downstream target gene regulated by the interaction of the protein-DNA-RNA triplet is obtained.

[0014] Furthermore, obtaining the DNA-RNA interaction data of the target cell includes:

[0015] The DNA-RNA interaction data is filtered to obtain filtered data.

[0016] Furthermore, the protein-DNA-RNA interaction score is obtained according to the DNA-RNA interaction score, the protein-DNA interaction data and the protein-RNA interaction data, comprising:

[0017] Scanning the protein binding sites in the DNA-RNA hybrid fragment according to the protein-DNA binding motif and the protein-RNA binding motif;

[0018] Calculating based on the protein binding sites in the DNA-RNA hybrid fragment to obtain a conservation score;

[0019] The conservation scores are screened to obtain confidence level scores for protein-DNA-RNA interactions.

[0020] Furthermore, the calculation based on the protein binding sites in the DNA-RNA hybrid fragment to obtain a conservation score includes:

[0021] The conservation score of the genomic coordinate sequence of the protein binding site was calculated using phyloP.

[0022] Furthermore, the conservation score is screened to obtain a confidence level score for the protein-DNA-RNA interaction, and the formula used includes:

[0023] T=S D ×S R ×S P.D ×S P.R (1)

[0024] Wherein, T represents the confidence level score of the protein-DNA-RNA interaction, S D and S R represent DNA binding score and RNA binding score, respectively, S P.D and S P.R represent the conservation scores of DNA binding sites and RNA binding sites, respectively.

[0025] Furthermore, the protein-DNA-RNA interaction score is obtained according to the DNA-RNA interaction score, the protein-DNA interaction data and the protein-RNA interaction data, comprising:

[0026] extracting the protein-DNA interaction experimental data and the protein-RNA interaction experimental data of the target cell line to obtain protein-DNA-RNA interaction regions;

[0027] The confidence level score of the protein-DNA-RNA interaction is calculated based on the protein-DNA-RNA interaction region.

[0028] Furthermore, the calculation based on the protein-DNA-RNA interaction region to obtain a confidence level score of the protein-DNA-RNA interaction includes:

[0029] T=S D.R ×S P.NA (2)

[0030] Wherein, T represents the confidence level score of the protein-DNA-RNA interaction, S D.R The peak score of DNA-RNA interaction representing the protein-DNA-RNA interaction region, S P.NA The peak score of protein-DNA interaction or the peak score of protein-RNA interaction represents the protein-DNA-RNA interaction region.

[0031] Furthermore, the functional information of downstream target genes regulated by the protein-DNA-RNA triplet interaction is obtained based on the prediction results of the protein-DNA-RNA triplet, including:

[0032] Obtaining relevant RNA data in the prediction results of the protein-DNA-RNA triplet;

[0033] Obtain lncRNA data and lncRNA regulatory information data;

[0034] Comparing the lncRNA data with the related RNA data to obtain a comparison result;

[0035] According to the comparison results and the lncRNA regulatory information data, functional information of the downstream target genes regulated by the protein-DNA-RNA triplet interaction is obtained.

[0036] Furthermore, the method of obtaining functional information of downstream target genes regulated by the protein-DNA-RNA triplet interaction based on the prediction results of the protein-DNA-RNA triplet also includes:

[0037] Obtaining the interaction binding site of the protein-DNA-RNA triplet of the target cell;

[0038] Obtaining genes whose interaction binding sites are no more than 1000 bp apart;

[0039] Obtaining genes whose interaction binding sites are no more than 5000 bp apart;

[0040] Based on the obtained genes, functional information of downstream target genes regulated by the protein-DNA-RNA triplet interaction is obtained.

[0041] On the other hand, an embodiment of the present invention further provides a computer prediction system for protein-DNA-RNA triplets and their functions, the system comprising:

[0042] The first module is used to obtain target cells from a target cell line; the target cell line is a group of cultured and screened cells;

[0043] The second module is used to obtain DNA-RNA interaction data of the target cell;

[0044] The third module is used to obtain protein-DNA interaction data of the target cell; the protein-DNA interaction data includes protein-DNA binding motifs and protein-DNA interaction experimental data of the target cell line;

[0045] The fourth module is used to obtain protein-RNA interaction data of the target cell; the protein-RNA interaction data includes protein-RNA binding motifs and protein-RNA interaction experimental data of the target cell line;

[0046] A fifth module is used to process the DNA-RNA interaction data of the target cell to obtain the DNA-RNA interaction score of the target cell;

[0047] a sixth module, configured to obtain a protein-DNA-RNA interaction score based on the DNA-RNA interaction score, the protein-DNA interaction data, and the protein-RNA interaction data;

[0048] The seventh module is used to obtain the prediction result of protein-DNA-RNA triplet according to the protein-DNA-RNA interaction score;

[0049] The eighth module is used to obtain the functional information of the downstream target genes regulated by the protein-DNA-RNA triplet interaction based on the prediction results of the protein-DNA-RNA triplet.

[0050] On the other hand, an embodiment of the present invention further provides an electronic device, including a processor and a memory;

[0051] The memory is used to store programs;

[0052] The processor executes the program to implement the aforementioned computer prediction method for protein-DNA-RNA triplet and its regulatory function.

[0053] On the other hand, an embodiment of the present invention further provides a computer storage medium storing a program executable by a processor, which, when executed by the processor, is used to implement the aforementioned computer prediction method for protein-DNA-RNA triplets and their functions.

[0054] The embodiments of the present application include at least the following beneficial effects: The present application provides a computer prediction method and system for protein-DNA-RNA triplets and their functions, which comprises obtaining target cells from a target cell line; the target cell line is a group of cultured and screened cells; obtaining DNA-RNA interaction data of the target cells; obtaining protein-DNA interaction data of the target cells; the protein-DNA interaction data includes protein-DNA binding motifs and protein-DNA interaction experimental data of the target cell line; obtaining protein-RNA interaction data of the target cells; the protein-RNA interaction data includes protein-RNA binding motifs and protein-RNA interaction experimental data of the target cell line; processing the DNA-RNA interaction data of the target cells to obtain a DNA-RNA interaction score of the target cells; obtaining a protein-DNA-RNA interaction score based on the DNA-RNA interaction score, the protein-DNA interaction data and the protein-RNA interaction data; obtaining a prediction result of the protein-DNA-RNA triplet based on the protein-DNA-RNA interaction score; and obtaining functional information of downstream target genes regulated by the protein-DNA-RNA triplet interaction based on the prediction result of the protein-DNA-RNA triplet. This invention can explore the functions and effects of the interactions between DNA, RNA and proteins in cells, improve our understanding of gene expression regulation, provide a new type of interactome data, and provide new insights into the role of basic molecular biology and gene regulation in disease. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 This is a flow chart of a method for computer prediction of protein-DNA-RNA triplets and their functions provided by an embodiment of the present invention;

[0056] Figure 2 This is a module connection diagram provided by an embodiment of the present invention;

[0057] Figure 3 This is a device connection diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0059] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0060] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0062] Before explaining the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0063] 1) The protein-DNA-RNA triplet is a complex molecular structure in biology, formed by the interaction of three biological molecules: protein, DNA, and RNA. This triplet structure plays an important biological role in cells, regulating gene transcription, translation, and post-transcription, thereby affecting protein synthesis and cell function. Specifically, the protein-DNA-RNA triplet usually involves the following components:

[0064] DNA (deoxyribonucleic acid): A molecule with a double helix structure that stores an organism's genetic information. DNA, through its sequence of bases, encodes the proteins and RNA needed in cells.

[0065] Proteins are among the most versatile molecules in living organisms, performing diverse functions within cells, such as catalyzing chemical reactions, providing structural support, and transmitting signals. Within the protein-DNA-RNA triplet, proteins are often transcription factors or other proteins involved in regulating gene expression.

[0066] RNA (ribonucleic acid): A type of nucleic acid molecule similar to DNA that performs multiple functions in cells, including gene regulation and protein synthesis. In the protein-DNA-RNA triplet, RNA is typically transcribed from DNA and may be mRNA (messenger RNA), tRNA (transfer RNA), or rRNA (ribosomal RNA).

[0067] 2) A target cell line is a cell line used in scientific research. These cell lines are selected or engineered to investigate specific biological processes, disease mechanisms, or drug effects. These cell lines typically possess properties that make them suitable for specific types of studies under experimental conditions.

[0068] 3) PhyloP, a statistical method for calculating the conservation of genomic variation in DNA sequences. This method compares DNA sequence variation between different species to identify regions that are conserved or highly conserved during evolution.

[0069] Specifically, the PhyloP method compares DNA sequences from multiple species to identify regions that have been conserved over evolution relative to random variation. Conservation is assessed by examining whether the base at each position remains unchanged or maintains consistent variation throughout evolution. If the base at a position remains unchanged or maintains the same variation across multiple species, then that position likely plays an important functional role during evolution and is therefore considered conserved.

[0070] 4) lncRNA data refers to data related to long non-coding RNA (lncRNA). This data typically includes information such as lncRNA sequence, expression level, structure, and function. lncRNAs are a class of non-coding RNA molecules longer than 200 nucleotides. They are not translated into proteins during transcription but play an important role in regulating gene expression, cellular processes, and development.

[0071] The embodiments of the present invention are further described below with reference to the accompanying drawings.

[0072] On the one hand, embodiments of the present invention provide a method for computer prediction of protein-DNA-RNA triplets and their functions. This method can explore the functions and roles of DNA, RNA, and protein interactions in cells, improve our understanding of gene expression regulation, and provide a new type of interactome data, providing new insights into basic molecular biology and the role of gene regulation in disease.

[0073] The embodiment of the present invention provides a computer prediction method for protein-DNA-RNA triplet and its function, specifically, referring to Figure 1 , the method comprises the following steps:

[0074] S100, obtaining target cells from a target cell line; the target cell line is a group of cultured and screened cells;

[0075] S200, obtaining DNA-RNA interaction data of target cells;

[0076] S300, obtaining protein-DNA interaction data of target cells; the protein-DNA interaction data includes protein-DNA binding motifs and protein-DNA interaction experimental data of target cell lines;

[0077] S400, obtaining protein-RNA interaction data of target cells; the protein-RNA interaction data includes protein-RNA binding motifs and protein-RNA interaction experimental data of target cell lines;

[0078] S500, processing the DNA-RNA interaction data of the target cell to obtain a DNA-RNA interaction score of the target cell;

[0079] S600, obtaining a protein-DNA-RNA interaction score based on the DNA-RNA interaction score, protein-DNA interaction data, and protein-RNA interaction data;

[0080] S700, obtaining the prediction results of protein-DNA-RNA triplet according to the protein-DNA-RNA interaction score;

[0081] S800. Based on the prediction results of the protein-DNA-RNA triplet, obtain the functional information of the downstream target genes regulated by the protein-DNA-RNA triplet interaction.

[0082] The embodiment of the present invention discloses step S200 of obtaining DNA-RNA interaction data of target cells, including:

[0083] S210 , filtering the DNA-RNA interaction data to obtain filtered data.

[0084] As an optional implementation, the embodiment of the present invention obtains DNA-RNA interaction data and protein-DNA / RNA interaction data of target cells, such as GRID-seq data; the DNA-RNA interaction data of target cells needs to be filtered to select high-quality data with high credibility; protein-DNA / RNA interaction data, including protein DNA / RNA binding motifs, and protein-DNA / RNA interaction experimental data of target cell lines, such as ChIP-seq data and protein-centric data.

[0085] The embodiment of the present invention discloses that step S600 obtains a protein-DNA-RNA interaction score based on the DNA-RNA interaction score, the protein-DNA interaction data, and the protein-RNA interaction data, including:

[0086] S610, scanning the protein binding sites in the DNA-RNA hybrid fragment according to the protein-DNA binding motif and the protein-RNA binding motif;

[0087] S620, calculating based on the protein binding sites in the DNA-RNA hybrid fragment to obtain a conservation score;

[0088] S630. Screen the conservation scores to obtain confidence level scores for protein-DNA-RNA interactions.

[0089] The embodiment of the present invention discloses that step S620 calculates the protein binding sites in the DNA-RNA hybrid fragment to obtain a conservation score, including:

[0090] S621. Use phyloP to calculate the conservation score of the genomic coordinate sequence of the protein binding site.

[0091] The embodiment of the present invention discloses that step S630 screens the conservation score to obtain a confidence level score of the protein-DNA-RNA interaction, and the formula used includes:

[0092] T=S D ×S R ×S P.D ×S P.R (1)

[0093] Among them, S represents the confidence level score of protein-DNA-RNA interaction, S D and S Rrepresent DNA binding score and RNA binding score, respectively, S P.D and S P.R represent the conservation scores of DNA binding sites and RNA binding sites, respectively.

[0094] As an optional embodiment, in step S600 of the embodiment of the present invention, when the DNA / RNA binding motif of the protein is used, the step also includes: scanning the protein binding sites in the DNA-RNA hybrid fragment; calculating the conservation of the screened binding sites, and extracting the DNA binding score, RNA binding score, DNA binding site conservation score and RNA binding site conservation score as the confidence score of the protein-DNA-RNA interaction pairing.

[0095] When the target cell line's protein-DNA / RNA interaction experimental data are collected, the following steps also include extracting each DNA-RNA interaction: The intersection of the protein and DNA / RNA interaction peaks is the protein-DNA-RNA interaction region. The peak matrices of the paired regions are multiplied together to provide a confidence score for the protein-DNA-RNA interaction.

[0096] The confidence score for protein-DNA-RNA interactions is calculated as follows:

[0097] When the DNA / RNA binding motif of a protein is used, phyloP is used to calculate the average conservation of the genomic coordinate sequence of the nucleic acid binding site, and the following calculation formula is constructed, namely formula (1):

[0098] T=S D ×S R ×S P.D ×S P.R (1)

[0099] Where T represents the confidence level score of protein-DNA-RNA interaction, S D and S R represent DNA binding score and RNA binding score, respectively, S P.D and S P.R represent the conservation scores of DNA binding sites and RNA binding sites, respectively.

[0100] The embodiment of the present invention discloses that step S600 obtains a protein-DNA-RNA interaction score based on the DNA-RNA interaction score, the protein-DNA interaction data, and the protein-RNA interaction data, including:

[0101] S640, extracting the protein-DNA interaction experimental data and the protein-RNA interaction experimental data of the target cell line to obtain protein-DNA-RNA interaction regions;

[0102] S650. Calculate the confidence level score of the protein-DNA-RNA interaction based on the protein-DNA-RNA interaction region.

[0103] The embodiment of the present invention discloses that step S650 calculates the confidence level score of the protein-DNA-RNA interaction based on the protein-DNA-RNA interaction region, including:

[0104] T=S D.R ×S P.NA (2)

[0105] Where, T represents the confidence level score of protein-DNA-RNA interaction, S D.R The peak fraction of DNA-RNA interactions representing the protein-DNA-RNA interaction region, S P.NA The peak fraction of protein-DNA interaction or the peak fraction of protein-RNA interaction represents the protein-DNA-RNA interaction region.

[0106] As an optional embodiment, in step S600 of the present embodiment, when using the experimental data of protein and DNA / RNA interaction of the target cell line, each DNA-RNA interaction is extracted separately: the intersection of the protein and DNA / RNA interaction peak is the protein-DNA-RNA interaction region, and the following calculation formula is constructed, namely formula (2):

[0107] T=S D.R ×S P.NA (2)

[0108] Where T represents the confidence level score of protein-DNA-RNA interaction, S D.R represents the peak fraction of DNA-RNA interaction at the interaction peak intersection, S P.NA The peak fraction representing the protein-DNA interaction or protein-RNA interaction at the interaction peak intersection.

[0109] The embodiment of the present invention discloses that step S800 obtains functional information of downstream target genes regulated by the protein-DNA-RNA triplet interaction based on the prediction results of the protein-DNA-RNA triplet, including:

[0110] S810, obtaining relevant RNA data in the prediction results of protein-DNA-RNA triplet;

[0111] S820, acquiring lncRNA data and lncRNA regulatory information data;

[0112] S830, comparing the lncRNA data and the related RNA data to obtain a comparison result;

[0113] S840. Based on the comparison results and lncRNA regulatory information data, obtain the functional information of the downstream target genes regulated by the protein-DNA-RNA triplet interaction.

[0114] As an optional implementation, an embodiment of the present invention obtains lncRNA data and lncRNA regulatory information data to determine whether the predicted protein-DNA-RNA triplet of the predicted target cell and the related RNA in its function is lncRNA, and the regulatory function of the RNA in the triplet interaction is considered to be the function of regulating the downstream target gene of the protein-DNA-RNA triplet interaction.

[0115] The embodiment of the present invention discloses that step S800 obtains functional information of downstream target genes regulated by the protein-DNA-RNA triplet interaction based on the prediction results of the protein-DNA-RNA triplet, including:

[0116] S850, obtaining the interaction binding site of the protein-DNA-RNA triplet of the target cell;

[0117] S860, obtain genes whose interaction binding sites are no more than 1000 bp apart;

[0118] S870, obtain genes whose interaction binding sites are no more than 5000 bp apart;

[0119] S880. Based on the acquired genes, obtain functional information of downstream target genes regulated by protein-DNA-RNA triplet interactions.

[0120] As an optional embodiment, in the embodiment of the present invention, based on the binding site of the protein-DNA-RNA triplet interaction of the target cell, genes within 1000bp or 5000bp downstream are considered to be potentially regulated by the interaction, and functional information of the downstream target genes regulated by the protein-DNA-RNA triplet interaction is obtained.

[0121] Table 1 shows a list of protein-DNA-RNA triplet interactions filtered by GRID-seq (i.e., DNA-RNA interaction data) and protein-RNA / DNA binding motifs (i.e., protein-RNA / DNA interaction data) for the breast cancer cell line MDA-MB-231 and the peripheral blood myeloma cell line MM.1S (i.e., target cell lines). The table shows the top 10 protein-DNA-RNA interactions.

[0122] Table 1 List of protein-DNA-RNA interactions screened by GRID-seq and NBPs binding motifs

[0123]

[0124]

[0125] * Column names from left to right are: protein name, DNA binding site conservation score, genes within 1k to 5k bp downstream of the interaction site, RNA name and ID, lncRNA identity category, protein-RNA binding ability score, RNA binding site conservation score, and protein-DNA-RNA interaction score.

[0126] It should be noted that DNA-RNA interaction sites were scanned using the protein's DNA- and RNA-binding motifs. The same protein and its DNA- and RNA-binding sites that matched the DNA-RNA interaction sequence pairs were referred to as protein-DNA-RNA interaction triplets. In this example, 79 such triplets were screened.

[0127] The GRID-seq MACS2 peak score, protein-DNA interaction score, protein-RNA interaction score, DNA binding site conservation score, and RNA binding site conservation score were multiplied together to form the triple interaction score (Table 1). Of the screened results, 54 were positive, indicating that the corresponding motifs had slower evolution and more conserved properties than expected under neutral drift; 25 were negative, with a high probability of being false positives.

[0128] The highest-scoring group is TIA1 (protein)-SNHG1 (RNA)-CBX3 (gene / DNA)-HNRNPA2B1 (gene / DNA). All members of this interaction group are highly expressed in the cerebral cortex. TIA1 is a protein associated with apoptosis, and mutations in this gene lead to various neurological diseases, such as amyotrophic lateral sclerosis and dementia. SNHG1 is a lncRNA that promotes the transcription of nearby genes through cis-regulation. CBX3 is a key factor in neural differentiation. HNRNPA2B1 is associated with neurodegeneration, and mutations in this gene can lead to various neurological diseases such as multisystem proteinopathy and amyotrophic lateral sclerosis, which can also be caused by mutations in TIA1. These results suggest that TIA1 and SNHG1 may interact, and this complex has a cis-regulatory effect on CBX3 and HNRNPA2B1, which in turn contributes to neuronal development and differentiation.

[0129] Table 2 shows a list of protein-DNA-RNA interactions screened from GRID-seq (i.e., DNA-RNA interaction data) and ChIP-seq data (i.e., protein-DNA interaction data) for the peripheral blood myeloma cell line MM.1S (i.e., target cell line). This table shows the top 15 protein-DNA-RNA interactions.

[0130] Table 2 List of protein-DNA-RNA interactions cross-screened by GRID-seq and ChIP-seq peaks in MM.1S

[0131]

[0132]

[0133] It should be noted that the examples cross-compare GRID-seq peak locations with ChIP-seq peak locations to extract cross-peaks. These genomic locations are considered to be protein-DNA-RNA interaction regions, and the involved proteins and RNAs, along with the genomic regions, are defined as protein-DNA-RNA interaction triplets. The confidence level of the triplets is obtained by multiplying the GRID-seq and ChIP-seq MACS2 peak scores.

[0134] The highest-scoring interaction in the MM.1S cell line was CTCF (protein)-DUSP22 (RNA)-IGHV3OR16-12 (gene / DNA)-IGHV3OR16-13 (gene / DNA)-AC136428.3 (gene / DNA) [Table 2]. CTCF (protein) and DUSP22 (RNA) play a role in regulating chromosomal rearrangements during the development and progression of hematological tumors and cancers. Chromosomal rearrangements are considered a key component of antibody production and the development of hematological malignancies. Regulated IGHV3OR16-12 and IGHV3OR16-13 are important components of immunoglobulin heavy chains. This series of protein-DNA-RNA interactions plays a crucial role in regulating the proliferation, differentiation, and antibody production of blood tissue cells. In fact, the majority of genes listed in Table 2 (11 / 15) are associated with immunoglobulin heavy chains, indicating that the protein-DNA-RNA interactions identified in this study are important in regulating antibody production.

[0135] Combined with the above description of Tables 1 and 2, after applying the present invention, proteins that bind to both DNA and RNA were computationally screened in the breast cancer cell line MDA-MB-231 and the peripheral blood myeloma cell line MM.1S, revealing the molecular mechanism of protein-DNA-RNA interactions in cancer cells.

[0136] This paper screens intracellular protein-DNA-RNA interactions through comprehensive analysis of multi-omics data, combined with GRID-seq, ChIP-seq and calculation of binding motif information of nucleic acid-binding proteins, to improve our understanding of the functions and regulatory mechanisms of protein-DNA-RNA interactions in cells, and explore the role of protein-DNA-RNA interactions in diseases.

[0137] Reference Figure 2 , shown is a method and system for computer screening and predicting protein-DNA-RNA triplets and their functions provided according to an embodiment of the present invention, including the following modules:

[0138] DNA-RNA interaction module 201, used to obtain high-quality DNA-RNA interaction data;

[0139] A protein-RNA interaction module 202 is used to obtain protein-RNA interaction data from the RNA binding motif of the protein or the protein-RNA interaction experimental data of the target cell line;

[0140] A protein-DNA interaction module 203 is used to obtain protein-DNA interaction data from the DNA binding motif of the protein or the protein-DNA interaction experimental data of the target cell line;

[0141] The protein-DNA-RNA interaction confidence level calculation module 204 is connected to the DNA-RNA interaction module 201, the protein-RNA interaction module 202, and the protein-DNA interaction module 203 to realize interaction, and is used to calculate the confidence score of the protein-DNA-RNA interaction based on the DNA-RNA interaction data, the protein-RNA interaction data, and the protein-DNA interaction data, and obtain a prediction result of the protein-DNA-RNA triplet based on the confidence score of the protein-DNA-RNA interaction;

[0142] The protein-DNA-RNA interaction function prediction module 205 is connected to the protein-DNA-RNA interaction confidence level calculation module 204 to realize interaction, and is used to predict the function of the protein-DNA-RNA in regulating downstream target genes based on whether the RNA in the protein-DNA-RNA interaction is lncRNA and the genes that may be regulated by the interaction.

[0143] On the other hand, an embodiment of the present invention further provides a computer prediction system for protein-DNA-RNA triplets and their functions, comprising:

[0144] The first module is used to obtain target cells from a target cell line; the target cell line is a group of cultured and screened cells;

[0145] The second module is used to obtain DNA-RNA interaction data of target cells;

[0146] The third module is used to obtain protein-DNA interaction data of target cells; protein-DNA interaction data includes protein-DNA binding motifs and protein-DNA interaction experimental data of target cell lines;

[0147] The fourth module is used to obtain protein-RNA interaction data of target cells; protein-RNA interaction data includes protein-RNA binding motifs and protein-RNA interaction experimental data of target cell lines;

[0148] The fifth module is used to process the DNA-RNA interaction data of the target cells to obtain the DNA-RNA interaction score of the target cells;

[0149] The sixth module is used to obtain the protein-DNA-RNA interaction score based on the DNA-RNA interaction score, protein-DNA interaction data and protein-RNA interaction data;

[0150] The seventh module is used to obtain the prediction results of protein-DNA-RNA triplets based on the protein-DNA-RNA interaction score;

[0151] The eighth module is used to obtain the functional information of the downstream target genes regulated by the protein-DNA-RNA triplet interaction based on the prediction results of the protein-DNA-RNA triplet.

[0152] It can be understood that the contents of the above method embodiments are applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0153] On the other hand, reference Figure 3 , an embodiment of the present invention further provides an electronic device, including a processor 301 and a memory 302;

[0154] The memory 302 is used to store programs;

[0155] The processor 301 executes the program to implement the aforementioned computer prediction method for protein-DNA-RNA triplet and its function.

[0156] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0157] Memory 302 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by the processor or other modules of the computer. Permanent storage devices may be readable and writable storage devices. Permanent storage devices may be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (e.g., a magnetic or optical disk, flash memory) as the permanent storage device. In other embodiments, the permanent storage device may be a removable storage device (e.g., a floppy disk, optical drive). System memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. System memory may store some or all instructions and data required by the processor during operation. In addition, memory 302 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 302 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.

[0158] The memory 302 stores executable codes. When the executable codes are processed by the processor 301 , the processor 301 can execute part or all of the above-mentioned methods.

[0159] On the other hand, an embodiment of the present invention further provides a computer storage medium storing a processor-executable program, which, when executed by the processor, is used to implement the aforementioned computer prediction method for protein-DNA-RNA triplets and their functions.

[0160] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0161] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A computer prediction method for protein-DNA-RNA triplets and their functions, characterized in that: The method comprises: Obtaining target cells from a target cell line; the target cell line is a group of cultured and screened cells; Acquiring DNA-RNA interaction data of the target cell; Obtaining protein-DNA interaction data of the target cell; the protein-DNA interaction data includes protein-DNA binding motifs and protein-DNA interaction experimental data of the target cell line; Obtaining protein-RNA interaction data of the target cell; the protein-RNA interaction data includes protein-RNA binding motifs and protein-RNA interaction experimental data of the target cell line; Processing the DNA-RNA interaction data of the target cell to obtain a DNA-RNA interaction score of the target cell; Obtaining a protein-DNA-RNA interaction score according to the DNA-RNA interaction score, the protein-DNA interaction data, and the protein-RNA interaction data; According to the protein-DNA-RNA interaction score, a prediction result of the protein-DNA-RNA triplet is obtained; According to the prediction results of the protein-DNA-RNA triplet, functional information of the downstream target gene regulated by the interaction of the protein-DNA-RNA triplet is obtained.

2. The method according to claim 1, characterized in that The obtaining of the DNA-RNA interaction data of the target cell comprises: The DNA-RNA interaction data is filtered to obtain filtered data.

3. The method according to claim 1, characterized in that Obtaining a protein-DNA-RNA interaction score according to the DNA-RNA interaction score, the protein-DNA interaction data, and the protein-RNA interaction data comprises: Scanning the protein binding sites in the DNA-RNA hybrid fragment according to the protein-DNA binding motif and the protein-RNA binding motif; Calculating based on the protein binding sites in the DNA-RNA hybrid fragment to obtain a conservation score; The conservation scores are screened to obtain confidence level scores for protein-DNA-RNA interactions.

4. The method according to claim 3, characterized in that The calculation based on the protein binding sites in the DNA-RNA hybrid fragment to obtain a conservation score includes: The conservation score of the genomic coordinate sequence of the protein binding site was calculated using phyloP.

5. The method according to claim 3, characterized in that The conservation score is screened to obtain a confidence level score for the protein-DNA-RNA interaction, and the formula used includes: T= S D ×S R ×S P.D ×S P.R (1) Wherein, T represents the confidence level score of the protein-DNA-RNA interaction, S D and S R represent DNA binding score and RNA binding score, respectively, S P.D and S P.R represent the conservation scores of DNA binding sites and RNA binding sites, respectively.

6. The method according to claim 1, characterized in that Obtaining a protein-DNA-RNA interaction score according to the DNA-RNA interaction score, the protein-DNA interaction data, and the protein-RNA interaction data comprises: extracting the protein-DNA interaction experimental data and the protein-RNA interaction experimental data of the target cell line to obtain protein-DNA-RNA interaction regions; The confidence level score of the protein-DNA-RNA interaction is calculated based on the protein-DNA-RNA interaction region.

7. The method according to claim 6, characterized in that The calculation based on the protein-DNA-RNA interaction region to obtain a confidence level score of the protein-DNA-RNA interaction includes: T=S D.R ×S P.NA (2) Wherein, T represents the confidence level score of the protein-DNA-RNA interaction, S D.R The peak score of DNA-RNA interaction representing the protein-DNA-RNA interaction region, S P.NA The peak score of protein-DNA interaction or the peak score of protein-RNA interaction represents the protein-DNA-RNA interaction region.

8. The method according to claim 1, characterized in that The function information of the downstream target gene regulated by the interaction of the protein-DNA-RNA triplet is obtained based on the prediction result of the protein-DNA-RNA triplet, including: Obtaining relevant RNA data in the prediction results of the protein-DNA-RNA triplet; Obtain lncRNA data and lncRNA regulatory information data; Comparing the lncRNA data with the related RNA data to obtain a comparison result; According to the comparison results and the lncRNA regulatory information data, functional information of the downstream target genes regulated by the protein-DNA-RNA triplet interaction is obtained.

9. The method according to claim 1, wherein obtaining functional information of downstream target genes regulated by the protein-DNA-RNA triplet interaction based on the prediction results of the protein-DNA-RNA triplet comprises: Obtaining the interaction binding site of the protein-DNA-RNA triplet of the target cell; Obtaining genes whose interaction binding sites are no more than 1000 bp apart; Obtaining genes whose interaction binding sites are no more than 5000 bp apart; Based on the obtained genes, functional information of downstream target genes regulated by the protein-DNA-RNA triplet interaction is obtained.

10. A computer prediction system for protein-DNA-RNA triplets and their functions, characterized in that: The system comprises: The first module is used to obtain target cells from a target cell line; the target cell line is a group of cultured and screened cells; The second module is used to obtain DNA-RNA interaction data of the target cell; The third module is used to obtain protein-DNA interaction data of the target cell; the protein-DNA interaction data includes protein-DNA binding motifs and protein-DNA interaction experimental data of the target cell line; The fourth module is used to obtain protein-RNA interaction data of the target cell; the protein-RNA interaction data includes protein-RNA binding motifs and protein-RNA interaction experimental data of the target cell line; A fifth module is used to process the DNA-RNA interaction data of the target cell to obtain the DNA-RNA interaction score of the target cell; a sixth module, configured to obtain a protein-DNA-RNA interaction score based on the DNA-RNA interaction score, the protein-DNA interaction data, and the protein-RNA interaction data; The seventh module is used to obtain the prediction result of protein-DNA-RNA triplet according to the protein-DNA-RNA interaction score; The eighth module is used to obtain the functional information of the downstream target genes regulated by the protein-DNA-RNA triplet interaction based on the prediction results of the protein-DNA-RNA triplet.

Citation Information

Patent Citations

  • Method for predicting protein-RNA interaction sites

    CN109949859A

  • RNA-protein interaction prediction method and device, medium and electronic equipment

    CN116529828A