Chromatin interaction information enhancement method based on RNA-DNA and DNA-DNA interaction information
The CNN network model integrates RNA-DNA interaction data and Hi-C data to generate an enhanced Hi-C matrix, which solves the problem of unknown RNA in the existing technology in chromatin regulation, and achieves efficient RNA information integration and Hi-C map enhancement.
Patent Information
- Application Number
- CN202510656424.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The prior art is difficult to effectively integrate RNA interaction data and Hi-C data, and cannot fully reveal the role of RNA in chromatin regulation.
Using a CNN network model method, RNA-DNA interaction data is fused with Hi-C data, and the enhanced Hi-C matrix is generated through Pearson correlation coefficient calculation and FAN integrity processing.
It realizes the effective integration of RNA information into chromatin interaction modeling, improves the spatial resolution of the Hi-C map, has good generalization ability across RNA types, and supports quantitative functional analysis and biological interpretation.
Smart Images

Figure CN120183510A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and particularly to a method for enhancing chromatin interaction information based on RNA-DNA and DNA-DNA interaction information. Background Art
[0002] In recent years, the research on the functions of nuclear RNAs in eukaryotic cells and their interactions with chromatin has been gradually deepened. RNA is no longer regarded as merely an intermediary in gene expression. It plays diverse and important roles in the dynamic regulation of chromatin conformation. There are various RNAs in cells, including coding RNAs and non-coding RNAs (such as lncRNAs, snoRNAs, snRNAs, etc.). These RNAs participate in gene transcription, epigenetic regulation, and cell fate determination by regulating the structure and function of chromatin. The study of the interaction between RNA and chromatin provides a new perspective for understanding gene expression regulation.
[0003] Currently, the interaction between RNA and chromatin is considered to play an important role in the organization within the cell nucleus. By binding to chromatin, RNA can regulate the open state of chromatin, affect the gene transcription program, and ensure that cells maintain a normal gene expression pattern during physiological processes. Existing genomic research techniques, such as Hi-C, eCLIP-seq, etc., have been able to provide preliminary data on the interaction between chromatin and RNA. However, most of the existing techniques focus on analyzing the three-dimensional structure of DNA and lack tools for integrating RNA interaction information, making it impossible to comprehensively reveal the role of RNA in chromatin regulation.
[0004] In the prior art, the MARGI (Mapping of RNA-Genome Interaction) technique, as an early method for studying RNA-DNA interaction, combines cross-linking immunoprecipitation and high-throughput sequencing and can preliminarily analyze the binding sites of RNA on the genome. However, this method still has certain limitations in terms of spatial accuracy and resolution. Subsequently, the developed iMARGI (in situ MARGI) technique has achieved significant improvements in in situ capture and spatial resolution, can more accurately locate RNA-DNA interaction sites, and conduct experiments under the condition of maintaining chromatin structure, enhancing its biological significance.
[0005] GRID-seq (Global RNA Interaction with DNA), as a deep sequencing technology, can capture the global interaction between RNA and DNA, revealing chromatin-related RNA sites and their specific binding positions to chromatin. This technology provides an important tool for studying the role of RNA in chromatin regulation. However, GRID-seq has not been able to comprehensively process multi-modal interaction data between different RNA types and chromatin, and lacks the ability to integrate with other genomic data (such as Hi-C, epigenetic data).
[0006] RIC-seq (RNA in situ conformation sequencing) provides a deeper understanding of the spatial organization of RNA and chromatin by measuring the secondary and tertiary structures of RNA, as well as RNA-RNA interactions, with high resolution. Although RIC-seq can provide rich information on RNA structure, the complexity of data interpretation and how to integrate these data with other data on chromatin interaction remain a technical challenge.
[0007] To fill these gaps, the RD-SPRITE technology can help reveal the role of RNA in chromatin spatial regulation by mapping the spatial interaction between RNA and DNA with high resolution. However, despite the fact that the above technologies provide rich tools and data sources for the study of RNA-chromatin interaction, due to the diversity of technical processes, there is still a lack of a systematic computational analysis framework that can effectively integrate RNA interaction data generated by different technologies with chromatin conformation data such as Hi-C. Therefore, developing an efficient analysis tool that can integrate RNA interaction and Hi-C data has become an urgent problem to be solved in the current field of bioinformatics. Summary of the Invention
[0008] In view of the above problems, the object of the present invention is to provide a method for enhancing chromatin interaction information based on RNA-DNA and DNA-DNA interaction information. This method integrates RNA interaction data and Hi-C data to construct a multi-modal data integration and analysis platform, which can make up for the deficiencies of the existing technology in multi-modal data integration and promote the further development of the research on RNA-chromatin interaction.
[0009] To solve the above technical problems, the present invention provides the following technical solutions:
[0010] A method for enhancing chromatin interaction information based on RNA-DNA and DNA-DNA interaction information, the method comprising the following steps:
[0011] S1. Collect the iMARGI data of RNA-DNA interactions and the original Hi-C data of DNA-DNA interactions for the same cell line;
[0012] S2. After feature extraction of the iMARGI data, calculate the Pearson correlation coefficient to obtain the DNA-DNA feature correlation coefficient matrix of the genome-wide fusion RNA information; After downsampling the original Hi-C data, perform FAN completion processing to obtain the completed Hi-C matrix;
[0013] S3. Construct a CNN network model, using the DNA-DNA feature correlation coefficient matrix of the fusion RNA information and the completed Hi-C matrix as the input of the model, train the model, and the output of the model is the enhanced Hi-C matrix;
[0014] S4. For the trained model, input data of different RNA types, and predict the corresponding Hi-C matrix for downstream analysis.
[0015] Optionally, the step S1 specifically includes:
[0016] Obtain the iMARGI data and the original Hi-C data from a public database or experimental data; Perform quality control on the obtained original data to remove noise data.
[0017] Optionally, the step S2 specifically includes:
[0018] Perform feature extraction on the RNA-DNA interaction data, and calculate the Pearson correlation coefficient for the extracted features to obtain the DNA-DNA feature correlation coefficient matrix of the fusion RNA information;
[0019] Perform KR normalization on the original Hi-C data, perform smoothing processing on the KR-normalized Hi-C data at downsampling rates of 1 / 25 and 1 / 100, and use the FAN method to perform completion processing on the downsampled Hi-C data to obtain the completed Hi-C matrix.
[0020] Optionally, in the step S3, the input of the model includes two parts: the DNA-DNA feature correlation coefficient matrix of the fusion RNA information and the completed Hi-C matrix;
[0021] Among them, the sample division is performed in the following way: Divide the DNA-DNA feature correlation coefficient matrix of the fusion RNA information and the completed Hi-C matrix into n submatrices of K×K dimensions as input samples;
[0022] Divide the input samples into a training set, a validation set, and a test set, which are used for model training, validation, and evaluation respectively.
[0023] Optionally, in step S3, for the completed Hi-C matrix, downsampling is performed using the DoubleConv module twice plus the DownConv layer, upsampling is performed using the DoubleConv module twice plus the UpConv layer, and finally a first result is obtained through a DoubleConv module;
[0024] For the DNA-DNA feature correlation coefficient matrix integrating RNA information, a second result is obtained using the DoubleConv module three times;
[0025] After the first result and the second result are merged, and then through two DoubleConv operations, the finally enhanced Hi-C matrix is obtained.
[0026] Optionally, in step S3, the error between the predicted Hi-C matrix and the actual Hi-C matrix is measured by a loss function;
[0027] The model weights are optimized by combining the L1 loss and the perceptual loss to handle the data imbalance problem and improve the performance of the model on sparse data;
[0028] The Adam optimizer is used to train the model, and through cross-validation and parameter tuning, the generalization ability of the model is improved.
[0029] Optionally, step S3 further includes:
[0030] The model is applied to data of different RNA types, with the preprocessed RNA data as the input and the predicted chromatin interaction enhancement matrix, i.e., the enhanced Hi-C matrix, as the output; by comparing with the actual Hi-C matrix, the prediction accuracy and stability of the model are verified.
[0031] Optionally, in step S4, the downstream analysis includes: Hi-C chromatin loop analysis mediated by different RNA types, verification of CLIP-seq data and ChIP-seq data, enrichment of chromatin loops in subnuclear structures and EP sites, the influence of different types of RNA on chromatin hierarchical structures, exploring the influence of RNA on chromatin conformation through model attribution, exploring the influence of RNA on TAD boundaries and chromatin loop anchors through model attribution, and visualization analysis of RNA and chromatin higher-order structures.
[0032] In the downstream analysis, the Hi-C chromatin loop analysis mediated by different RNA types includes: identifying chromatin loops in the original Hi-C data and the predicted Hi-C data using the mustache tool, and performing veen analysis on the chromatin loops in the original data and the predicted data;
[0033] Validation of CLIP-seq data and ChIP-seq data includes: using CLIP-seq data to evaluate the specific co-localization of RBPs that bind to nascent RNAs near chromatin loop anchors mediated by different types of RNAs; combining public ChIP-seq data to analyze whether chromatin regulatory factors are enriched at the chromatin loop anchors in the prediction results.
[0034] Enrichment of chromatin loops in subnuclear structures and EP sites includes: calculating the enrichment of Hi-C data predicted by different types of RNAs in subnuclear structures and EP sites respectively, revealing that lncRNAs are frequently involved in the construction of promoter-enhancer interactions, and the roles of snRNAs and snoRNAs in the nucleolus and nuclear speckle regions.
[0035] The effects of different types of RNAs on chromatin hierarchical structures include: calculating the effects of different types of RNAs on chromatin loops, topologically associating domains, and A / B compartments of Hi-C data respectively, revealing that lncRNAs are involved in local fine regulation, while snRNAs / snoRNAs have advantages in long-distance interactions and larger-scale structural remodeling.
[0036] In the downstream analysis, exploring the effects of RNAs on chromatin conformation through model attribution includes: calculating the contribution characteristics of lncRNAs, mRNAs, snRNAs, and snoRNAs at TAD boundary regions and chromatin loop anchors, revealing the promoting effect of RNAs on chromatin loop structures, as well as negative or neutral regulation at TAD boundary regions, indicating that the role of RNAs in chromatin conformation is scale-dependent and more inclined to mediate small-scale and specific chromatin loop structures.
[0037] Exploring the effects of RNAs on TAD boundaries and chromatin loop anchors through model attribution includes: calculating the contribution characteristics of lncRNAs, mRNAs, snRNAs, and snoRNAs at TAD boundary regions and chromatin loop anchors, and showing interaction distribution maps with position information where chromatin interactions at a certain site are affected and not affected by RNAs.
[0038] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0039] (1) Achieving effective integration of RNA information into chromatin interaction modeling: By constructing a DNA-DNA feature correlation coefficient matrix that fuses RNA-DNA interaction features, the present invention regards RNAs as "endogenous factors" for chromatin conformation regulation, and for the first time transforms the asymmetric RNA-DNA interaction matrix system into a symmetric matrix that can be embedded in chromatin analysis, overcoming the problem that RNA localization information cannot directly reflect DNA-DNA interaction relationships.
[0040] (2) Enhancing the spatial resolution of Hi-C maps by integrating multi-omics data: By introducing a deep learning model based on a dual-channel convolutional neural network and combining RNA-DNA interaction information with the Hi-C contact matrix, the interaction signals in Hi-C data are effectively enhanced, especially showing stronger generalization ability and prediction accuracy in the context of low-resolution or sparse data, thus improving the modeling accuracy of chromatin spatial structures.
[0041] (3) Exhibiting good generalization ability across different RNA types: The model demonstrates stable prediction performance across various RNA types (such as lncRNA, mRNA, snRNA, snoRNA), capable of parsing the regulatory characteristics of different RNAs in chromatin structures, capturing their differential roles in different hierarchical structures (such as TAD, Loop, Compartment), and providing a theoretical basis for revealing the diversity of RNA regulatory mechanisms.
[0042] (4) Achieving model attribution to support quantitative functional analysis: The attribution analysis module in this invention can trace the source contributions of the model's predictions for specific chromatin structures, thereby enabling quantitative assessment of the regulatory effects of different RNAs in key regions such as TAD boundaries and chromatin loop anchors, revealing the dominant role of RNA in promoting chromatin loop formation and stability, as well as the potential mechanism of a certain negative regulatory effect at TAD boundaries, demonstrating the unique advantages of this method in functional attribution and scale-dependent analysis.
[0043] (5) Supporting visualization and cross-omics verification to enhance biological interpretability: By integrating CLIP-seq, ChIP-seq data, and subnuclear structure annotation information, it is verified whether RNA-mediated chromatin loop anchors are enriched in RNA-binding proteins and epigenetic markers; further visualizing the spatial enrichment characteristics of different RNA types in chromatin regulation makes the prediction results more biologically interpretable and verifiable.
[0044] (6) Promoting the refined development of chromatin three-dimensional structure and function research: This invention not only provides the ability to predict enhanced Hi-C maps but also clarifies the regulatory patterns of RNA in promoter-enhancer interactions, nucleolus / nuclear speckle enrichment, long-distance chromatin association, etc. through downstream analysis, providing key tool and method support for exploring the coupling relationship between RNA regulation of chromatin structure and function.
[0045] (7) Having broad applicability and scalability: The method of this invention is applicable to RNA-DNA interactions and Hi-C data in different cell types and various public databases, with good generality, and can be extended to multiple bioinformatics application scenarios such as epigenetic state prediction, regulatory element function identification, and RNA molecule function annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0047] Figure 1 is a schematic flowchart of a chromatin interaction information enhancement method based on RNA-DNA and DNA-DNA interaction information provided by an embodiment of the present invention;
[0048] Figure 2 is a schematic diagram of the calculation of the correlation coefficient matrix provided by an embodiment of the present invention;
[0049] Figure 3 is a schematic diagram of sample division provided by an embodiment of the present invention;
[0050] Figure 4 is a schematic diagram of the network model structure provided by an embodiment of the present invention;
[0051] Figure 5 is a curve graph showing the change of the model loss function with the number of training rounds provided by an embodiment of the present invention;
[0052] Figure 6 is a schematic diagram of the coincidence between the Hi-C prediction results of four RNA types and the chromatin loops of the original data provided by an embodiment of the present invention;
[0053] Figure 7 is a co-localization clustering heat map of ChIP-seq data of different epigenetic signals at the chromatin loop anchors in the original data and the Hi-C data predicted by four RNAs provided by an embodiment of the present invention;
[0054] Figure 8 is a co-localization clustering heat map of different CLIP-seq data at the chromatin loop anchors in the original Hi-C data and the Hi-C data predicted by four RNAs provided by an embodiment of the present invention;
[0055] Figure 9 is a schematic diagram of the enrichment of subnuclear structures and EP provided by an embodiment of the present invention;
[0056] Figure 10 is a schematic diagram of the distribution characteristics of Hi-C data of different RNA types at different scales provided by an embodiment of the present invention;
[0057] Figure 11 is a schematic diagram of the influence of RNA on TAD boundaries provided by an embodiment of the present invention;
[0058] Figure 12 It is a schematic diagram of the influence of RNA on chromatin loop anchors provided by an embodiment of the present invention;
[0059] Figure 13 It is a schematic diagram of the positional distribution of the influence of RNA on TAD boundary interactions provided by an embodiment of the present invention;
[0060] Figure 14 It is a schematic diagram of the positional distribution of the influence of RNA on chromatin loop anchors provided by an embodiment of the present invention;
[0061] Figure 15 It is a multi-omics visualization of the influence of lncRNA on the SMYD3 locus provided by an embodiment of the present invention: a schematic diagram of ChIP-seq co-localization;
[0062] Figure 16 It is a multi-omics visualization of the influence of lncRNA on the SMYD3 locus provided by an embodiment of the present invention: a schematic diagram of CLIP-seq co-localization. Detailed implementation manners
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] An embodiment of the present invention provides a method for enhancing chromatin interaction information based on RNA-DNA and DNA-DNA interaction information. This method implements a prediction model based on deep learning (HiClip model), which can infer the Hi-C map affected by RNA from RNA-DNA interaction data and Hi-C data.
[0065] Refer to Figure 1 As shown, this method includes the following steps:
[0066] S1. Collect iMARGI data of RNA-DNA interactions and original Hi-C data of DNA-DNA interactions of the same cell line.
[0067] In this step, the iMARGI data and the original Hi-C data can be obtained from public databases or experimental data, and quality control can be performed on the obtained original data to remove noise data and low-quality data.
[0068] S2. After extracting the features of the iMARGI data, calculate the Pearson correlation coefficient to obtain the DNA-DNA feature correlation coefficient matrix of the genome-wide fusion RNA-DNA interaction information; downsample the original Hi-C data and then perform FAN completion processing to obtain the completed Hi-C matrix.
[0069] Specifically, as shown in Figure 2, extract the features of the RNA-DNA interaction data (iMARGI data), and calculate the Pearson correlation coefficient for the extracted features to obtain the DNA-DNA feature correlation coefficient matrix that integrates RNA information.
[0070] Perform KR normalization on the original Hi-C data to eliminate technical noise and batch effects; use downsampling rates of 1 / 25 and 1 / 100 to smooth the KR-normalized Hi-C data to ensure sufficient features are captured from the low-resolution matrix; use the FAN method to perform completion processing on the downsampled Hi-C data to obtain the completed Hi-C matrix.
[0071] It should be noted that the Hi-C data, Hi-C matrix, and Hi-C map mentioned in the embodiments of the present invention can sometimes be used interchangeably. When not emphasizing their differences, the meanings they express are the same.
[0072] S3. Construct a CNN network model, use the DNA-DNA feature correlation coefficient matrix that integrates RNA information and the completed Hi-C matrix as the input of the model, train the model, and the output of the model is the enhanced Hi-C matrix.
[0073] The input of the CNN network model includes two parts: the DNA-DNA feature correlation coefficient matrix that integrates RNA information and the completed Hi-C matrix.
[0074] Among them, the sample division is carried out in the following way: divide the DNA-DNA feature correlation coefficient matrix that integrates RNA information and the completed Hi-C matrix into n K×K-dimensional sub-matrices as input samples, as shown in Figure 3 Figure. Divide the input samples into a training set, a validation set, and a test set, which are used for model training, validation, and evaluation respectively.
[0075] The structure of the model is as shown in Figure 4As shown in the figure, for the completed Hi-C matrix, downsampling is performed using the DoubleConv module twice followed by the DownConv layer, and upsampling is performed using the DoubleConv module twice followed by the UpConv layer. Finally, a first result c5 is obtained through a DoubleConv module. For the DNA-DNA feature correlation coefficient matrix integrating RNA information, a second result e3 is obtained using the DoubleConv module three times. After merging the first result c5 and the second result e3, and then passing through the DoubleConv twice, the final enhanced Hi-C matrix is obtained.
[0076] The constructed CNN network model above is trained, and the error between the predicted Hi-C matrix and the actual Hi-C matrix is measured through the loss function.
[0077] Specifically, the model weights are optimized by combining the L1 loss and the perceptual loss to handle the data imbalance problem and improve the performance of the model on sparse data. The Adam optimizer is used to train the model, and through cross-validation and parameter tuning, the generalization ability of the model is improved.
[0078] The prediction and verification of the model include: applying the model to data of different RNA types, with the preprocessed RNA data as the input and the predicted chromatin interaction enhancement matrix, i.e., the enhanced Hi-C matrix, as the output. The prediction accuracy and stability of the model are verified by comparing with the actual Hi-C matrix.
[0079] Among them, the experimental results on data of different RNA types are as Figure 5 shown. Through cross-validation and multiple experiments, it can be ensured that the model has good stability and accuracy.
[0080] S4. For the trained model, data of different RNA types are input, and the corresponding (enhanced) Hi-C matrix is predicted for downstream analysis.
[0081] In the embodiments of the present invention, the downstream analysis includes: Hi-C chromatin loop analysis mediated by different RNA types, verification of CLIP-seq data and ChIP-seq data, enrichment of chromatin loops in subnuclear structures and EP sites, the impact of different types of RNA on chromatin hierarchical structures, exploring the impact of RNA on chromatin conformation through model attribution, exploring the impact of RNA on TAD boundaries and chromatin loop anchors through model attribution, and visual analysis of the high-order structure of RNA and chromatin.
[0082] Among them, the Hi-C chromatin loop analysis mediated by different RNA types includes: identifying chromatin loops in the original Hi-C data and the predicted Hi-C data using the mustache tool, performing a veen analysis on the chromatin loops of the original data and the predicted data, and the results are as Figure 6 shown.
[0083] The verification of CLIP-seq data and ChIP-seq data includes: using CLIP-seq data to evaluate whether there is specific co-localization of RBPs that bind to nascent RNAs near the chromatin loop anchors mediated by different types of RNAs; combining public ChIP-seq data to analyze whether more chromatin regulatory factors (such as transcription factors, histone modification marks) are enriched at the chromatin loop anchors in the prediction results to further verify the accuracy and reliability of the model, as Figure 7 and Figure 8 shown.
[0084] The enrichment of chromatin loops in subnuclear structures and EP sites includes: calculating the enrichment of the predicted Hi-C data of different types of RNAs in subnuclear structures and EP sites respectively, revealing the frequent participation of lncRNAs in the construction of promoter-enhancer interactions, and the important roles of snRNAs and snoRNAs in the nucleolus and nuclear speckle regions, as Figure 9 shown.
[0085] The effects of different types of RNAs on chromatin hierarchical structures include: calculating the effects of different types of RNAs on chromatin loops, topologically associating domains, and A / B compartments of Hi-C data respectively, revealing that lncRNAs mainly participate in local fine regulation, while snRNAs / snoRNAs have obvious advantages in long-distance interactions and larger-scale structural remodeling, as Figure 10 shown.
[0086] The present invention also designs a model attribution module. Specifically, exploring the effects of RNAs on chromatin conformation through model attribution includes: calculating the contribution characteristics of lncRNAs, mRNAs, snRNAs, and snoRNAs at TAD boundary regions and chromatin loop anchors, revealing the promoting effects of RNAs in chromatin loop structures, and there is certain negative or neutral regulation at TAD boundary regions, indicating that the effects of RNAs on chromatin conformation have scale dependence and are more inclined to mediate small-scale, specific chromatin loop structures, as Figure 11 and Figure 12 shown.
[0087] Investigating the impact of RNA on TAD boundaries and chromatin loop anchors through model attribution includes: calculating the contribution characteristics of lncRNA, mRNA, snRNA, and snoRNA at TAD boundary regions and chromatin loop anchors, and presenting interaction distribution maps with location information showing whether chromatin interactions at a certain site are affected by RNA or not. As Figure 13 and Figure 14 shown, there are more positions where TAD boundaries are not affected by RNA, and they are distributed in the lower left of this chromatin interaction site and within the genome of this chromatin interaction, while fewer chromatin loop anchors are not affected by RNA.
[0088] Furthermore, in the visualization analysis of RNA and higher-order chromatin structures, taking the SMYD3 gene as an example, it shows the promoting effect of lncRNA on specific chromatin sites and its possible molecular mechanism, further verifying that lncRNA actively regulates the three-dimensional chromatin structure by mediating promoter-enhancer and CTCF-dependent chromatin interactions; while in the transcription factory region, lncRNA does not participate in forming obvious chromatin interactions, but may participate in the post-transcriptional processing and modification of nascent RNA by aggregating RBPs, as Figure 15 and Figure 16 shown.
[0089] In summary, the present invention provides a method for enhancing chromatin interaction information based on RNA-DNA and DNA-DNA interaction information. The core lies in introducing RNA as an endogenous influencing factor regulating chromatin spatial structure into the chromatin interaction modeling process, accurately predicting and enhancing the Hi-C map through a deep learning model, and further analyzing the specific regulatory effects of RNA on multi-scale chromatin hierarchical structures (such as chromatin loops, TADs, compartments, etc.).
[0090] Compared with the prior art, the method of the present invention has the following advantages:
[0091] The present invention constructs a DNA-DNA feature correlation coefficient matrix integrating RNA-DNA interaction features, regards RNA as an "endogenous factor" for chromatin conformation regulation, and for the first time transforms the asymmetric RNA-DNA interaction matrix system into a symmetric matrix that can be embedded in chromatin analysis, overcoming the problem that RNA localization information cannot directly reflect the DNA-DNA interaction relationship.
[0092] The present invention introduces a deep learning model based on a two-channel convolutional neural network, combines RNA-DNA interaction information with the Hi-C contact matrix, effectively enhances the interaction signal of Hi-C data, especially shows stronger generalization ability and prediction accuracy in the background of low-resolution or sparse data, and improves the modeling accuracy of chromatin spatial structure.
[0093] The prediction model of the present invention exhibits stable prediction performance on various types of RNA (such as lncRNA, mRNA, snRNA, snoRNA), can analyze the regulatory characteristics of different RNAs in chromatin structure, capture their functional differences in different hierarchical structures (such as TAD, Loop, Compartment), and provide a theoretical basis for revealing the diversity of RNA regulatory mechanisms.
[0094] The present invention designs an attribution analysis module, which can trace the source contribution of the model to the prediction of specific chromatin structures, thereby realizing the quantitative evaluation of the regulatory effects of different RNAs in key regions such as TAD boundaries and chromatin loop anchors, revealing the dominant role of RNA in promoting chromatin loop formation and stability, as well as the potential mechanism of certain negative regulatory effects at TAD boundaries, and reflecting the unique advantages of this method in functional attribution and scale-dependent analysis.
[0095] The present invention combines CLIP-seq, ChIP-seq data and subnuclear structure annotation information to verify whether RNA-mediated chromatin loop anchors are enriched in RNA-binding proteins and epigenetic marks; further visualize the spatial enrichment characteristics of different RNA types in chromatin regulation, making the prediction results have higher biological interpretability and verifiability.
[0096] The present invention not only provides the ability to predict enhanced Hi-C maps, but also clarifies the regulatory patterns of RNA in promoter-enhancer interactions, nucleolus / nuclear speckle enrichment, long-distance chromatin association, etc. through downstream analysis, providing key tool and method support for exploring the coupling relationship between RNA regulation of chromatin structure and function.
[0097] The present invention is applicable to RNA-DNA interactions and Hi-C data in different cell types and various public databases, has good generality, and can be extended to a variety of bioinformatics application scenarios such as epigenetic state prediction, regulatory element function identification, and RNA molecule function annotation.
[0098] It should be noted that in this article, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device including the said element.
[0099] The terms "one embodiment", "an embodiment", "an exemplary embodiment", "some embodiments", etc. in the specification indicate that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when describing a specific feature, structure or characteristic in combination with an embodiment, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0100] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0101] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0102] It should be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0103] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0104] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0105] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit.
[0106] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0107] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the essence and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components, and circuits, etc., are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for enhancing chromatin interaction information based on RNA-DNA and DNA-DNA interaction information, characterized in that: The following steps are involved: S1. Collect iMARGI data of RNA-DNA interactions and original Hi-C data of DNA-DNA interactions of the same cell line; S2. After feature extraction, the Pearson correlation coefficient of the iMARGI data is calculated to obtain a DNA-DNA feature correlation coefficient matrix of the fusion RNA information of the whole genome; after downsampling the original Hi-C data, FAN completion processing is performed to obtain a complete Hi-C matrix; S3, constructing a CNN network model, taking the DNA-DNA feature correlation coefficient matrix of the fused RNA information and the complete Hi-C matrix as the input of the model, training the model, and outputting the model as the enhanced Hi-C matrix; S4. For the trained model, input data of different RNA types and predict the corresponding Hi-C matrix for downstream analysis.
2. The method according to claim 1, characterized in that: The step S1 specifically includes: The iMARGI data and the original Hi-C data are obtained from public databases or experimental data; quality control is performed on the obtained original data to remove noise data.
3. The method according to claim 1, characterized in that The step S2 specifically includes: Feature extraction was performed on RNA-DNA interaction data, and the Pearson correlation coefficient of the extracted features was calculated to obtain a DNA-DNA feature correlation coefficient matrix that fused RNA information. The original Hi-C data were KR normalized, and the KR normalized Hi-C data were smoothed using downsampling rates of 1 / 25 and 1 / 100. The downsampled Hi-C data were completed using the FAN method to obtain a complete Hi-C matrix.
4. The method according to claim 1, characterized in that: In step S3, the input of the model includes two parts: a DNA-DNA feature correlation coefficient matrix of fused RNA information and a complete Hi-C matrix; The sample division is performed in the following manner: the DNA-DNA feature correlation coefficient matrix of the fused RNA information and the complete Hi-C matrix are divided into n K×K dimensional sub-matrices as input samples; The input samples are divided into training set, validation set and test set for model training, validation and evaluation respectively.
5. The method according to claim 1, characterized in that In step S3, for the completed Hi-C matrix, two DoubleConv modules plus a DownConv layer are used for downsampling, two DoubleConv modules plus an UpConv layer are used for upsampling, and finally a DoubleConv module is used to obtain the first result; For the DNA-DNA feature correlation coefficient matrix fused with RNA information, the second result was obtained by using the DoubleConv module three times; After merging the first result and the second result, the final enhanced Hi-C matrix is obtained through two DoubleConv operations.
6. The method according to claim 1, characterized in that In step S3, the error between the predicted Hi-C matrix and the actual Hi-C matrix is measured by a loss function; Combine L1 loss and perceptual loss to optimize model weights to handle data imbalance and improve model performance on sparse data; The Adam optimizer is used to train the model, and the generalization ability of the model is improved through cross-validation and parameter tuning.
7. The method according to claim 1, characterized in that The step S3 further comprises: The model was applied to data of different RNA types, with the input being the preprocessed RNA data and the output being the predicted chromatin interaction enhancement matrix, i.e., the enhanced Hi-C matrix; the prediction accuracy and stability of the model were verified by comparison with the actual Hi-C matrix.
8. The method according to claim 1, characterized in that In step S4, downstream analysis includes: Hi-C chromatin loop analysis mediated by different RNA types, CLIP-seq data and ChIP-seq data verification, enrichment of chromatin loops in sub-nuclear structures and EP sites, effects of different types of RNA on chromatin hierarchical structure, exploration of the effects of RNA on chromatin conformation through model attribution, exploration of the effects of RNA on TAD boundaries and chromatin loop anchors through model attribution, and visualization analysis of RNA and chromatin high-order structure.
9. The method according to claim 8, characterized in that In the downstream analysis, the Hi-C chromatin loop analysis mediated by different RNA types includes: identifying chromatin loops using the mustache tool on the original Hi-C data and the predicted Hi-C data, and performing veen analysis on the chromatin loops of the original data and the chromatin loops of the predicted data; CLIP-seq data and ChIP-seq data validation include: using CLIP-seq data to evaluate whether RBPs bound to nascent RNAs specifically co-localize near chromatin loop anchors mediated by different types of RNAs; combining public ChIP-seq data to analyze whether more chromatin regulatory factors are enriched at the chromatin loop anchors in the predicted results; The enrichment of chromatin loops in subcellular nuclear structures and EP sites includes: calculating the enrichment of Hi-C data predicted by different types of RNA in subcellular nuclear structures and EP sites, revealing that lncRNAs frequently participate in the construction of promoter-enhancer interactions, and the roles of snRNAs and snoRNAs in nucleolus and nuclear speckle regions; The effects of different types of RNA on the hierarchical structure of chromatin include: calculating the effects of different types of RNA on chromatin loops, topological domains and AB compartments of Hi-C data, revealing that lncRNA is involved in local fine regulation, while snRNA / snoRNA has advantages in long-distance interactions and larger-scale structural remodeling.
10. The method according to claim 8, characterized in that In the downstream analysis, the influence of RNA on chromatin conformation was explored through model attribution, including: calculating the contribution characteristics of lncRNA, mRNA, snRNA and snoRNA in TAD boundary regions and chromatin loop anchor points, revealing the promotion effect of RNA in chromatin loop structure, as well as negative or neutral regulation in TAD boundary regions, indicating that the effect of RNA on chromatin conformation is scale-dependent, and is more inclined to mediate small-scale, specific chromatin loop structures; The influence of RNA on TAD boundaries and chromatin loop anchors was explored through model attribution, including: calculating the contribution characteristics of lncRNA, mRNA, snRNA and snoRNA in TAD boundary regions and chromatin loop anchors, and displaying the interaction distribution map containing position information of chromatin interactions at a certain site affected by RNA and not affected by RNA.
Citation Information
Patent Citations
RNA and protein binding site recognition method based on deep learning
CN113178229A
Protein-DNA-RNA triad and computer prediction method and system for functions of protein-DNA-RNA triad
CN118116465A
Single-cell Hi-C map prediction method based on single-cell RNA expression data
CN118645154A
Construction method of TaDRIM-seq sequencing library
CN118834932A
Methods of changing transcriptional output
US20210163930A1