A method for quantitatively identifying the interaction strength of genes

By combining Hi-C data and annotation files, the interaction intensity of transformed chromatin is the interaction intensity between genes, which solves the problem that it is difficult to quantitatively identify gene interaction intensity in the existing technology, and realizes quantitative evaluation of intergenic interaction relationships and strength, providing technical support for gene function research and the construction of gene regulation networks.

CN119943159BActive Publication Date: 2025-06-13NANCHANG CAMPUS OF EAST CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510425447.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-13
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing Hi-C data analysis methods are difficult to effectively quantify the interaction intensity between genes, and there is a lack of methods to identify direct interactions and interaction intensity between genes.

Method used

By using Hi-C data and combining annotation files, the chromatin interaction intensity is converted into the interaction intensity between genes based on the three-dimensional structural data of chromatin. The steps of Hi-C library construction, sequencing, quality control, alignment, cleavage and re-comparison are adopted, and the reads count of the gene-related regions of self-built python scripts were used for quantitative analysis.

Benefits of technology

Quantitative evaluation of the intensity of interactions between genes is achieved, and interaction relationships between different genes can be identified and the intensity of these interactions is quantified, providing reliable technical support for gene function research and the construction of gene regulatory networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943159B_ABST
    Figure CN119943159B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for quantitatively identifying the interaction strength of genes, comprising the following steps: S1, construction and sequencing of Hi-C libraries; S2, data quality control and alignment; S3, cutting and re-alignment of unaligned reads; S4, counting reads in gene-related regions; S5, quantitative analysis of interaction strength; By converting chromosomal interaction data into the interaction strength between genes, the present invention can not only identify the interaction relationships between different genes, but also quantify the strength of these interactions, providing a quantitative analysis means for in-depth understanding of gene regulatory networks and functional research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and particularly to a method for quantitatively identifying the interaction strength between genes. Background Art

[0002] With the rapid development of genomics and molecular biology technologies, the importance of three-dimensional structural information at the genomic level in gene function and gene regulation research has been increasing. As a powerful tool for studying the three-dimensional structure of chromosomes, Hi-C technology can generate interaction data between all regions in the genome, thereby revealing the spatial organization characteristics of chromatin.

[0003] However, traditional Hi-C data analysis mainly focuses on the interactions between chromosomal regions or specific genomic regions, calculating the interactions between chromosomal regions using a sliding window method, and there is a lack of effective quantitative methods for the direct interactions between genes and the strength of these interactions.

[0004] Therefore, we need to design a method for quantitatively identifying the interaction strength between genes to solve the above problems. Summary of the Invention

[0005] In view of the above problems, the present invention provides a new method for quantitatively identifying the interaction strength between genes by using Hi-C data and combining with an annotation file. This method is based on chromatin three-dimensional structure data and combines detailed information about genes and gene regulatory regions in the annotation file, and can convert chromatin interaction strength into the interaction strength between genes. The method of the present invention can not only identify gene interactions, but also quantify these interactions, providing reliable technical support for fields such as gene function research, construction of gene regulatory networks, and genetic breeding.

[0006] According to one aspect of the present invention, there is provided a method for quantitatively identifying the interaction strength between genes, comprising the following steps:

[0007] S1. Construction and sequencing of Hi-C library

[0008] Construct a Hi-C library for the sample to be analyzed, and use a high-throughput sequencing platform to obtain Hi-C data. The obtained data are paired-end sequencing reads;

[0009] S2. Data quality control and alignment

[0010] Perform quality control on the obtained Hi-C sequencing data, filter out low-quality reads, and use bowtie2 software to align the quality-controlled reads to the reference genome to obtain the first alignment result, and retain the unaligned reads;

[0011] S3. Cutting and re-aligning the unaligned reads

[0012] Cut the unaligned reads at the restriction enzyme cleavage sites used in Hi-C library construction. Re-align the remaining reads to the reference genome using the bowtie2 software to obtain the second alignment result. Then, use the samtools software to merge the first and second alignment results to generate a new bam file and sort it;

[0013] S4. Counting reads in gene-related regions

[0014] Use the annotation file to label gene-related regions, which include the promoter region, downstream enhancer region, gene body region, CDS region, exon region, and intron region of the gene. Statistically analyze the positions according to the region coordinates. Use a self-built python script to screen the bam file. According to the gene regulation mode, ensure that one end of the paired reads is located in one or more of the above regions of a certain gene, and the other end is located in one or more of the above regions of other genes. Count the reads count or the normalized reads count located in these gene regions to evaluate the interaction regulation strength between a certain gene and other genes.

[0015] In some embodiments, the method further includes the following steps:

[0016] S5. Quantitative analysis of interaction strength

[0017] Quantify the interaction strength between different genes through reads count or standardize according to the quantification data, and draw an interaction network between genes based on the statistical results to provide a basis for the analysis of the gene regulatory network.

[0018] In some embodiments, in step S2, use the bowtie2 software to align the quality-controlled reads to the reference genome. To ensure the alignment accuracy, the parameters are set as: --very-sensitive -L 30 --score-min L,-0.6, -0.2 --end-to-end --reorder.

[0019] In some embodiments, in step S3, re-align the remaining reads to the reference genome using the bowtie2 software, and the parameters are set as: --very-sensitive -L 20 --score-min L, -0.6, -0.2--end-to-end --reorder.

[0020] In some embodiments, in step S4, the normalized reads count is obtained by normalizing the reads count.

[0021] In some embodiments, in step S3, the annotated information is used as a location region to count the reads count or the normalized reads count of Hi-C data, for evaluating the interaction regulation strength from a certain gene to other genes.

[0022] Advantages of the present invention:

[0023] 1. By converting chromosome interaction data into the interaction strength between genes, the present invention can not only identify the interaction relationships between different genes, but also quantify the strength of these interactions, providing a quantitative analysis means for in-depth understanding of gene regulatory networks and functional research.

[0024] 2. The present invention improves the analysis method of Hi-C sequencing data, and further updates the calculation method of the interaction strength between chromosome segments based on resolution parameters to a new method that can reveal the interaction frequency between single genes, expanding the research on gene functions from the quantity of single or several genes to the conditions of studying regulatory networks.

[0025] 3. This method is simple to operate, applicable to Hi-C data analysis of different species, and can be widely applied to genomics research. Description of the Drawings

[0026] Figure 1 is a flowchart of the present invention;

[0027] Figure 2 is a heatmap visualization of the bipartite structure of the mouse X chromosome drawn by Hi-Browse using in situ DNase Hi-C data;

[0028] Figure 3 is a flowchart of the prior art;

[0029] Figure 4 is the inter-gene interaction network of the present invention. Detailed Embodiments

[0030] The present invention will be further described in detail below with reference to the embodiments.

[0031] Reference Figure 1, the present invention provides a method for quantitatively identifying the interaction strength of genes, comprising the following steps: First, construct and sequence a Hi-C library; align the obtained paired-end sequences to the reference genome respectively; then count the number of reads related to genes and gene regulatory regions according to the annotation file, so as to convert the regional interaction strength of chromatin into the interaction strength between genes. In this way, not only can the interaction relationships between different genes be identified, but also the strengths of these interactions can be quantified.

[0032] Specifically, it includes the following steps:

[0033] S1. Construction and sequencing of Hi-C library

[0034] Construct a Hi-C library for the sample to be analyzed, and use a high-throughput sequencing platform to obtain Hi-C data. The obtained data are paired-end sequencing reads;

[0035] S2. Data quality control and alignment

[0036] Perform quality control on the obtained Hi-C sequencing data, filter out low-quality reads, and use the bowtie2 software to align the quality-controlled reads to the reference genome. In order to ensure the alignment accuracy, the parameter settings '--very-sensitive -L 30 --score-min L, -0.6, -0.2 --end-to-end --reorder' are adopted to obtain the first alignment result, and the unaligned reads are retained;

[0037] S3. Cutting and re-aligning of unaligned reads

[0038] Cut the unaligned reads at the restriction enzyme cleavage sites used in Hi-C library construction, and re-align the cut reads to the reference genome using the bowtie2 software with the parameter settings '--very-sensitive -L 20 --score-min L, -0.6, -0.2 --end-to-end --reorder' to obtain the second alignment result. Then use the samtools software to merge the first alignment result and the second alignment result to generate a new bam file and sort it;

[0039] S4. Reads counting in gene regions

[0040] Reads are counted according to the gene promoter region, downstream enhancer region, gene body region, CDS region, exon region, and intron region annotated in the annotation file. The positions are statistically counted according to the region coordinates, and a self-built Python script is used to screen the reads in the bam file to ensure that one end is located in the above-mentioned region of a certain gene and the other end is located in the corresponding region of other genes, so as to count the number of interacting reads between genes.

[0041] S5. Quantitative analysis of interaction strength

[0042] The interaction strength between different genes is quantified by the number of reads or standardized according to the quantified data, and then an interaction network between genes is drawn according to the statistical results, providing a basis for the analysis of the gene regulatory network. Example

[0043] The Hi-C technique was first proposed by the research team of Professor Job Dekker at the University of Massachusetts Medical School in the United States in 2009. The interpretation of Hi-C data is challenging because the data is large and cannot be presented using a standard genome browser. An effective Hi-C visualization tool must provide multiple visualization modes and be able to view the data in combination with existing complementary data.

[0044] A key parameter in the Hi-C analysis pipeline is the effective resolution of the analyzed data, and the resolution is generally in the size of 1-10 Mb. In the prior art, this method is often combined with the sliding window technique to display the interaction level between genomic chromosomal fragments.

[0045] It can be seen from Figures 2 - 3 that for the Hi-C data analyzed by this resolution-based method, only the interaction strength between chromosomal fragments based on the resolution parameter can be presented, and there is no way to calculate the direct interaction frequency between genes.

[0046] Taking mice as the research species, the method of the present invention is used to calculate the interaction frequency of gene A regulating gene B.

[0047] S1. Construction and sequencing of Hi-C library

[0048] Construct a Hi-C library for the sample to be analyzed, and use a high-throughput sequencing platform to obtain Hi-C data. The obtained data is paired-end sequencing reads;

[0049] S2. Data quality control and alignment

[0050] The obtained Hi-C sequencing data was subjected to quality control to filter out low-quality reads. The quality-controlled reads were aligned to the mouse genome sequence using the bowtie2 software. To ensure the alignment accuracy, the parameter settings '--very-sensitive -L 30 --score-min L, -0.6, -0.2 --end-to-end --reorder' were adopted to obtain the first alignment result, and the unaligned reads were retained;

[0051] S3. Cutting and realignment of unaligned reads

[0052] The unaligned reads were cut using the restriction enzyme cleavage sites used during Hi-C library construction. The cut reads were realigned to the mouse genome sequence using the bowtie2 software with the parameter settings '--very-sensitive -L 20 --score-min L, -0.6, -0.2 --end-to-end --reorder' to obtain the second alignment result. Then, the samtools software was used to merge the first and second alignment results to generate a new bam file and sort it;

[0053] S4. Reads counting in gene regions

[0054] Reads counting was performed according to the gene promoter region, downstream enhancer region, gene body region, CDS region, exon region, and intron region annotated in the annotation file. The positions were statistically counted according to the region coordinates. Reads in the bam file were screened through a self-built python script to ensure that one end was located in the above-mentioned region of a certain gene and the other end was located in the corresponding region of another gene to statistically count the number of interacting reads between genes (see Table 1). For example, the number of interacting reads between the gene body of gene A and the promoter region of gene B (defining the promoter regulatory region as 2 - 4 kb before the translation start site) was calculated as the interaction frequency of gene A regulating gene B.

[0055] Table 1. Reads count from gene to gene regulatory region based on the annotation file

[0056] chromosome strand promoter start promoter end enhancer start enhancer end gene start gene end gene name chr14 + 55204104 55206103 55211921 55213920 55206104 55211920 Mhrt chr14 - 55232084 55234083 55206141 55208140 55208141 55232083 Myh7 chr14 + 55208486 55210485 55210817 55212816 55210486 55210816 Gm29015 chr14 + 55214850 55216849 55229059 55231058 55216850 55229058 Gm31251

[0057] S5. Quantitative analysis of interaction intensity

[0058] The interaction intensity between different genes (e.g., gene A and gene B) was quantified by the number of reads or standardized according to the quantified data (see Table 2). Subsequently, an interaction network between genes was drawn based on the statistical results (such as Figure 4as shown in the figure), providing a basis for the analysis of gene regulatory networks.

[0059] Table 2 Interaction intensity between genes

[0060] regulatory chromosome regulatory gene regulated chromosome regulated gene interaction frequency chr6 Ptn chr14 Myh7 20 chr6 Dgki chr14 Myh7 51 chr14 Gm49130 chr14 Myh7 27 chr14 Mhrt chr14 Myh7 1292 chr14 Gm31251 chr14 Myh7 338

[0061] The present invention transforms the interaction intensity of chromosomal regions into the interaction intensity between genes, realizing the quantitative evaluation of the interaction intensity between genes. It provides a practical technical approach for gene function research, genetic breeding, and the construction of gene regulatory networks.

[0062] The above are only some embodiments of the present invention. For those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention.

Claims

1. A method for quantitatively identifying the strength of gene interactions, characterized in that: The following steps are involved: S1. Hi-C library construction and sequencing Construct Hi-C libraries for the samples to be analyzed, and use a high-throughput sequencing platform to obtain Hi-C data. The data obtained are double-end sequencing reads; S2. Data quality control and comparison The obtained Hi-C sequencing data was quality controlled to filter out low-quality reads, and the reads after quality control were aligned to the reference genome using bowtie2 software to obtain alignment result 1, and the unaligned reads were retained; S3. Cutting and realignment of unaligned reads The unaligned reads were cut using the restriction endonuclease sites used in Hi-C library construction, and the remaining reads were re-aligned to the reference genome using bowtie2 software to obtain alignment result 2. Then, alignment result 1 and alignment result 2 were merged using samtools software to generate a new bam file and sorted; S4. Read counts of gene-related regions Gene-related regions were annotated using annotation files. Gene-related regions included the promoter region, downstream enhancer region, gene body region, CDS region, exon region and intron region of the gene. The positions were counted according to the regional coordinates. A self-built python script was used to screen the bam files. According to the gene regulation pattern, one end of the paired reads was ensured to be located in one or more regions of a certain gene, and the other end was located in one or more regions of other genes. The read counts or standardized read counts located in these gene regions were counted to evaluate the strength of the interactive regulation from a certain gene to other genes.

2. The method for quantitatively identifying gene interaction strength according to claim 1, characterized in that: The method further comprises the following steps: S5. Quantitative analysis of interaction strength The interaction strength between different genes is quantified by reads count or standardized according to the quantitative data, and the gene interaction network is drawn according to the statistical results to provide a basis for the analysis of gene regulatory networks.

3. The method for quantitatively identifying gene interaction strength according to claim 1, characterized in that: In step S2, bowtie2 software was used to align the quality-controlled reads to the reference genome. To ensure alignment accuracy, the parameters were set as follows: --very-sensitive -L 30 --score-min L, -0.6, -0.2 --end-to-end --reorder.

4. The method for quantitatively identifying gene interaction strength according to claim 3, characterized in that: In step S3, the remaining reads were aligned to the reference genome using bowtie2 software again, with the parameters set as: --very-sensitive -L 20 --score-min L, -0.6, -0.2 --end-to-end --reorder.

5. The method for quantitatively identifying gene interaction strength according to claim 1, characterized in that: In the step S4, the normalized reads count is obtained by normalizing the reads count.

6. The method for quantitatively identifying gene interaction strength according to claim 1, characterized in that: In step S4, the annotation information is used as the position region to count the read counts of the Hi-C data or the normalized read counts, so as to evaluate the strength of the interactive regulation from a certain gene to other genes.

Citation Information

Patent Citations

  • Key gene for duck meat and egg character differentiation and method for excavating key gene based on whole genome sequencing

    CN119464325A

  • Methods for determining absolute genome-wide copy number variations of complex tumors

    EP2844771A1