Method for quantitatively identifying gene interaction strength

By combining Hi-C data and annotation files, the interaction intensity of transformed chromatin is the interaction intensity between genes, which solves the problem that it is difficult to quantitatively identify gene interaction intensity in the existing technology, and realizes the identification and quantification of intergenic interaction relationships, providing technical support for the research of gene functions and the construction of regulatory networks.

CN119943159AActive Publication Date: 2025-05-06NANCHANG CAMPUS OF EAST CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510425447.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing Hi-C data analysis methods are difficult to effectively quantify the interaction intensity between genes, and there is a lack of methods to identify direct interactions and interaction intensity between genes.

Method used

By using Hi-C data and combining annotation files, the chromatin interaction intensity is converted into the interaction intensity between genes based on the three-dimensional structural data of chromatin. The quantitative analysis of the interaction intensity between genes is achieved by using Hi-C library construction, sequencing, quality control, alignment and gene-related region reads counting steps.

Benefits of technology

The identification of interaction relationships between genes and the quantification of interaction intensity is achieved, providing reliable technical support for gene function research, the construction of gene regulation networks and genetic breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943159A_ABST
    Figure CN119943159A_ABST
Patent Text Reader

Abstract

The invention discloses a method for quantitatively identifying gene interaction strength. The method comprises the following steps: S1, constructing and sequencing a Hi-C library; s2, data quality control and comparison; s3, performing cutting and re-comparison on the reads which are not compared; s4, performing reads counting on a gene related region; s5, quantitative analysis of interaction strength; compared with the prior art, the chromosome interaction data is converted into the interaction strength between the genes, so that the interaction relationship between different genes can be identified, the interaction strength can be quantified, and a quantitative analysis means is provided for deeply understanding a gene regulation and control network and researching functions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular to a method for quantitatively identifying gene interaction strength. Background Art

[0002] With the rapid development of genomics and molecular biology technologies, the importance of three-dimensional structural information at the genome level in the study of gene function and gene regulation is increasing. As a powerful tool for studying the three-dimensional structure of chromosomes, Hi-C technology can generate interaction data between all regions in the genome, thereby revealing the spatial organization characteristics of chromatin.

[0003] However, traditional Hi-C data analysis mainly focuses on the interactions between chromosome regions or specific genomic regions, using a sliding window approach to calculate the interactions between chromosome regions. However, there is a lack of effective quantitative methods for the direct interactions between genes and the strength of the interactions.

[0004] Therefore, we need to design a method to quantitatively identify the strength of gene interactions to solve the above problems. Summary of the invention

[0005] In view of the above problems, the present invention provides a new method for quantitatively identifying the interaction strength between genes by using Hi-C data and combining with annotation files. This method is based on chromatin three-dimensional structure data, and combined with detailed information about genes and gene regulatory regions in annotation files, it can convert chromatin interaction strength into gene-gene interaction strength. The method of the present invention can not only identify the interaction between genes, but also quantify these interactions, providing reliable technical support for the fields of gene function research, construction of gene regulatory networks, and genetic breeding.

[0006] According to one aspect of the present invention, a method for quantitatively identifying gene interaction strength is provided, comprising the following steps: S1. Hi-C library construction and sequencing Construct Hi-C libraries for the samples to be analyzed, and use a high-throughput sequencing platform to obtain Hi-C data. The data obtained are double-end sequencing reads; S2. Data quality control and comparison The obtained Hi-C sequencing data was quality controlled to filter out low-quality reads, and the reads after quality control were aligned to the reference genome using bowtie2 software to obtain alignment result 1, and the unaligned reads were retained; S3. Cutting and realignment of unaligned reads The unaligned reads were cut using the restriction endonuclease sites used in Hi-C library construction, and the remaining reads were re-aligned to the reference genome using bowtie2 software to obtain alignment result 2. Then, alignment result 1 and alignment result 2 were merged using samtools software to generate a new bam file and sorted; S4. Read counts of gene-related regions Gene-related regions were annotated using annotation files. Gene-related regions included the promoter region, downstream enhancer region, gene body region, CDS region, exon region and intron region of the gene. The positions were counted according to the regional coordinates. A self-built python script was used to screen the bam files. According to the gene regulation pattern, one end of the paired reads was ensured to be located in one or more regions of a certain gene, and the other end was located in one or more regions of other genes. The read counts or standardized read counts located in these gene regions were counted to evaluate the strength of the cross-regulation from a certain gene to other genes.

[0007] In some embodiments, the method further comprises the following steps: S5. Quantitative analysis of interaction strength The interaction strength between different genes is quantified by reads count or standardized according to the quantitative data, and the gene interaction network is drawn according to the statistical results to provide a basis for the analysis of gene regulatory networks.

[0008] In some embodiments, in step S2, bowtie2 software is used to align the quality-controlled reads to the reference genome. To ensure alignment accuracy, the parameters are set as follows: --very-sensitive -L 30 --score-min L,-0.6, -0.2 --end-to-end --reorder.

[0009] In some embodiments, in step S3, the remaining reads after cutting are re-aligned to the reference genome using bowtie2 software, and the parameters are set as follows: --very-sensitive -L 20 --score-min L, -0.6, -0.2 --end-to-end --reorder.

[0010] In some embodiments, in step S4, the standardized reads count is obtained by performing a standardization process on the reads count.

[0011] In some embodiments, in step S3, the annotation information is used as the position region to count the read counts of Hi-C data or the normalized read counts to evaluate the strength of the interactive regulation from a gene to other genes.

[0012] Beneficial effects of the present invention: 1. The present invention converts chromosome interaction data into the interaction strength between genes, which can not only identify the interaction relationship between different genes, but also quantify the strength of these interactions, providing a quantitative analysis method for in-depth understanding of gene regulatory networks and functional research.

[0013] 2. The present invention improves the analysis method of Hi-C sequencing data, and further updates the calculation method of the interaction strength between chromosome fragments based on resolution parameters to a new method that can reveal the interaction frequency between single genes, expanding the study of gene function from a single or a few numbers to the conditions that can study regulatory networks.

[0014] 3. This method is simple to operate, suitable for Hi-C data analysis of different species, and can be widely used in genomics research. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flow chart of the present invention; Figure 2 Heatmap visualization of the bipartite structure of mouse chromosome X using in situ DNase Hi-C data for Hi-Browse; Figure 3 A flow chart of the prior art; Figure 4 This is the gene interaction network of the present invention. DETAILED DESCRIPTION

[0016] The present invention will be further described in detail below in conjunction with the embodiments.

[0017] refer to Figure 1 The present invention provides a method for quantitatively identifying the strength of gene interactions, comprising the following steps: first constructing and sequencing a Hi-C library; aligning the obtained double-end sequences to the reference genome; and then counting the number of reads related to genes and gene regulatory regions according to the annotation file, thereby converting the regional interaction strength of chromatin into the interaction strength between genes. In this way, not only can the interaction relationship between different genes be identified, but also the strength of these interactions can be quantified.

[0018] The specific steps include: S1. Hi-C library construction and sequencing Construct Hi-C libraries for the samples to be analyzed, and use a high-throughput sequencing platform to obtain Hi-C data. The data obtained are double-end sequencing reads; S2. Data quality control and comparison The obtained Hi-C sequencing data were quality controlled to filter out low-quality reads, and bowtie2 software was used to align the quality-controlled reads to the reference genome. To ensure the alignment accuracy, the parameters were set to '--very-sensitive -L 30 --score-min L, -0.6, -0.2 --end-to-end --reorder' to obtain alignment result 1, and the unaligned reads were retained; S3. Cutting and realignment of unaligned reads The unaligned reads were cut using the restriction endonuclease sites used in Hi-C library construction. The cut reads were re-aligned to the reference genome using bowtie2 software with parameters set to '--very-sensitive-L 20 --score-min L, -0.6, -0.2 --end-to-end --reorder' to obtain alignment result 2. Then, samtools software was used to merge alignment result 1 and alignment result 2 to generate a new bam file and sort it. S4. Gene region read counts Reads were counted according to the gene promoter region, downstream enhancer region, gene body region, CDS region, exon region and intron region marked in the annotation file. The positions were counted according to the regional coordinates, and the reads in the bam file were screened by a self-built python script to ensure that one end was located in the above region of a gene and the other end was located in the corresponding region of other genes, so as to count the number of interactive reads between genes.

[0019] S5. Quantitative analysis of interaction strength The interaction strength between different genes is quantified by the number of reads or standardized according to the quantitative data, and then the gene interaction network is drawn according to the statistical results to provide a basis for the analysis of the gene regulatory network. Example

[0020] Hi-C technology was first proposed by the research team of Job Dekker, a professor at the University of Massachusetts Medical School, in 2009. The interpretation of Hi-C data is challenging because the data is large and cannot be presented using a standard genome browser. Effective Hi-C visualization tools must provide multiple visualization modes and be able to view the data in conjunction with existing complementary data.

[0021] A key parameter in the Hi-C analysis pipeline is the effective resolution of the analysis data, which is generally 1-10Mb in size. In the prior art, this method is often combined with sliding window technology to display the level of interaction between genomic chromosome fragments.

[0022] Depend on Figure 2-Figure 3 It can be seen that for Hi-C data analyzed based on this resolution method, only the interaction strength between chromosome fragments based on resolution parameters can be presented, and there is no way to calculate the direct interaction frequency between genes.

[0023] In the following, mice are used as research species, and the method of the present invention is used to calculate the interaction frequency of gene A regulating gene B.

[0024] S1. Hi-C library construction and sequencing Construct Hi-C libraries for the samples to be analyzed, and use a high-throughput sequencing platform to obtain Hi-C data. The data obtained are double-end sequencing reads; S2. Data quality control and comparison The obtained Hi-C sequencing data were quality controlled to filter out low-quality reads, and bowtie2 software was used to align the quality-controlled reads to the mouse genome sequence. To ensure the alignment accuracy, the parameters were set to '--very-sensitive -L 30 --score-min L, -0.6, -0.2 --end-to-end --reorder' to obtain alignment result 1, and the unaligned reads were retained; S3. Cutting and realignment of unaligned reads The unaligned reads were cut using the restriction endonuclease sites used in Hi-C library construction. The cut reads were re-aligned to the mouse genome sequence using bowtie2 software with the parameters set to '--very-sensitive -L 20 --score-min L, -0.6, -0.2 --end-to-end --reorder' to obtain alignment result 2. Then, samtools software was used to merge alignment result 1 and alignment result 2 to generate a new bam file and sort it. S4. Gene region read counts Reads were counted according to the gene promoter region, downstream enhancer region, gene body region, CDS region, exon region and intron region marked in the annotation file. According to the regional coordinate statistical position, the reads in the bam file were screened by a self-built python script to ensure that one end was located in the above region of a gene and the other end was located in the corresponding region of other genes to count the number of interactive reads between genes (see Table 1). For example, the number of interactive reads of the gene body of gene A and the promoter region of gene B (the 2-4 kb before the translation start site is defined as the promoter regulatory region) was calculated as the interaction frequency of gene A regulating gene B.

[0025] Table 1. Count of reads from genes to gene regulatory regions based on annotation files chromosome chain Promoter start End of promoter Enhancer initiation Enhancer end Gene start Gene end Gene name chr14 + 55204104 55206103 55211921 55213920 55206104 55211920 Mh chr14 - 55232084 55234083 55206141 55208140 55208141 55232083 Myh7 chr14 + 55208486 55210485 55210817 55212816 55210486 55210816 Gm29015 chr14 + 55214850 55216849 55229059 55231058 55216850 55229058 Gm31251 S5. Quantitative analysis of interaction strength The interaction strength between different genes (e.g., gene A and gene B) is quantified by the number of reads or standardized according to the quantified data (see Table 2), and then the gene interaction network is drawn according to the statistical results (e.g. Figure 4 This provides a basis for the analysis of gene regulatory networks.

[0026] Table 2 Interaction strength between genes Regulating chromosomes Regulatory genes Regulated chromosome Controlled genes Interaction frequency chr6 Ptn chr14 Myh7 20 chr6 Dj chr14 Myh7 51 chr14 Gm49130 chr14 Myh7 27 chr14 Mh chr14 Myh7 1292 chr14 Gm31251 chr14 Myh7 338 The present invention converts the chromosome region interaction strength into the gene-gene interaction strength, realizes the quantitative evaluation of the gene-gene interaction strength, and provides a practical technical approach for gene function research, genetic breeding, and the construction of gene regulatory networks.

[0027] The above are only some embodiments of the present invention. For those skilled in the art, several modifications and improvements may be made without departing from the inventive concept of the present invention, which all fall within the protection scope of the present invention.

Claims

1. A method for quantitatively identifying the strength of gene interactions, characterized in that: The following steps are involved: S1. Hi-C library construction and sequencing Construct Hi-C libraries for the samples to be analyzed, and use a high-throughput sequencing platform to obtain Hi-C data. The data obtained are double-end sequencing reads; S2. Data quality control and comparison The obtained Hi-C sequencing data was quality controlled to filter out low-quality reads, and the reads after quality control were aligned to the reference genome using bowtie2 software to obtain alignment result 1, and the unaligned reads were retained; S3. Cutting and realignment of unaligned reads The unaligned reads were cut using the restriction endonuclease sites used in Hi-C library construction, and the remaining reads were re-aligned to the reference genome using bowtie2 software to obtain alignment result 2. Then, alignment result 1 and alignment result 2 were merged using samtools software to generate a new bam file and sorted; S4. Read counts of gene-related regions Gene-related regions were annotated using annotation files. Gene-related regions included the promoter region, downstream enhancer region, gene body region, CDS region, exon region and intron region of the gene. The positions were counted according to the regional coordinates. A self-built python script was used to screen the bam files. According to the gene regulation pattern, one end of the paired reads was ensured to be located in one or more regions of a certain gene, and the other end was located in one or more regions of other genes. The read counts or standardized read counts located in these gene regions were counted to evaluate the strength of the cross-regulation from a certain gene to other genes.

2. The method for quantitatively identifying gene interaction strength according to claim 1, characterized in that: The method further comprises the following steps: S5. Quantitative analysis of interaction strength The interaction strength between different genes is quantified by reads count or standardized according to the quantitative data, and the gene interaction network is drawn according to the statistical results to provide a basis for the analysis of gene regulatory networks.

3. The method for quantitatively identifying gene interaction strength according to claim 1, characterized in that: In step S2, bowtie2 software was used to align the quality-controlled reads to the reference genome. To ensure alignment accuracy, the parameters were set as follows: --very-sensitive -L 30 --score-min L, -0.6, -0.2 --end-to-end --reorder.

4. The method for quantitatively identifying gene interaction strength according to claim 3, characterized in that: In step S3, the remaining reads were aligned to the reference genome using bowtie2 software again, with the parameters set as: --very-sensitive -L 20 --score-min L, -0.6, -0.2 --end-to-end --reorder.

5. The method for quantitatively identifying gene interaction strength according to claim 1, characterized in that: In the step S4, the normalized reads count is obtained by normalizing the reads count.

6. The method for quantitatively identifying gene interaction strength according to claim 1, characterized in that: In step S4, the annotation information is used as the position region to count the read counts of the Hi-C data or the normalized read counts, so as to evaluate the strength of the interactive regulation from a certain gene to other genes.

Citation Information

Patent Citations

  • Identification method and system of gene-regulatory chromatin interaction as well as application

    CN108220394A

  • Key gene for duck meat and egg character differentiation and method for excavating key gene based on whole genome sequencing

    CN119464325A

  • Methods for determining absolute genome-wide copy number variations of complex tumors

    EP2844771A1

  • Bs and rrbs sequencing-based bioinformatics analysis method and device

    WO2013097061A1