Motif-Hi-C: A Fast Method for Evaluating the Quality of Hi-C Data Based on Proximity Ligation Motif Sequences and Its Applications
The Motif-Hi-C method is used to classify and differentiate Hi-C data, which solves the problem of time-consuming processing of existing software in large Hi-C data sets, and achieves rapid and accurate data quality evaluation and analysis efficiency improvement.
Patent Information
- Application Number
- CN202510413063.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing Hi-C data quality evaluation software takes a long time to process large Hi-C data sets and occupies a lot of computing resources, making it difficult to quickly judge data quality.
Motif-Hi-C data quality evaluation method based on adjacent connected Motif sequences is adopted. Motif matching is performed through KMP algorithm and AC automata algorithm, and classified into matched.fastq and unmatched.fastq files. Different types of data are used to calculate the ratio of effective interactions to the total read pair to evaluate the data quality.
It significantly improves the efficiency and accuracy of Hi-C data quality evaluation, and is especially suitable for large-scale Hi-C data analysis, reducing the amount of upstream data analysis, and improving the analysis speed and efficiency.
Smart Images

Figure CN119920309B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of genomic data analysis, and particularly relates to a fast method for evaluating the quality of Hi-C data based on proximity ligation Motif sequences, namely Motif-Hi-C, and its application. Background Art
[0002] The high-throughput chromatin conformation capture (Hi-C) technology is based on the "proximity ligation" of the 3C technology. By combining high-throughput sequencing technology with bioinformatics analysis, it can reveal the three-dimensional structure of the genome on a genome-wide scale. The Hi-C technology uses restriction endonucleases to recognize specific cleavage sites to cut the genome. Subsequently, through biotin filling and proximity ligation, two spatially adjacent but discontinuous DNA sequences are combined into a chimeric fragment (this fragment is Motif). After constructing a complete Hi-C library and sequencing, the original Hi-C data is generated, stored in the fastq format, and used for subsequent analysis and processing.
[0003] The Hi-C technique has strict requirements for experimental operations. Therefore, in Hi-C datasets, a large amount of data noise often occurs due to the principles of Hi-C experiments and the reaction efficiency of reagents, such as random fragments, self-cycling (Self Cycle, where DNA region fragments connect to themselves), and dangling ends (Dangling Ends, where the biotinylated end-complemented DNA sequences fail to ligate successfully to adjacent sequences) (Eagen et al., 2018, Trends in Biochemical Sciences). These noises will seriously affect the accuracy of the analysis results (Wang et al., 2018, Nature Plants). Therefore, it is necessary to process the original fastq files. Currently, the Hi-C data processing pipeline is mainly divided into two key stages: the upstream stage is to extract the effective interaction information data of Hi-C from the fastq files (Touray et al., 2023, Bio Protocol) (upstream analysis); the downstream stage uses the effective interaction information data of Hi-C to analyze different chromatin structures, such as the identification and mapping analysis of chromatin three-dimensional structures (spatial conformations), including chromatin territories, compartments, topologically associating domains (TADs), and chromatin loops (downstream analysis). The main goal of the upstream stage of Hi-C data is to remove noise from the original data and extract effective interaction information. Currently, commonly used Hi-C data quality assessment software, such as Hi-C Pro, HiCUP, Juicer, Hicexplorer, etc., can effectively evaluate Hi-C data. Among them, Hi-C Pro provides a complete upstream analysis pipeline (Servant et al., 2015, Genome Biology), HiCUP can more quickly extract the effective interaction data in the fastq files (Wingett et al., 2015, F1000Research), Juicer can support both upstream and downstream analysis steps (Durand et al., 2016, Cell System), and HiCExplorer has high data quality control capabilities (Wolff et al., 2018, Nucleic Acids Research). Although the currently commonly used Hi-C data quality assessment software can effectively process Hi-C data, in the face of large Hi-C datasets or low server performance, these software have problems such as long processing time and large computational resource consumption in the upstream analysis stage of Hi-C data, resulting in researchers being unable to quickly judge the quality of Hi-C datasets.
[0004] Hi-C technology can capture all intra-chromatin and inter-chromatin spatial interaction information across the entire genome, thus generating extremely large and complex datasets. For large eukaryotic genomes, the number of theoretical interactions grows exponentially with genome size. At the same time, since the original fastq files obtained from Hi-C experiments contain a large amount of data noise, it is crucial for subsequent analysis to efficiently and accurately process these vast Hi-C data to obtain effective interaction information within reasonable time, resource, and storage requirements.
[0005] Therefore, it is necessary to develop a rapid Hi-C data quality assessment method applicable to large-scale Hi-C data. Summary of the Invention
[0006] The objective of the present invention is to provide a rapid Hi-C data quality assessment method Motif-Hi-C based on proximity ligation Motif sequences and its application to solve the above problems.
[0007] According to the first aspect of the present invention, there is provided a Hi-C data quality assessment method Motif-Hi-C based on proximity ligation Motif sequences, the method comprising the following steps:
[0008] S1: Perform Motif matching on the Hi-C data stored in the fastq file: For single-digested Hi-C data, use the KMP algorithm for Motif matching, and for multi-digested Hi-C data, use the Aho-Corasick automaton algorithm for Motif matching;
[0009] S2: Classify the original Hi-C data stored in the fastq file: According to the restriction enzyme cleavage sites and Motif characteristics, classify the original Hi-C data stored in the fastq file into matched.fastq and unmatched.fastq files;
[0010] S3: Denoise the data in the unmatched.fastq file: After performing alignment, filtering, and deduplication steps on the data in the unmatched.fastq file, obtain the denoised data in the unmatched.fastq file and the evaluation results;
[0011] S4: Simulate deduplication of the data in the matched.fastq file: Calculate the PCR amplification deduplication factor based on the deduplication situation of the data in the unmatched.fastq file, and then use the PCR amplification deduplication factor obtained from the data in the unmatched.fastq file to perform simulated PCR amplification deduplication on the evaluation results of the matched.fastq file;
[0012] S5: Hi-C data quality assessment: Combine the data assessment results of the unmatched.fastq file after noise reduction in step S3 with the data assessment results of the matched.fastq file after simulated deduplication in step S4, and calculate the ratio of the number of valid interactions (valid interaction rmdup) to the total number of read pairs (total read pairs). Evaluate the Hi-C data quality through this ratio.
[0013] Thus, after classifying the Hi-C data, this method adopts different strategies for subsequent processing of different types of data, which can greatly improve the efficiency of Hi-C data quality assessment. Moreover, this method is efficient, scientific, and accurate. The application of Motif-Hi-C provides new ideas and solutions for high-throughput sequencing data analysis, especially suitable for large-scale Hi-C data quality analysis, and can greatly improve the analysis efficiency and speed.
[0014] In some embodiments, in step S2 of the method, the method for classifying the original Hi-C data storage fastq file includes the following steps:
[0015] a1: Screen for read IDs containing a specific Motif: Read the fastq file containing Hi-C sequencing reads. For each read in the R1.fastq and R2.fastq files, the program searches for read IDs with a specific Motif in their IDs.
[0016] a2: Output the read IDs that meet the conditions: For reads whose IDs contain the specific Motif, the program outputs their IDs to two new files respectively.
[0017] a3: Merge and remove duplicate IDs: The program merges the two folders in step a2 into a new file and renames it to ensure that all read IDs that meet the conditions are integrated. Subsequently, the merged ID list is deduplicated, and the deduplicated ID list is saved as a new file.
[0018] a4: Classify the fastq files according to the ID list: Using the ID list of the deduplicated file in step a3, classify the original R1.fastq and R2.fastq files. If the ID of a read exists in the deduplicated file in step a3, classify the read as "matched", and for R1.fastq and R2.fastq, generate R1.matched.fastq and R2.matched.fastq files respectively; if the ID of a read does not exist in the deduplicated file in step a3, classify the read as "unmatched", and for R1.fastq and R2.fastq, generate R1.unmatched.fastq and R2.unmatched.fastq files respectively.
[0019] By classifying the original Hi-C data into two major types of data files, namely matched.fastq and unmatched.fastq, only the unmatched.fastq data needs to be denoised subsequently, while the matched.fastq data only needs to be simulated for deduplication, thus completing the quality assessment of Hi-C data. Thereby, the amount of upstream data analysis can be greatly reduced, the data analysis speed and efficiency can be improved, and this method is verified to be accurate and reliable, especially suitable for the quality assessment of large-scale Hi-C data.
[0020] In some embodiments, the method for matching Motif for single-digested Hi-C data in step S1 of the method includes the following steps:
[0021] 1) Construct a partial match table according to the Motif sequence for quickly moving the pointer during the matching process, and initialize a partial match table with the same length as the Motif, with all initial values being 0;
[0022] 2) Starting from the second character of the Motif, calculate the length of the longest common prefix and suffix at each position in turn and record it in the partial match table. During the matching process, use the constructed partial match table to search for the Motif in the reads sequence;
[0023] 3) Initialize two pointers, respectively pointing to the starting positions of the reads sequence and the Motif;
[0024] 4) Move the pointers in the reads sequence one by one, compare the base characters at the corresponding positions, and adjust the position of the Motif according to the partial match table: If the current base character matches successfully, move both pointers one position backward; if the current base character does not match and the value of the partial match table is greater than 0, move the Motif to the right and keep the pointer of the reads sequence unchanged; if the current base character does not match and the value of the partial match table is 0, move both the Motif and the pointer of the reads sequence one position to the right;
[0025] 5) Repeat the above process until the matching is completed or the pointer of the reads sequence moves to the end.
[0026] The KMP algorithm is a classic string matching algorithm suitable for single - pattern matching. For single - enzyme - digested Hi - C data, since only one specific enzyme - cutting site (i.e., one specific Motif) is involved, the KMP algorithm can efficiently identify all matching sites without dealing with complex pattern combinations. For the processing of single - enzyme - digested Hi - C data, the KMP algorithm helps reduce code complexity and potential errors.
[0027] In some embodiments, the method for matching Motifs for multi - enzyme - digested Hi - C data in step S1 using the Aho - Corasick automaton algorithm includes the following steps:
[0028] 1) Construct a Trie tree as the search data structure of the Aho - Corasick automaton;
[0029] 2) Construct fail pointers to jump to the base character with the longest common prefix and suffix when the current base character fails to match for continued matching;
[0030] 3) Scan the reads sequence for matching.
[0031] The Aho - Corasick automaton algorithm is an efficient multi - pattern string matching algorithm. This algorithm can handle multiple patterns simultaneously and output the starting positions and pattern indices of all matches. For multi - enzyme - digested Hi - C data, the Aho - Corasick automaton algorithm can identify all relevant enzyme - cutting sites (i.e., multiple Motifs) at once without performing separate matching operations for each pattern. The Aho - Corasick automaton algorithm is more advantageous in processing multi - enzyme - digested Hi - C data due to its multi - pattern matching ability, efficiency, and scalability.
[0032] In some embodiments, the Bowtie2 alignment software is used for data alignment in step S3. Thus, through the selection of partial Hi-C data for alignment performance testing, it is found that although Bowtie2 takes longer than BWA in constructing the reference genome index, the overall running time of Bowtie2 in sequence alignment is generally shorter than that of BWA. Especially when dealing with large-scale data sets, the speed advantage of Bowtie2 is more obvious. Bowtie2 becomes the preferred sequence alignment tool in the Motif-Hi-C analysis method, which can optimize the data processing flow and improve the overall analysis efficiency.
[0033] According to the second aspect of the present invention, there is provided an application of a Hi-C data quality assessment method Motif-Hi-C based on proximity ligation Motif sequences in the analysis of Hi-C high-throughput sequencing data. Thus, the application of Motif-Hi-C provides new ideas and solutions for high-throughput sequencing data analysis, especially suitable for large-scale Hi-C data quality analysis, which can greatly improve the analysis efficiency and speed.
[0034] According to the third aspect of the present invention, there is provided a method for analyzing Hi-C high-throughput sequencing data, the method comprising the following steps:
[0035] S1: The Hi-C high-throughput sequencing data is preliminarily evaluated by using the Hi-C data quality assessment method Motif-Hi-C based on proximity ligation Motif sequences;
[0036] S2: According to the result after preliminary evaluation in step S1, if the evaluation result allows for subsequent analysis, the matched.fastq file data is subjected to alignment, filtering, and deduplication operations to obtain specific valid interaction reads information;
[0037] S3: The processed matched.fastq file data in step S2 and the unmatched.fastq file data after noise reduction are separately subjected to downstream analysis or merged and then subjected to downstream analysis.
[0038] Thus, after quickly judging the quality of Hi-C data, if it is determined according to the analysis results of the Motif-Hi-C method that this set of data sets can continue downstream analysis, it is not necessary to re-analyze all the data, but only to complete the quality control analysis on the data in the matched.fastq file. The data in the matched.fastq file and the unmatched.fastq file can be separately subjected to downstream analysis, or they can be merged and then subjected to downstream analysis. This not only retains the advantage of quickly evaluating the quality of the data set, but also expands the possibility of in-depth analysis of the data, providing a flexible solution for efficient and accurate Hi-C data processing and analysis.
[0039] According to the fourth aspect of the present invention, there is provided an application of a method for evaluating the quality of Hi-C data, Motif-Hi-C, based on proximity ligation Motif sequences in the evaluation of Hi-C data quality. Thus, by adopting this method, the efficiency of evaluating the quality of Hi-C data can be improved, and the quality of Hi-C data can be evaluated quickly and accurately.
[0040] According to the fifth aspect of the present invention, there is provided a method for quickly analyzing large-scale Hi-C high-throughput sequencing data, and the method includes the following steps:
[0041] S1: Perform Motif matching on the fastq files storing Hi-C data: For single-digested Hi-C data, use the KMP algorithm for Motif matching, and for multi-digested Hi-C data, use the AC automaton algorithm for Motif matching;
[0042] S2: Classify the fastq files storing the original Hi-C data: According to the restriction enzyme cleavage sites and Motif characteristics, classify the fastq files storing the original Hi-C data into matched.fastq and unmatched.fastq files;
[0043] S3: Only perform downstream analysis on the matched.fastq data in step S2 after quality control.
[0044] Thus, by classifying large-scale original data and only performing downstream analysis on the matched.fastq data after quality control, not only can effective analysis results be obtained, but also the volume of data quality control processing and analysis is greatly reduced, and the efficiency of Hi-C data analysis is greatly improved.
[0045] In some ways, the method for classifying the fastq files storing the original Hi-C data in step S2 includes the following steps:
[0046] a1: Screening read IDs containing a specific Motif: Read the fastq file containing Hi-C sequencing reads. For each read in the R1.fastq and R2.fastq files, the program searches for read IDs with a specific Motif in their IDs.
[0047] a2: Outputting eligible read IDs: For reads with a specific Motif in their IDs, the program outputs their IDs to two new files respectively.
[0048] a3: Merging and removing duplicate IDs: The program merges the two folders in step a2 into a new file and renames it to ensure that all eligible read IDs are integrated. Subsequently, the merged ID list is de-duplicated, and the de-duplicated ID list is saved as a new file.
[0049] a4: Classifying fastq files according to the ID list: Using the ID list of the file de-duplicated in step a3, classify the original R1.fastq and R2.fastq files. If the ID of a read exists in the file de-duplicated in step a3, classify the read as "matched", and generate R1.matched.fastq and R2.matched.fastq files for R1.fastq and R2.fastq respectively; if the ID of a read does not exist in the file de-duplicated in step a3, classify the read as "unmatched", and generate R1.unmatched.fastq and R2.unmatched.fastq files for R1.fastq and R2.fastq respectively.
[0050] Advantages of the present invention:
[0051] The present invention discloses a Hi-C data quality assessment method Motif-Hi-C based on proximity ligation Motif sequences. After classifying the Hi-C data, different strategies are adopted for subsequent processing of different types of data, which can greatly improve the efficiency of Hi-C data quality assessment. Moreover, this method is efficient, scientific, and accurate. The application of Motif-Hi-C provides new ideas and solutions for high-throughput Hi-C sequencing data analysis, especially suitable for large-scale Hi-C data quality analysis, and can greatly improve the analysis efficiency and speed. Description of the drawings
[0052] Figure 1 It is a schematic diagram of the formation of a specific Motif;
[0053] Figure 2It is the functional flowchart for quickly judging the quality of Hi-C data by the Motif-Hi-C method;
[0054] Figure 3 It is the graph of the total proportion of data noise in the matched.fastq and unmatched.fastq files;
[0055] Figure 4 It is the statistical graph of the running time of the Motif-Hi-C method and four Hi-C data analysis software (small Hi-C dataset);
[0056] Figure 5 It is the statistical graph of the running time of the Motif-Hi-C method and four Hi-C data analysis software (large Hi-C dataset);
[0057] Figure 6 It is the statistical graph of the analysis results of the Motif-Hi-C method and four Hi-C data analysis software;
[0058] Figure 7 It is the statistical graph of the trend of the analysis results of the Motif-Hi-C method and four Hi-C data analysis software;
[0059] Figure 8 It is the flowchart of the complete upstream analysis of the Motif-Hi-C method (including supplementary process design);
[0060] Figure 9 It is the statistical graph of the analysis results of the matched.fastq and unmatched.fastq files;
[0061] Figure 10 It is the data visualization result graph of the matched.fastq and unmatched.fastq files in the C2C12_4_i_M dataset: Among them, the red area in the graph represents the interaction between chromatin, the yellow square represents TADs (only one square is indicated by a yellow arrow, but it does not mean there is only one), and the black dots represent loops (only one black dot is indicated by a black arrow, but it does not mean there is only one);
[0062] Figure 11 It is the data visualization result graph of the matched.fastq and unmatched.fastq files in the 293T_2_i_M dataset: Among them, the red area in the graph represents the interaction between chromatin, the yellow square represents TADs (only one square is indicated by a yellow arrow, but it does not mean there is only one), and the black dots represent loops (only one black dot is indicated by a black arrow, but it does not mean there is only one);
[0063] Figure 12Visualization results of the data in the matched.fastq and unmatched.fastq files in the PC-3_4_i_M dataset: Among them, the red area in the figure represents the interaction between chromatin, the yellow boxes represent TADs (only one of the boxes is indicated by a yellow arrow, but it does not mean there is only one), and the black dots represent loops (only one of the black dots is indicated by a black arrow, but it does not mean there is only one). Detailed implementation mode
[0064] Example 1 Design and process of the Motif-Hi-C method.
[0065] Based on the analysis process of existing Hi-C data analysis software, most of the time is consumed in the analysis and alignment in the upstream stage. Therefore, how to shorten the time required for analysis and alignment in the upstream stage will be the key to improving the quality control efficiency of Hi-C data.
[0066] When performing Hi-C experiments, restriction enzymes are used to identify and cut specific restriction sites on DNA fragments. Then, through biotin filling and ligation steps, originally non-adjacent DNA fragments are ligated together to form a unique ligation sequence Motif. Since each restriction enzyme has its specific restriction site, the chimeric short sequence (Motif sequence) formed by biotin filling of the digested DNA fragments through proximity ligation is also specific. Therefore, in the obtained DNA fragments with effective interactions, this specific Motif must be included (as Figure 1 shown, taking the HpyCH4IV endonuclease as an example, the restriction site recognized by this enzyme is ACGT, and the formed Motif is ACGCGT. All DNA fragments that are digested by HpyCH4IV and have effective interactions must include this specific Motif).
[0067] Combined with the Hi-C sequencing data type, the present invention proposes a Hi-C data quality assessment method based on the proximity ligation Motif sequence (hereinafter referred to as the "Motif-Hi-C" method). This method classifies the original Hi-C data storage fastq files into matched.fastq (containing the Motif sequence) and unmatched.fastq (not containing the Motif sequence) files according to the Motif. Then, each part of the classified file data is processed separately according to its data characteristics to reduce the amount of data required for alignment analysis in the upstream stage, improve the data analysis efficiency and speed, and achieve the purpose of quickly judging the quality of Hi-C data. Figure 2It is the flow chart of this method, mainly including: matching and classifying Motif for the original fastq file, aligning, filtering, deduplicating the data in the unmatched.fastq file, and merging and counting the analysis results.
[0068] 1.1 Matching Motif for the fastq file storing Hi-C data.
[0069] For single-digested Hi-C data, the KMP algorithm is used for Motif matching, and for multi-digested Hi-C data, the Aho-Corasick automaton algorithm is used for Motif matching.
[0070] (1) Matching strategy for single-digested Hi-C data based on the KMP algorithm: The KMP algorithm is a string matching algorithm used to find the occurrence position of a pattern string in a main text string, and it is an efficient algorithm with linear time complexity. The classification strategy for single-digested Hi-C data based on the KMP algorithm is mainly used to improve the efficiency of finding Motif. This algorithm makes full use of the length of the longest common prefix and suffix before each position of Motif, that is, uses the next array to realize the movement and backtracking of Motif, while the reads sequence does not backtrack, so as to reduce the number of backtracks when the matching fails and achieve the maximum movement amount, thereby improving the matching efficiency. Its time complexity is , where N is the length of the reads sequence and M is the length of Motif. The execution process of this classification strategy is as follows:
[0071] 1) Construct a partial match table based on Motif for quickly moving the pointer during the matching process. And initialize a partial match table with the same length as Motif, and the initial values are all 0.
[0072] 2) Starting from the second character of Motif, calculate the length of the longest common prefix and suffix for each position in turn and record it in the partial match table. During the matching process, use the constructed partial match table to search for Motif in the reads sequence.
[0073] 3) Initialize two pointers, which point to the starting positions of the reads sequence and Motif respectively.
[0074] 4) Move the pointer in the reads sequence in turn, compare the base characters at the corresponding positions, and adjust the position of Motif according to the partial match table. If the current base character matches successfully, both pointers move one position backward. If the current base character does not match and the value of the partial match table is greater than 0, move Motif to the right (skip the already matched part), and keep the pointer of the reads sequence unchanged. If the current base character does not match and the value of the partial match table is 0, move both the Motif and the pointer of the reads sequence one position to the right.
[0075] 5) Repeat the above process until the matching is completed or the pointer of the reads sequence moves to the end.
[0076] (2) Classification strategy for multi-enzyme digestion Hi-C data based on the Aho-Corasick automaton algorithm: The Aho-Corasick automaton algorithm is a multi-pattern string matching algorithm that is widely used in the field of multi-pattern matching, such as finding and replacing prohibited words, searching for keywords, etc. Its time complexity is , where N is the length of the reads sequence, M is the total length of all Motifs, and Z is the number of matching results. The execution process of this classification strategy is as follows:
[0077] 1) Construct a Trie tree as the search data structure of the Aho-Corasick automaton.
[0078] 2) Construct the fail pointer to jump to the base character with the longest common prefix and suffix when the current base character fails to match. Similar to the KMP algorithm, when the Aho-Corasick automaton fails to match the current character during matching, it uses the fail pointer to make a jump.
[0079] 3) Scan the reads sequence for matching.
[0080] 1.2 Classify the fastq files storing the original Hi-C data.
[0081] The fastq files obtained after the second-generation sequencing of the Hi-C library are in units of four lines, representing the information of the sequencing data. The contents of the four lines are as follows: The first line is the ID of the reads, the second line is the sequence information of the reads, the third line is the "+" separator, and the fourth line is the sequence quality of the reads.
[0082] According to the restriction enzyme cleavage sites and Motif characteristics, classify the fastq files storing the Hi-C data after Motif matching into matched.fastq (containing Motif sequences) and unmatched.fastq (not containing Motif sequences) files. The specific classification method is as follows:
[0083] 1) Screen the read IDs containing specific Motifs: First, read the fastq files (i.e., R1.fastq and R2.fastq) containing Hi-C sequencing reads. For each read in the R1.fastq and R2.fastq files, the program will search for the read ID (such as ACGCGT) that contains a specific Motif (matching the motif sequence) in its ID (usually after the @ symbol at the beginning of each line of the read sequence).
[0084] 2) Output read IDs that meet the criteria: For reads whose IDs contain a specific Motif, the program will output their IDs to two new files respectively. Suppose these two files are named R1_ACGCGT_ID.txt and R2_ACGCGT_ID.txt. In this way, R1_ACGCGT_ID.txt contains all the read IDs that meet the criteria in R1.fastq, while R2_ACGCGT_ID.txt contains the corresponding IDs in R2.fastq.
[0085] 3) Merge and remove duplicate IDs: The program will merge the IDs in R1_ACGCGT_ID.txt and R2_ACGCGT_ID.txt into a new file, for example, named ACGCGT_combined_ID.txt, to ensure that all read IDs that meet the criteria are integrated. Subsequently, the merged ID list will be de-duplicated, and the de-duplicated ID list will be saved as a new file, for example, named ACGCGT_unique_ID.txt.
[0086] 4) Classify the fastq files according to the ID list: Using the ID list in unique_ID.txt, classify the original R1.fastq and R2.fastq files. If the ID of a read exists in unique_ID.txt, then classify this read (including its sequence and quality information) as "matched". For R1.fastq and R2.fastq, generate R1.matched.fastq and R2.matched.fastq files respectively. Similarly, if the ID of a read does not exist in unique_ID.txt, then classify this read as "unmatched". For R1.fastq and R2.fastq, generate R1.unmatched.fastq and R2.unmatched.fastq files respectively.
[0087] Subsequently, process the matched.fastq and unmatched.fastq files separately.
[0088] 1.3 Selection of sequence paired-end alignment software.
[0089] In the Hi-C data analysis process, sequence alignment is a crucial and most time-consuming step. Therefore, choosing an efficient alignment software is essential for optimizing the time efficiency of the entire analysis process. Currently, the widely used alignment software mainly includes BWA and Bowtie2, both of which are implemented based on the BWT algorithm. To evaluate the performance of these two software and determine the alignment tool most suitable for the data analysis requirements of the present invention, the present invention conducted a series of performance tests. These tests were carried out in a twenty-thread environment. Specifically, for the alignment mode settings, BWA used the mem mode, while Bowtie2 used the --end-to-end mode for single-end alignment, and the remaining parameter settings followed the default configurations of their respective modes. This performance evaluation covered two main aspects: First, the time efficiency of both in constructing the reference genome index was compared; second, their running times when performing sequence alignment tasks were evaluated. The final test results revealed that in the index construction step, Bowtie2 took longer than BWA, while for most of the data, in terms of sequence alignment, Bowtie2 had a shorter running time than BWA. Overall, however, Bowtie2 performed better than BWA in terms of running speed. Also, through the final alignment rate results, it was found that although BWA was slightly better than Bowtie2, the difference between the two was not significant. Finally, considering the performance of running speed and alignment rate comprehensively, Bowtie2 was ultimately selected as the sequence alignment tool in the Motif-Hi-C analysis method to optimize the data processing flow of Motif-Hi-C and improve the overall analysis efficiency.
[0090] 1.4 Processing strategies for the data in the matched.fastq and unmatched.fastq files.
[0091] A total of 12 Hi-C datasets were selected for analysis in this invention. The relevant datasets cover different cell types and restriction endonucleases commonly used in Hi-C research (all 12 Hi-C datasets were provided by the Institute of Agricultural Genomics, Chinese Academy of Agricultural Sciences). The above datasets were divided into two categories: 5 small Hi-C datasets (C2C12_1_D_S, 293T_1_i_S, PC-3_1_i_S, PC-3_2_i_S, PC-3_3_i_S) and 7 large Hi-C datasets (C2C12_2_D_M, C2C12_3_D_M, C2C12_4_i_M, 293T_2_i_M, PC-3_4_i_M, PC-3_5_i_M, PC-3_6_i_M). The detailed information is shown in Table 1. The relevant Hi-C data has been uploaded to the NCBI database (website: https: / / www.ncbi.nlm.nih.gov / ). The original data can be obtained by entering the project numbers PRJNA1203007 and PRJNA1201070 (where PRJNA1203007 contains mouse Hi-C library data, namely the library data of C2C12_1_D_S, C2C12_2_D_M, C2C12_3_D_M, C2C12_4_i_M; PRJNA1201070 contains human Hi-C library data, namely the library data of 293T_1_i_S, 293T_2_i_M, PC-3_1_i_S, PC-3_2_i_S, PC-3_3_i_S, 293T_2_i_M, PC-3_4_i_M, PC-3_5_i_M, PC-3_6_i_M).
[0092] Table 1 Detailed information of 12 Hi-C datasets
[0093] Dataset Name Cell Type Fragment Length (bp) Ligation System Motif Sequence C2C12_1_D_S C2C12 50-1000 dilution ligation(Mouse_C2C12_DD) 5’-GATCCGT-3’ 5’-ACGGATC-3’ 5’-GATCGATC-3’ 5’-ACGCGT-3’ 293T_1_i_S HEK293T 80-680 in situ ligation(Human_293T_DD) 5’-GATCCGT-3’ 5’-ACGGATC-3’ 5’-GATCGATC-3’ 5’-ACGCGT-3’ PC-3_1_i_S Human_PC-3 150-700 in situ ligation(Human_PC-3_i1_SD) 5’-GATCGATC-3’ PC-3_2_i_S Human_PC-3 150-700 in situ ligation(Human_PC-3_i2_SD) 5’-GATCGATC-3’ PC-3_3_i_S Human_PC-3 150-700 in situ ligation(Human_PC-3_c2_SD) 5’-GATCGATC-3’ C2C12_3_D_M C2C12 50-1000 dilution ligation(Mouse_C2C12_DD) 5’-GATCCGT-3’ 5’-ACGGATC-3’ 5’-GATCGATC-3’ 5’-ACGCGT-3’ C2C12_4_i_M C2C12 50-1000 in situ ligation(Mouse_C2C12_DD) 5’-GATCGATC-3’ C2C12_2_D_M C2C12 50-700 dilution ligation(Mouse_C2C12_SD) 5’-GATCCGT-3’ 5’-ACGGATC-3’ 5’-GATCGATC-3’ 5’-ACGCGT-3’ 293T_2_i_M HEK293T 80-680 in situ ligation(Human_293T_DD) 5’-GATCCGT-3’ 5’-ACGGATC-3’ 5’-GATCGATC-3’ 5’-ACGCGT-3’ PC-3_4_i_M Human_PC-3 50-900 in situ ligation(Human_PC-3_c1_SD) 5’-GATCGATC-3’ PC-3_5_i_M Human_PC-3 50-900 in situ ligation(Human_PC-3_i1_1_SD) 5’-GATCGATC-3’ PC-3_6_i_M Human_PC-3 50-900 in situ ligation(Human_PC-3_i1_2_SD) 5’-GATCGATC-3’
[0094] (1)Processing strategy for unmatched.fastq file data (noise reduction): For the reads in the unmatched.fastq file, considering that it contains a large amount of data noise, use the classic Hi-C Pro (Servant et al., 2015, Genome Biology) Hi-C data processing process that incorporates the Bowtie2 alignment process to perform steps such as alignment, filtering, and deduplication. And calculate the PCR amplification deduplication multiple according to the deduplication situation of the unmatched.fastq file data.
[0095] Processing strategy for the data in the matched.fastq file (simulated deduplication): In the Motif-Hi-C method, the actual alignment, filtering, and deduplication steps are not performed on the data in the matched.fastq file. Instead, the PCR amplification deduplication factor calculated from the data in the unmatched.fastq file is used to simulate PCR amplification deduplication for the evaluation results of the matched.fastq file (specifically: when calculating various evaluation values from the data in the matched.fastq file, the final evaluation value of the matched.fastq is obtained by directly dividing the quick evaluation value by the PCR amplification deduplication factor. The PCR amplification deduplication factor is obtained from the true PCR amplification deduplication factor calculated after performing true PCR deduplication on the data in the unmatched.fastq file). Therefore, in order to verify the accuracy of this processing strategy, first, a statistical analysis was performed on the total noise proportion of the data in the matched.fastq and unmatched.fastq files storing 12 groups of Hi-C datasets, as Figure 3 shown. The results indicate that: compared to the reads in the unmatched.fastq file, the reads in the matched.fastq file of the dataset contain very little noise. Therefore, the reads in the matched.fastq file are highly likely to be valid interaction reads. For example, Hi-CPro and HiCUP were used to analyze the data in the matched.fastq and unmatched.fastq files of the PC-3_1_i_S dataset. The results show that in the reads data of the matched.fastq file, the total noise is 137,476 (pairs), accounting for 8.55% of the total noise in this dataset. In the reads data of the unmatched.fastq file, the total noise significantly increases to 1,470,858 (pairs), accounting for 91.45% of the total noise in this dataset. Therefore, in order to quickly evaluate the quality of Hi-C data and reduce the time consumed by operations such as alignment, filtering, and deduplication, when performing quick evaluation in Motif-Hi-C for the data processing of the matched.fastq file, the noise can not be filtered, and the reads obtained after classification are regarded as the valid interaction count (validinteraction rmdup).
[0096] After completing the above steps, the data processing evaluation results of the matched.fastq and unmatched.fastq files are merged for comprehensive analysis, and the ratio of the number of valid interactions (valid interaction rmdup) to the total number of read pairs (total readpairs) is calculated. This ratio serves as an important indicator for evaluating the overall quality of the Hi-C dataset, enabling researchers to quickly conduct a preliminary assessment of the quality of the Hi-C dataset. Through this method, although the data in the matched.fastq file has not undergone quality control and may slightly affect the accuracy, the method is still within an acceptable error range, and this method significantly saves analysis time. This allows researchers to quickly determine whether the dataset is suitable for further downstream analysis, thereby improving the analysis efficiency while ensuring data quality.
[0097] Example 2 Verification of the feasibility of the Motif-Hi-C method classification strategy.
[0098] Although the research in Example 1 shows that the noise level of the data in the matched.fastq file is extremely low, in order to further verify the feasibility of the Motif-Hi-C method for the fastq file classification strategy, the present invention adopts two evaluation methods. First, considering that the classification strategy of the Motif-Hi-C method does not further subdivide the reads containing Motifs according to the number of Motifs, which may affect the accuracy of the analysis results. Therefore, the present invention evaluates the number of reads containing two or more Motifs to verify the rigor of the classification strategy of this method. Second, the present invention compares the repeatability of the valid interaction reads obtained by adopting this classification strategy with the reads obtained after quality control analysis using a conventional analysis software. Here, the HiCUP tool is used by the present invention. In this way, the rationality and accuracy of this method are evaluated.
[0099] 2.1 Evaluation of the number of reads containing multiple Motifs.
[0100] In the present invention, although according to the principle of Hi-C experiments, it is known that the probability of generating reads containing multiple Motifs in conventional Hi-C experiments is low. However, to ensure the rigor of the experiments, the present invention analyzed three groups of Hi-C datasets to calculate the number of reads containing two or more Motifs in these data. Taking the PC-3_1_i_S dataset as an example, the proportion of reads with different Motif numbers in the total number of reads was explored to judge the impact on the overall data under each classification. The results showed that in the R1.fastq file, the number of reads containing one Motif was 2,649,049, the number of reads containing two Motifs was 1,646, and the number of reads containing more than two Motifs was only 3. Among them, the number of reads containing two or more Motifs accounted for 0.06% of the total number of reads. Similarly, in the R2.fastq file, the number of reads containing one Motif was 2,104,521, the number of reads containing two Motifs was 1,337, and no reads containing more than two Motifs were found. Among them, the number of reads containing two Motifs accounted for 0.06% of the total number of reads.
[0101] On the other hand, considering the case where read1 and read2 exactly cover the sequencing fragment in the type of Hi-C sequencing data, in this case, the Motif is exactly divided equally by read1 and read2, and the Motif cannot be directly detected in read1 or read2. Therefore, the present invention further analyzed the connection situation between the reverse complementary sequences of the reads in R2.fastq and the corresponding read sequences in R1.fastq, and found that the number of reads containing one Motif was 4,504,923, the number of reads containing two Motifs was 130,312, and the number of reads containing more than two Motifs was 294. Among them, the number of reads containing two or more Motifs accounted for 2.81% of the total number of reads. Therefore, it can be concluded that the proportion of reads containing two or more Motifs in the overall reads is extremely low, and some of them may contain invalid reads. This indicates that not classifying this part of the reads separately has little impact on the accuracy of evaluating the quality of Hi-C data.
[0102] 2.2 Repeatability test of the analysis results of Motif-Hi-C and HiCUP.
[0103] To evaluate the accuracy of the valid reads retained during the analysis of the Motif-Hi-C method, the present invention compared the repeatability of the valid reads retained by the Motif-Hi-C and HiCUP methods. Specifically, the present invention compared the IDs of the reads finally retained by the two methods, aiming to evaluate the proportion of the same reads retained by the two methods among the reads after quality control of both, so as to verify the accuracy of Motif-Hi-C. It should be noted that in the processing strategy of the data in the matched.fastq file, the Motif-Hi-C method only performs simulated deduplication based on the data in the unmatched.fastq file and does not actually execute the deduplication step. Therefore, when performing the repeatability detection with the HiCUP software, the compared IDs are the reads IDs that have been filtered for noise but not yet deduplicated. The research results show that in most Hi-C data sets, the reads retained by the Motif-Hi-C method almost cover the reads retained after filtration by the HiCUP method, and the repeatability with the reads in the HiCUP control group is at least above 75%. This indicates that the valid interacting reads retained by the Motif-Hi-C method are mostly true and valid in most cases.
[0104] Example 3: Using Motif-Hi-C to quickly judge the quality of Hi-C data.
[0105] After the feasibility verification, the actual data test is carried out to evaluate the performance of the Motif-Hi-C method in two aspects: quickly judging the quality of Hi-C data and the reliability of its analysis results. By comparing the analysis results of the Motif-Hi-C method with those of four mainstream Hi-C data analysis software, this example reveals the advantages of Motif-Hi-C in terms of analysis speed.
[0106] First, to verify the advantage of Motif-Hi-C in terms of analysis speed, the present invention selected 5 small Hi-C datasets (C2C12_1_D_S, 293T_1_i_S, PC-3_1_i_S, PC-3_2_i_S, PC-3_3_i_S) and 5 large Hi-C datasets (C2C12_2_D_M, C2C12_3_D_M, C2C12_4_i_M, 293T_2_i_M, PC-3_5_i_M) from the above 12 Hi-C datasets according to cell type and the type of restriction endonuclease used. And to ensure the consistency of the experiment, the present invention set the number of task threads for the small Hi-C datasets to two threads, and the number of task threads for the large Hi-C datasets to twenty threads, and calculated the running time of each dataset. Then, the present invention compared the performance of Motif-Hi-C with the four Hi-C data analysis software used in the above research in terms of running time. The results are as Figure 4 , Figure 5 shown. After analyzing 10 Hi-C datasets, the research results show that the Motif-Hi-C method adopted by the present invention is significantly superior to the other four Hi-C data analysis software in terms of running time. For example, for the 293T_1_i_S dataset, the analysis time of Motif-Hi-C is almost the shortest among the four Hi-C data analysis software, about half of that of the Juicer software. And, compared with the Hi-C Pro and HiCUP software with relatively high reliability of analysis results, Motif-Hi-C also significantly reduces the time required for data analysis and quality control in the test of 10 Hi-C datasets. For example, for the PC-3_1_i_S Hi-C dataset, Motif-Hi-C only needs 9,589 seconds to evaluate the data quality, while Hi-C Pro and HiCUP respectively need 26,389 seconds and 17,596 seconds. It is worth noting that this advantage in time efficiency will be more obvious in datasets with a high Motif ratio, such as the 293T_2_i_M, C2C12_2_D_M, C2C12_4_i_M datasets, etc. This proves the unique advantage of the Motif-Hi-C method in the research of quickly judging the quality of Hi-C data, and further emphasizes the application potential of Motif-Hi-C in efficiently processing Hi-C data containing specific Motifs.
[0107] Furthermore, to verify the reliability of the Motif-Hi-C analysis results, the quality of Hi-C data was judged by the ratio of valid interactions to total reads (valid interaction rmdup / total read pairs). Therefore, 12 Hi-C datasets were used to compare the performance of the Motif-Hi-C method and four Hi-C data analysis software in terms of the ratio of valid interactions to total reads (valid interaction rmdup / total read pairs). The results showed that, as observed from Figure 6 , there were no significant differences between the metrics obtained by the rapid quality control analysis using the Motif-Hi-C method and the other four Hi-C data analysis software, and these differences were as expected, which further demonstrated the feasibility of using the Motif-Hi-C method to rapidly judge the quality of Hi-C data. Moreover, the present invention analyzed the changing trends of the Motif-Hi-C method and four Hi-C data analysis software for 12 Hi-C datasets in terms of the ratio of valid interactions to total reads (valid interaction rmdup / total read pairs), and the results were as shown in Figure 7 . The obtained results showed that the changing trend of the Motif-Hi-C analysis results was very close to that of the other four Hi-C data analysis software. This result indicated that when it was necessary to screen out the dataset with the highest quality from multiple groups of data, the Motif-Hi-C method could provide a rapid and relatively accurate selection.
[0108] Example 4 Hi-C high-throughput sequencing data analysis method (supplementary process based on Motif-Hi-C analysis).
[0109] Although the above Motif-Hi-C method for rapidly judging the quality of Hi-C data can effectively achieve a rapid preliminary assessment of the quality of Hi-C datasets, since the alignment step is not performed on the reads data of the matched.fastq file, the exact position information of these reads on the reference genome has not been determined, so it cannot be directly used for downstream analysis. To address this limitation, the present invention proposes a supplementary process: performing alignment, filtering, and deduplication operations on the data of the matched.fastq file to obtain specific valid interaction reads information. In this way, after researchers rapidly judge the quality of Hi-C data, if they determine according to the analysis results of the Motif-Hi-C method that this group of datasets can continue downstream analysis, they do not need to re-analyze all the data, but only need to complete the quality control analysis on the data of the matched.fastq file, as shown in Figure 8As shown below, it specifically includes the following steps: S1: Based on the Motif-Hi-C analysis method, perform a preliminary evaluation of the Hi-C high-throughput sequencing data; S2: According to the results of the preliminary evaluation in step S1, if the evaluation results allow for subsequent analysis, perform alignment, filtering, and deduplication operations on the data in the matched.fastq file to obtain specific valid interaction reads information; S3: Perform downstream analysis separately on the processed matched.fastq file data and the unmatched.fastq file data after noise reduction in step S2, or perform downstream analysis after merging them.
[0110] In addition, to explore the potential association between Motif and different chromatin structures. After classifying the initial fastq files, the present invention separately performs operations such as alignment, filtering, and deduplication on the data in the matched.fastq and unmatched.fastq files to obtain independent analysis results for both parts. Through this method, researchers can perform independent downstream analysis on the matched.fastq and unmatched.fastq data of the same Hi-C dataset to analyze the relationship between the distribution of Motif and chromatin structure. Finally, the Motif-Hi-C method can also use the Samtools tool to merge the results of both parts to provide complete upstream analysis data for further analysis. Such a method not only retains the advantage of quickly evaluating the quality of the dataset but also expands the possibility of in-depth data analysis, providing a flexible solution for efficient and accurate Hi-C data processing and analysis.
[0111] Example 5 Use Motif-Hi-C to analyze the impact of Motif distribution on Hi-C data.
[0112] After performing a rapid quality assessment of the Hi-C dataset using the Motif-Hi-C method, this embodiment further studied the impact of the presence or absence of motifs on Hi-C data under the Motif-Hi-C classification strategy, that is, analyzed the differences in the data of the two parts of the matched.fastq and unmatched.fastq files. On the other hand, considering that large Hi-C datasets contain relatively rich chromatin structure information, which is more conducive to subsequent downstream analysis comparisons. Therefore, the present invention selected 7 large Hi-C datasets (C2C12_2_D_M, C2C12_3_D_M, C2C12_4_i_M, 293T_2_i_M, PC-3_4_i_M, PC-3_5_i_M, PC-3_6_i_M), and successfully obtained the matched.fastq part and the unmatched.fastq part of each Hi-C dataset using Motif-Hi-C. In addition, to conduct in-depth analysis, it is necessary to understand information such as the specific alignment positions of the data on the reference genome and its alignment quality. Therefore, the present invention used the supplementary process of the Motif-Hi-C method (Example 4) to further perform upstream analysis steps such as alignment, filtering, and deduplication on the data of the matched.fastq and unmatched.fastq files respectively.
[0113] Through this series of processes, the present invention generated two quality-controlled BAM files for the matched.fastq and unmatched.fastq parts respectively. Further, the present invention used the HiCUP2Juicer script to convert the two BAM result files of the matched.fastq part and the unmatched.fastq part into files that can be used for downstream analysis.
[0114] 5.1 Quality control analysis and evaluation of the impact of motif distribution on Hi-C data.
[0115] By adopting the Motif-Hi-C analysis method, the present invention performed a comprehensive and detailed quality control analysis and evaluation on the data of the matched.fastq and unmatched.fastq files in the 7 large Hi-C datasets studied previously. For the detailed analysis results, see Table 2-8.
[0116] Table 2 Statistical results of quality control analysis and evaluation of the matched.fastq and unmatched.fastq of the PC-3_5_i_M dataset
[0117] Analysis indicators matched.fastq unmatched.fastq Total_read pairs_processed 150,696,704 175,001,088 Unique_paired_alignments 90,782,504 84,799,630 Mapping_rate(%) 60.2 48.5 Valid_interaction_pairs 89,208,249 71,262,153 Dangling / Internal_end_pairs 5,659 5,762,524 Religation_pairs 499,785 3,394,689 Self Cycle pairs 216,703 191,919 Dumped_pairs 315,115 2,739,162 Valid_interaction_rmdup 20,478,527 16,611,070 Trans_interaction 3,420,380 2,821,445 Cis_interaction 17,058,147 13,789,625 Cis_shortRange 2,482,785 1,976,044 Cis_longRange 14,575,362 11,813,581 Valid_interaction_rmdup / Total_read pairs(%) 13.6 9.5
[0118] Table 3 Statistical results of quality control analysis and evaluation of matched.fastq and unmatched.fastq in PC-3_6_i_M dataset
[0119] Analysis indicators matched.fastq unmatched.fastq Total_read pairs_processed 170,133,672 184,531,187 Unique_paired_alignments 104,717,823 91,679,225 Mapping_rate(%) 61.6 49.7 Valid_interaction_pairs 102,898,627 77,522,777 Dangling / Internal_end_pairs 6,750 6,093,613 Religation_pairs 576,968 3,549,110 Self Cycle pairs 250,712 204,389 Dumped_pairs 366,406 2,812,044 Valid_interaction_rmdup 32,647,838 25,186,139 Trans_interaction 5,456,118 4,261,023 Cis_interaction 27,191,720 20,925,116 Cis_shortRange 3,966,399 3,002,199 Cis_longRange 23,225,321 17,922,917 Valid_interaction_rmdup / Total_read pairs(%) 19.2 13.6
[0120] Table 4 Statistical results of quality control analysis and evaluation of matched.fastq and unmatched.fastq in PC-3_4_i_M dataset
[0121] Analysis indicators matched.fastq unmatched.fastq Total_read pairs_processed 109,219,079 151,511,239 Unique_paired_alignments 73,417,531 67,998,606 Mapping_rate(%) 67.2 44.9 Valid_interaction_pairs 71,780,966 41,754,798 Dangling / Internal_end_pairs 6,927 12,790,722 Religation_pairs 543,246 6,502,823 Self Cycle pairs 178,686 105,642 Dumped_pairs 328,135 4,425,571 Valid_interaction_rmdup 50,060,865 28,979,624 Trans_interaction 7,930,430 5,024,262 Cis_interaction 42,130,435 23,955,362 Cis_shortRange 6,684,073 3,738,807 Cis_longRange 35,446,362 20,216,555 Valid_interaction_rmdup / Total_read pairs(%) 45.8 19.1
[0122] Table 5 Quality control analysis and evaluation results of matched.fastq and unmatched.fastq in 293T_2_i_M dataset
[0123] Analysis metrics matched.fastq unmatched.fastq Total_read pairs_processed 262,067,352 130,196,103 Unique_paired_alignments 191,992,047 96,045,752 Mapping_rate(%) 73.3 73.8 Valid_interaction_pairs 185,856,147 70,049,775 Dangling / Internal_end_pairs 20,425 9,174,466 Religation_pairs 737,638 9,915,890 Self Cycle pairs 143,699 65,878 Dumped_pairs 375,121 862,802 Valid_interaction_rmdup 133,623,495 50,570,505 Trans_interaction 20,526,884 7,653,309 Cis_interaction 113,096,611 42,917,196 Cis_shortRange 15,469,569 5,843,044 Cis_longRange 97,627,042 37,074,152 Valid_interaction_rmdup / Total_read pairs(%) 51.0 38.8
[0124] Table 6 Statistical results of quality control analysis and evaluation of matched.fastq and unmatched.fastq in C2C12_3_D_M dataset
[0125] Analysis metrics matched.fastq unmatched.fastq Total_read pairs_processed 99,443,969 110,724,395 Unique_paired_alignments 57,403,247 52,471,550 Mapping_rate(%) 57.7 47.4 Valid_interaction_pairs 56,262,780 37,551,549 Dangling / Internal_end_pairs 14,561 3,414,703 Religation_pairs 158,418 5,525,878 Self Cycle pairs 142,627 125,974 Dumped_pairs 628,568 627,833 Valid_interaction_rmdup 39,973,160 26,264,916 Trans_interaction 21,591,017 14,825,196 Cis_interaction 18,382,143 11,439,720 Cis_shortRange 729,277 473,954 Cis_longRange 17,652,866 10,965,766 Valid_interaction_rmdup / Total_read pairs(%) 40.2 23.7
[0126] Table 7 Statistical results of quality control analysis and evaluation of matched.fastq and unmatched.fastq in C2C12_2_D_M dataset
[0127] Analysis metrics matched.fastq unmatched.fastq Total_read pairs_processed 334,641,526 156,490,829 Unique_paired_alignments 199,701,363 92,247,230 Mapping_rate(%) 59.7 58.9 Valid_interaction_pairs 195,058,882 48,702,706 Dangling / Internal_end_pairs 33,534 6,517,301 Religation_pairs 504,085 19,420,256 Self Cycle pairs 233,502 84,980 Dumped_pairs 3,415,916 1,907,807 Valid_interaction_rmdup 18,083,463 6,197,428 Trans_interaction 9,852,420 3,297,341 Cis_interaction 8,231,043 2,900,087 Cis_shortRange 302,976 203,072 Cis_longRange 7,928,067 2,697,015 Valid_interaction_rmdup / Total_read pairs(%) 5.4 4.0
[0128] Table 8 Statistical results of quality control analysis and evaluation of matched.fastq and unmatched.fastq in C2C12_4_i_M dataset
[0129] Analysis metrics matched.fastq unmatched.fastq Total_read pairs_processed 92,786,532 59,894,785 Unique_paired_alignments 66,054,238 44,122,260 Mapping_rate(%) 71.2 73.7 Valid_interaction_pairs 64,762,493 41,818,893 Dangling / Internal_end_pairs 5,348 409,842 Religation_pairs 326,079 859,329 Self Cycle pairs 82,559 59,220 Dumped_pairs 462,292 248,084 Valid_interaction_rmdup 35,558,112 23,833,974 Trans_interaction 4,062,338 2,614,296 Cis_interaction 31,495,774 21,219,678 Cis_shortRange 3,556,681 2,408,887 Cis_longRange 27,939,093 18,810,791 Valid_interaction_rmdup / Total_read pairs(%) 23.6 13.6
[0130] The analysis results in Table 2-8 show that the reads in the matched.fastq file contain more valid interactions compared to the reads in the unmatched.fastq file. Specifically, when analyzing the interaction types in detail, it was found that for both transinteraction and cis interaction, the number of interactions in the Motif-containing part of the matched.fastq was significantly higher than that in the Motif-lacking part of the unmatched.fastq. Especially in the cis interaction type, more abundant interactions were detected in the matched.fastq file data. In addition, the index of valid interaction rmdup / total read pairs in the matched.fastq file was also significantly higher than that in the unmatched.fastq part (as Figure 9 shown). This indicates that the Motif-containing part of the matched.fastq has higher data quality than the Motif-lacking part of the unmatched.fastq. On the other hand, regarding the Hi-C data noise, it was observed that the overall noise level in the matched.fastq part was lower. However, in the case of the self cycle pairs type of noise, although its proportion was low in both types of data, the matched.fastq part was slightly higher than the unmatched.fastq part. This phenomenon may be due to the fact that a Motif is generated during the formation of self cycle pairs, and this Motif can also be recognized by the Motif-Hi-C method just like the Motif contained in the valid DNA fragments, which is consistent with the experimental principle.
[0131] By deeply analyzing the quality control results of the matched and unmatched.fastq file data, the association between Motif and data quality and interaction characteristics can be better understood, providing a solid foundation for subsequent chromatin structure analysis and functional research.
[0132] 5.2 Chromatin structure analysis of the influence of Motif distribution on Hi-C data.
[0133] After completing the quality control analysis, in order to explore whether there is a potential correlation between these Motifs and chromatin structures at various levels, downstream chromatin structure analysis was performed on the data of the matched.fastq and unmatched.fastq files. Through detailed analysis, the research observed significant differences between the two parts of the data. The matched.fastq part containing Motifs carried richer chromatin structure information, as shown in Tables 9 and 10. Whether in the detection of loops or TADs, the matched.fastq part significantly exceeded the unmatched.fastq part. Taking the 293T_2_i_M dataset as an example, the unmatched.fastq part detected 915 loops and 758 TADs respectively, while the matched.fastq part detected 3,208 loops and 2,583 TADs respectively. Even in the C2C12_3_D_M and C2C12_2_D_M datasets with poor data quality, it was found that the vast majority of chromatin structure information was concentrated in the matched.fastq part. In addition, by comparing the compartments of the two parts of the data, the research found that the values of the first principal component calculated for the matched.fastq part were generally higher than those calculated for the unmatched.fastq part, which also indicated that the data in the matched.fastq file contained stronger compartment signals.
[0134] Table 9 Statistics of the number of TADs detected in Matched.fastq and unmatched.fastq
[0135] Dataset name (TAD) matched.fastq unmatched.fastq PC-3_5_i_M 67 33 PC-3_6_i_M 247 116 PC-3_4_i_M 666 189 293T_2_i_M 2583 758 C2C12_3_D_M 20 1 C2C12_2_D_M 0 0 C2C12_4_i_M 277 72
[0136] Table 10 Statistics of the number of loops detected in Matched.fastq and unmatched.fastq
[0137] Dataset name (Loop) matched.fastq unmatched.fastq PC-3_5_i_M 409 251 PC-3_6_i_M 1,069 612 PC-3_4_i_M 2,356 960 293T_2_i_M 3,208 915 C2C12_3_D_M 4 1 C2C12_2_D_M 1 0 C2C12_4_i_M 597 243
[0138] After completing the downstream chromatin structure analysis, in order to more intuitively see the difference in the number of TADs and loops contained in the matched.fastq and unmatched.fastq files. The present invention comprehensively considered factors such as cell type and restriction endonuclease type, and took the region from 0 mb to 90 mb of chromosome 2 in the three groups of data of 293T_2_i_M, PC-3_4_i_M, and C2C12_4_i_M as an example for visualization, and the results are as Figures 10 - 12 shown below: Among them, Figure 10 is the visualization result of the C2C12_4_i_M data, Figure 11It is the visualization result of 293T_2_i_M data, Figure 12 It is the visualization result of PC-3_4_i_M data, Figures 10 - 12 In the figure, the red area represents the interaction between chromatin, the yellow square represents TADs, and the black dots represent loops. The darker the red color, the higher the interaction frequency within the chromatin. As can be seen from Figures 10 - 12 it, the matched.fastq part of the three groups of data has a higher interaction frequency than the unmatched.fastq part, and contains more TADs and loops. This further proves that the matched.fastq part has richer chromatin interaction information than the unmatched.fastq part.
[0139] In summary, combined with the results of quality control analysis, it is considered that when analyzing Hi-C data and chromatin structure, it is very important to identify and deeply analyze DNA fragments containing Motif. Moreover, the present invention also believes that increasing the proportion of reads containing Motif can help improve the overall quality of the data set, thereby helping researchers obtain more three-dimensional chromatin structure information. In addition, the research of the present invention also shows that through the data classification processing strategy, it is possible to provide a new perspective for analyzing Hi-C data. Especially in the future, when there is potential to be applied to large-scale sample data, in order to save analysis time and computing resources, it is also feasible to only analyze the DNA interaction information containing Motif (matched.fastq data). If only the matched.fastq data is subjected to quality control and analysis, not only can effective analysis results (chromatin interaction information) be obtained, but also the data volume of data quality control processing and analysis is greatly reduced, which can greatly improve the efficiency of Hi-C data analysis, especially suitable for the rapid analysis of large-scale Hi-C high-throughput sequencing data. When rapidly analyzing large-scale Hi-C high-throughput sequencing data, the method can be as follows: S1: Perform Motif matching on the Hi-C data stored in the fastq file: For single-digested Hi-C data, use the KMP algorithm for Motif matching, and for multi-digested Hi-C data, use the Aho-Corasick automaton algorithm for Motif matching (the specific matching method can be seen in Example 1); S2: Classify the original Hi-C data stored in the fastq file: According to the restriction enzyme cleavage sites and Motif characteristics, classify the original Hi-C data stored in the fastq file into matched.fastq and unmatched.fastq files (the specific classification method can be seen in Example 1); S3: Only perform downstream analysis on the matched.fastq data in step S2 after quality control.
[0140] In summary, the present invention has developed an innovative method named Motif-Hi-C method. This method is mainly used for quickly evaluating the quality of Hi-C data. The core of this method lies in using Motif to effectively classify the original fastq data and then adopting a differential processing strategy according to the Hi-C data characteristics of each part. Thereby significantly saving the alignment time and greatly improving the efficiency of data analysis in the upstream stage. Further, a feasibility analysis of the classification strategy adopted by Motif-Hi-C was carried out to ensure its scientificity and rationality. Then, actual data tests proved that the Motif-Hi-C method has a significant speed advantage compared with four mainstream Hi-C data analysis software. In addition, based on this innovative classification strategy, the present invention first performed upstream analysis on the data of 7 groups of large-scale Hi-C datasets, namely matched.fastq and unmatched.fastq files, to explore the differences in the interaction information contained in the two parts of the data. And, through downstream analysis, it also deeply explored and compared the information of different levels of chromatin structure in these two parts of the data, revealing the influence of Motif distribution on Hi-C data. This innovative classification strategy also provides a new perspective for improving the quality of Hi-C data. Through this series of analyses, the efficiency and innovation of the Motif-Hi-C method in Hi-C data processing and analysis are demonstrated.
[0141] In summary, the present invention selects 5 groups of small Hi-C datasets and 7 groups of large Hi-C datasets for analysis. The above 12 datasets cover different cell types and restriction endonucleases in Hi-C research, ensuring that the Motif-Hi-C method can be used for the analysis of different Hi-C datasets. The present invention compares the difference in running time of Motif-Hi-C and four Hi-C data analysis software in processing 12 Hi-C datasets. The running time results show that, compared with the four Hi-C data analysis software, the Motif-Hi-C method can process Hi-C data more quickly. The ratio of the number of valid pairs after deduplication to the total read pairs is a key indicator for evaluating the quality of Hi-C datasets. By comparing the ratio of the number of valid interactions (valid interaction rmdup) to the total read pairs (totalread pairs) obtained by Motif-Hi-C and four Hi-C data analysis software after processing 12 Hi-C datasets, the results show that there is no significant difference between the indicators obtained by using the Motif-Hi-C method and the other four Hi-C data analysis software, further indicating that it is feasible to use the Motif-Hi-C method to quickly judge the quality of Hi-C data. Moreover, the present invention analyzes the change trend of the Motif-Hi-C method and four Hi-C data analysis software for 12 datasets in terms of the index of valid interaction rmdup / total read pairs. The results show that the change trend of the Motif-Hi-C analysis results is extremely close to that of the other four Hi-C data analysis software. The above results indicate that when performing quality control and usability analysis on Hi-C data, and when it is necessary to screen out the dataset with the highest quality from multiple groups of Hi-C data, the Motif-Hi-C method can provide a quick and accurate choice. The process of quickly judging the quality of Hi-C data by Motif-Hi-C can effectively achieve a preliminary and rapid assessment of the quality of Hi-C datasets. However, since the alignment step is not performed on the reads data of the matched.fastq file, the exact position information of these reads on the reference genome has not been determined, so it cannot be directly used for downstream chromatin structure analysis. When the quality of Hi-C data is preliminarily evaluated using the Motif-Hi-C method, if the evaluation result allows for subsequent analysis, then perform alignment, filtering, and deduplication operations on the data of the matched.fastq file to obtain specific valid interaction reads information, and make comparisons in multiple aspects between the data of the matched.fastq and unmatched.fastq files.Meanwhile, in the present invention, independent downstream analyses are performed on the two parts of data, namely matched.fastq and unmatched.fastq, of the same Hi-C dataset, further analyzing the potential association between Motif and different chromatin structures. The results show that, compared with the reads data without Motif (unmatched.fastq), the reads data containing Motif (matched.fastq) exhibits higher data quality and contains more chromatin structure information. The present invention not only retains the advantage of rapid assessment of the dataset quality but also expands the possibility of in-depth data analysis, providing a flexible solution for efficient and accurate Hi-C data processing and analysis.
Claims
1. The Motif-Hi-C method for evaluating the quality of Hi-C data based on proximity ligation Motif sequences, characterized in that The method includes the following steps: S1: Motif matching for Hi-C data stored in fastq files: For single-digested Hi-C data stored in fastq files, the KMP algorithm is used for Motif matching. For multi-digested Hi-C data stored in fastq files, the Aho-Corasick automaton algorithm is used for Motif matching; S2: Classification of original Hi-C data stored in fastq files: According to the restriction enzyme cleavage sites and Motif characteristics, the original Hi-C data stored in fastq files are classified into matched.fastq and unmatched.fastq files; S3: Noise reduction processing of unmatched.fastq file data: After performing alignment, filtering, and deduplication steps on the unmatched.fastq file data, the noise-reduced unmatched.fastq file data and the evaluation results are obtained; S4: Simulated deduplication of matched.fastq file data: Calculate the PCR amplification deduplication factor based on the deduplication situation of the unmatched.fastq file data, and then use the PCR amplification deduplication factor obtained from the unmatched.fastq file data to perform simulated PCR amplification deduplication on the evaluation results of the matched.fastq file; S5: Hi-C data quality assessment: Combine the evaluation results of the noise-reduced unmatched.fastq file data in step S3 with the evaluation results of the simulated deduplicated matched.fastq file data in step S4, and calculate the ratio of the number of valid interactions to the total number of read pairs. The quality of Hi-C data is evaluated through this ratio.
2. The method according to claim 1, wherein The method for classifying the original Hi-C data stored in fastq files in step S2 includes the following steps: a1: Screening read IDs containing specific Motifs: Read the fastq files containing Hi-C sequencing reads. For each read in the R1.fastq and R2.fastq files, the program searches for read IDs with specific Motifs in their IDs; a2: Outputting eligible read IDs: For reads with specific Motifs in their IDs, the program outputs their IDs to two new files respectively; a3: Merging and removing duplicate IDs: The program merges the two folders in step a2 into a new file and renames it to ensure that all eligible read IDs are integrated. Subsequently, duplicate removal is performed on the merged ID list, and the deduplicated ID list is saved as a new file; a4: Classify the fastq files according to the ID list: Using the ID list of the deduplicated file in step a3, classify the original R1.fastq and R2.fastq files. If the ID of a read exists in the deduplicated file in step a3, classify the read as "matched", and for R1.fastq and R2.fastq, generate R1.matched.fastq and R2.matched.fastq files respectively; if the ID of a read does not exist in the deduplicated file in step a3, classify the read as "unmatched", and for R1.fastq and R2.fastq, generate R1.unmatched.fastq and R2.unmatched.fastq files respectively.
3. The method according to claim 1, wherein The method for matching Motif for the fastq file storing single-digested Hi-C data in step S1 using the KMP algorithm includes the following steps: 1) Construct a partial match table based on the Motif sequence for quickly moving the pointer during the matching process, and initialize a partial match table with the same length as the Motif, with all initial values being 0; 2) Starting from the second character of the Motif, calculate the length of the longest common prefix and suffix for each position in turn and record it in the partial match table. During the matching process, use the constructed partial match table to search for the Motif in the reads sequence; 3) Initialize two pointers, pointing to the starting positions of the reads sequence and the Motif respectively; 4) Move the pointer in the reads sequence in turn, compare the base characters at the corresponding positions, and adjust the position of the Motif according to the partial match table: If the current base character matches successfully, move both pointers one position backward; If the current base character does not match and the value of the partial match table is greater than 0, move the Motif to the right and keep the pointer of the reads sequence unchanged; if the current base character does not match and the value of the partial match table is 0, move both the Motif and the pointer of the reads sequence one position to the right; 5) Repeat the above process until the matching is completed or the pointer of the reads sequence moves to the end.
4. The method according to claim 1, characterized in that The method for matching Motif for the fastq file storing multi-digested Hi-C data in step S1 using the Aho-Corasick automaton algorithm includes the following steps: 1) Construct a Trie tree as the search data structure of the Aho-Corasick automaton; 2) Construct a fail pointer to jump to the base character with the longest common prefix and suffix to continue matching when the current base character fails to match; 3) Scan the reads sequence for matching.
5. The method according to claim 1, characterized in that In step S3, the Bowtie2 comparison software is used for data comparison.
6. Application of the Hi-C data quality assessment method Motif-Hi-C based on the proximity ligation Motif sequence in the analysis of Hi-C high-throughput sequencing data as described in any one of claims 1-5.
7. A Hi-C high-throughput sequencing data analysis method, characterized in that, The method includes the following steps: S1: Use the Motif-Hi-C method for Hi-C data quality assessment based on proximity ligation Motif sequences described in any one of claims 1-5 to preliminarily evaluate the Hi-C high-throughput sequencing data; S2: According to the results after the preliminary evaluation in step S1, if the evaluation results can be used for subsequent analysis, perform alignment, filtering, and deduplication operations on the matched.fastq file data to obtain specific valid interaction reads information; S3: Perform downstream analysis separately on the processed matched.fastq file data and the noise-reduced unmatched.fastq file data in step S2, or perform downstream analysis after merging them.
8. Application of the Motif-Hi-C method for Hi-C data quality assessment based on proximity ligation Motif sequences described in any one of claims 1-5 in Hi-C data quality assessment.
9. A rapid analysis method for large-scale Hi-C high-throughput sequencing data, characterized in that, The method includes the following steps: S1: Perform Motif matching on the Hi-C data storage fastq file: For the single-digested Hi-C data storage fastq file, use the KMP algorithm for Motif matching, and for the multi-digested Hi-C data storage fastq file, use the Aho-Corasick automaton algorithm for Motif matching; S2: Classify the original Hi-C data storage fastq file: According to the restriction enzyme cleavage sites and Motif characteristics, classify the original Hi-C data storage fastq file into matched.fastq and unmatched.fastq files; S3: Only perform downstream analysis on the quality-controlled matched.fastq data in step S2.
10. The method according to claim 9, characterized in that, The method for classifying the original Hi-C data storage fastq file in step S2 includes the following steps: a1: Screen for read IDs containing specific Motifs: Read the fastq file containing Hi-C sequencing reads. For each read in the R1.fastq and R2.fastq files, the program will search for read IDs with specific Motifs in their IDs; a2: Output the qualified read IDs: For reads with specific Motifs in their IDs, the program will output their IDs to two new files respectively; a3: Merge and remove duplicate IDs: The program will merge the two folders in step a2 into a new file and rename it to ensure that all qualified read IDs are integrated. Subsequently, the merged ID list will be deduplicated, and the deduplicated ID list will be saved as a new file; a4: Classify the fastq files according to the ID list: Using the ID list of the deduplicated file in step a3, classify the original R1.fastq and R2.fastq files. If the ID of a read exists in the deduplicated file in step a3, classify this read as "matched", and generate R1.matched.fastq and R2.matched.fastq files for R1.fastq and R2.fastq respectively; if the ID of a read does not exist in the deduplicated file in step a3, classify this read as "unmatched", and generate R1.unmatched.fastq and R2.unmatched.fastq files for R1.fastq and R2.fastq respectively.
Citation Information
Patent Citations
Method for eliminating three-dimensional genomics technology noise by utilizing exonuclease combination
CN108804872A
Gene sequencing data quality evaluation method, system and equipment and storage medium
CN118629512A