Methods for identifying cleavage sites of nucleic acid nicking enzymes

By performing mNGS sequencing and data processing on samples before and after cleavage with the cleavage enzyme, and analyzing the data using FastP and Bowtie2 software, the problems of long time and high cost in identifying cleavage sites of nucleic acid cleavage enzymes in existing technologies have been solved, and rapid and accurate identification of cleavage sites has been achieved.

CN116741277BActive Publication Date: 2026-05-29GUANGDONG GENERAL HOSPITAL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG GENERAL HOSPITAL
Filing Date
2023-06-12
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies require identifying the cleavage site of nucleic acid cleavage enzymes, which is time-consuming and costly.

Method used

mNGS next-generation sequencing was performed on samples before and after enzyme digestion. Data was cropped and filtered using FASTP software, and compared using Bowtie2 software. The cleavage site was determined by combining Log2 Coverage Ratio values ​​and line graph analysis.

Benefits of technology

This enables rapid and low-cost identification of cleavage sites of nickases, reduces the need for specificity of the nucleic acid sequence to be cleaved, and improves the speed and accuracy of analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0004281482720000011
    Figure HDA0004281482720000011
Patent Text Reader

Abstract

The present application relates to the field of biotechnology, and particularly relates to a method for identifying a cleavage site of a nucleic acid nicking enzyme. The present application discloses a method for identifying a cleavage site of a nicking enzyme, which comprises the following steps: obtaining nucleic acid sequence data before and after enzyme cleavage by using mNGS sequencing, analyzing second-generation sequencing data by using Bowite2 and samtools software, and obtaining the cleavage site of the nicking enzyme according to single-base sequencing depth and analysis. The present application does not need to determine the sequence of the nucleic acid to be cleaved, has short time consumption and low budget, and has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and more specifically to a method for identifying the cleavage sites of nucleic acid cleavage enzymes. Background Technology

[0002] Nucleic acid testing is widely used in food safety, biomedical testing, and environmental monitoring. Nucleic acid sequence-specific isothermal polymerase chain reaction (PCR) amplification technology is a prime example of the rapid development of molecular diagnostics in recent years. Clippases can efficiently replace DNA double-stranded restriction enzymes for rapid isothermal amplification, ushering in a new stage of development and accelerating practical applications. The key to using cleavage enzymes lies in identifying their specific cleavage sites, which is crucial for primer design and kit development. However, current methods for determining cleavage enzyme sites require first identifying the target nucleic acid sequence, which is time-consuming and costly. Summary of the Invention

[0003] In view of this, the technical problem to be solved by the present invention is to provide a method for identifying the cleavage sites of nucleic acid cleavage enzymes.

[0004] This invention provides a method for identifying the cleavage site of a cleavage enzyme, comprising the following steps:

[0005] Step 1: Sequencing the samples before and after cleavage with the cleavage enzyme to obtain sample data 1 before cleavage with the cleavage enzyme and sample data 2 after cleavage with the cleavage enzyme.

[0006] Step 2: Trim and filter the sample data 1 and sample data 2 respectively to obtain quality control data 1 and quality control data 2;

[0007] Step 3: After comparing and sorting quality control data 1 and quality control data 2, the sequencing depth of the base sites is obtained;

[0008] Step 4: Calculate the Log2 Coverage Ratio value based on the sequencing depth of the base sites and then create a line graph. Determine the cleavage site of the nicking enzyme based on the Log2 Coverage Ratio value and the line graph.

[0009] In the identification method described in this invention,

[0010] The Log2 Coverage Ratio is the base 2 logarithm of the ratio of the sequencing depth of the base sites after enzyme digestion to the sequencing depth of the base sites before enzyme digestion.

[0011] The specific formula is as follows:

[0012] Log2 Coverage Ratio = log2(Depth of sample after enzyme digestion / Depth of sample before enzyme digestion)

[0013] The criteria for determining the cleavage site of the cleavage enzyme are as follows: the lowest point of the cleavage with a Log2 Coverage Ratio value less than 0.06 and a significant trough is the cleavage site of the cleavage enzyme.

[0014] In step 1 of the identification method of the present invention, the sample is a double-stranded DNA sample.

[0015] The sequencing was mNGS next-generation sequencing.

[0016] In step 2 of the identification method of the present invention, the cutting parameters are set as follows: remove adapter, set -5 to 20, set -3 to 20.

[0017] In step 2 of the identification method of the present invention, the filtering parameters are set as follows: -q is set to 20, -n is set to 15, and -l is set to 80.

[0018] In the identification method described in this invention, reads that do not match need to be removed after the comparison and before sorting.

[0019] In the identification method of the present invention, the comparison software is bowtie2, the comparison mode is end-to-end mode, and the parameters of the end-to-end mode are set as follows: very-sensitive, -L is set to 30, and -score-min is set to L, -0.6, -0.2.

[0020] In the identification method of the present invention, the nicking enzyme is a restriction endonuclease that creates single-strand gaps inside or near a DNA-specific sequence.

[0021] The identification method described in this invention can accurately locate the cleavage site of the nicking enzyme. Changes in parameters or steps will affect the analysis speed and accuracy of the identification results. In this invention, the first step of removing the adapter reduces the risk that the adapter sequence is difficult to remove completely after subsequent processing steps. Both -3 and -5 are set to 20 to remove bases with quality below the threshold in the reads, resulting in better accuracy of the analysis results compared with other values. Bowtie2 software is used for alignment, and very-sensitive is selected in end-to-end mode. This further improves the data quality and obtains more accurate alignment results compared with sensitive. At the same time, the nicking enzyme has a single cleavage site, and the precise and rigorous quality control parameter settings reduce the amount of data retained and speed up the analysis.

[0022] This invention discloses a method for identifying the action site of nicking enzymes. This method uses mNGS sequencing to obtain nucleic acid sequence data before and after enzyme digestion, and uses Bowite2 and samtools software to analyze the second-generation sequencing data. Based on the single-base sequencing depth and analysis, the action site of the nicking enzyme is obtained. This invention does not require specifying the sequence of the nucleic acid to be digested, has a short time consumption and low budget, and has good application prospects. Attached Figure Description

[0023] Figure 1 Linear diagram of cleavage sites of cleavage enzymes. Detailed Implementation

[0024] This invention provides a method for identifying the cleavage sites of nucleases. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the desired result. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art can clearly modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to implement and apply the technology of this invention.

[0025] The test materials used in this invention are all common commercial products and can be purchased on the market.

[0026] The present invention will be further illustrated below with reference to the embodiments:

[0027] Example 1: Method for identifying the cleavage sites of nucleases

[0028] Nucleic acids before and after enzyme digestion were subjected to mNGS next-generation sequencing. Bioinformatics analysis was performed using the following steps, primarily using FASTP. The main FASTP parameters were set to -5 20 -3 20-q20-n 15-l 80, with other parameters using default values. Reads that did not meet the requirements were removed.

[0029] (1) Adapter filtering: -A disables adapter trimming. By default, the software will cut out the adapter. If -A is set, this function will be disabled. In this project, the adapter removal function is enabled.

[0030] -a specifies an adapter sequence file; this project uses the software's built-in adapter library.

[0031] (2) Global cropping options

[0032] -f,--trim_front1 trims the first few bases of read1; the default value in this project is 0.

[0033] -F,--trim_front2 trims the first few bases of read2; the default value in this project is 0.

[0034] -t,--trim_tail1 trims the number of bases at the end of read1; the default value in this project is 0.

[0035] -T,--trim_tail2 trims how many bases from the end of read2; the default value in this project is 0.

[0036] -b,--max_len1 If read1 is longer than max_len1, a segment will be truncated from the end to make read1 the same length as max_len1. In this project, the default value of 0 is used, which means there is no limit.

[0037] -B,--max_len2 If read2 is longer than max_len2, a segment will be truncated from the end to make read2 the same length as max_len2. In this project, the default value of 0 is used, which means there is no limit.

[0038] (3) PolyG tail trimming.

[0039] -g,--trim_poly_g truncates the end of polyG;

[0040] --poly_g_min_len detects the length of polyG at the end of the read; the default value of 10 is used in this project.

[0041] (4) polyX tail trimming

[0042] -x,--trim_poly_x truncates the 3' end of the polyX string.

[0043] --poly_x_min_len detects the length of polyX at the end of the read; this project uses the default value of 10.

[0044] (5) Trim each read based on its quality value.

[0045] -5,--cut_front moves the window from the 5' end of the read to the end, removing bases in the window whose average quality value is less than the '<' threshold. 20 is used in this project.

[0046] -3,--cut_tail moves the window from the 3' end of the read to the beginning, removing bases in the window whose average quality value is less than the '<' threshold. In this project, 20 is used.

[0047] -r,--cut_right moves the window from the beginning to the end of the read. If the average quality value of a window is less than the threshold, the bases in the window and the part to its right are removed, and the process stops.

[0048] -W,--cut_window_size sliding window filtering, which is similar to calculating kmer, 1 to 1000, this project uses the default of 4 bases;

[0049] -M, in the selected window, the average base quality value, ranging from 1 to 36. This project uses the default Q20. If the average value of this area is lower than 20, it is considered a low-quality area and will be removed.

[0050] -Q controls whether to remove low-quality items. This project uses the default automatic removal, meaning this parameter is not specified.

[0051] -q sets the standard for low quality. This project uses 20, which means that a quality value less than 20 is considered a low quality base, commonly known as Q20.

[0052] -u is the percentage of low-quality bases. It does not mean that a read will be dropped if it contains low-quality bases. Instead, it sets a certain percentage. This project uses the default of 40, which means that 40% of the 150bp reads will be dropped if they contain more than 60 low-quality bases. If a read does not meet the condition, it will be dropped in pairs.

[0053] -n filters reads with too many N bases. If the N base content is greater than n, the read / pair will be discarded. This project uses 15, which means that if more than 10% of N is present in a 150bp read, it will be removed.

[0054] (6) Length filtering option - L

[0055] -l is followed by a length value. Reads shorter than this length will be discarded. This project uses 80, meaning that segments shorter than 80bp will be filtered out.

[0056] (7) Low complexity filtering -y,--low_complexity_filter, uses low complexity filtering. Here, low complexity is defined as the proportion of bases that are different from the next base (base[i]!=base[i+1]). -Y,--complexity_threshold, the threshold for low complexity (0~100), this project uses the default 30;

[0057] (8) Parameters related to the filtering results report

[0058] -j,--json outputs a JSON-formatted report file (string[=fastp.json]).

[0059] The -h, --html option outputs an HTML-formatted report filename, which can be viewed directly in a browser (e.g., string[=fastp.html]).

[0060] -R,--report_title should be quoted with'or",default is"fastp report"(string[=fastp report])

[0061] (9) Threads used in the filtering process

[0062] -w,--thread Number of threads to use

[0063] (10) The filtered data were compared with the genome before and after enzyme digestion using bowtie2. The comparison parameters were --very-sensitive-L 30--score-min L,-0.6,-0.2--end-to-end;

[0064] (11) After removing the misaligned reads using samtools view-F 4, convert them to BAM format and sort them according to the physical location of the genome;

[0065] (12) After sorting, use samtools depth to obtain the depth of each position on the gene;

[0066] (13) Use R software to compare the depth of each position between the two groups: Log2 Coverage Ratio = log2(Depth of the sample after enzyme digestion / Depth of the sample before enzyme digestion); and plot the Log2 Coverage Ratio as a straight line graph; use the straight line graph to determine the location of the double-strand break. In the broken line graph, the segment with Log2 Coverage Ratio less than 0.06 and with obvious troughs is judged as the double-strand break segment, and the break position is judged as the lowest trough.

[0067] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for identifying the cleavage site of a cleavage enzyme, characterized in that, Includes the following steps: Step 1: Sequencing the samples before and after cleavage with the cleavage enzyme to obtain sample data 1 before cleavage with the cleavage enzyme and sample data 2 after cleavage with the cleavage enzyme. Step 2: Trim and filter the sample data 1 and sample data 2 respectively to obtain quality control data 1 and quality control data 2; Step 3: After comparing and sorting quality control data 1 and quality control data 2, the sequencing depth of the base sites is obtained; Step 4: Calculate the Log2 Coverage Ratio value based on the sequencing depth of the base sites and then create a line graph. Determine the cleavage site of the nicking enzyme based on the Log2 Coverage Ratio value and the line graph. The Log2 Coverage Ratio value is: the logarithm of the ratio of the sequencing depth of the base site after enzyme digestion to the sequencing depth of the base site before enzyme digestion, with base 2 as the base. The criteria for determining the cleavage site of the cleavage enzyme are as follows: the lowest point of the breakage with a Log2 Coverage Ratio value less than 0.06 and a clear trough is the cleavage site of the cleavage enzyme. The clipping parameters are set as follows: remove adapter, set -5 to 20, set -3 to 20; The filtering parameters are set as follows: -q is set to 20, -n is set to 15, and -l is set to 80. The comparison mode is end-to-end mode, and the parameters of the end-to-end mode are set as follows: very-sensitive, -L is set to 30, and -score-min is set to L, -0.6, -0.2; After the comparison and before sorting, reads that do not match need to be removed.

2. The identification method according to claim 1, characterized in that, In step 1, the sample is a double-stranded DNA sample.

3. The identification method according to claim 1, characterized in that, In step 1, the sequencing is mNGS next-generation sequencing.

4. The identification method according to claim 1, characterized in that, The nicking enzyme is a restriction endonuclease that creates single-strand gaps inside or near a specific DNA sequence.