Enzyme digestion artificial sequence filtering method, gene sequencing equipment and readable storage medium

By detecting and removing the reverse complementary artificial sequence produced by the enzyme cleavage interruption method, the problem of false positive mutations caused by the enzyme cleavage interruption method is solved, and the accuracy of somatic mutation detection is improved.

CN120613005APending Publication Date: 2025-09-09SHANGHAI BAIZHEN BIOTECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510702439.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

The artificial sequences generated by the enzyme fragmentation method during genome sequencing lead to false positive low-frequency mutations, affecting the accuracy of somatic mutation detection results.

Method used

By obtaining the alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome, the target sequence with reverse complementary artificial sequence characteristics is detected and removed based on the reverse complementary sequence characteristics, thereby filtering out false positive artificial sequences.

Benefits of technology

The accuracy of somatic mutation detection results is improved, and false positive mutations in the enzyme digestion library construction process are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613005A_ABST
    Figure CN120613005A_ABST
Patent Text Reader

Abstract

The invention provides an enzyme digestion artificial sequence filtering method, gene sequencing equipment and a readable storage medium, and relates to the technical field of gene sequencing. After an original comparison file of original enzyme digestion library-building sequencing data relative to a reference genome is obtained, reverse complementary artificial sequence feature detection is performed on all to-be-detected sequences in the original comparison file based on the reference genome according to sequence feature conditions of reverse complementary sequences; according to the method, an original comparison file is obtained, so that a target sequence with reverse complementary artificial sequence features is determined in the original comparison file, and then all the target sequences existing in the original comparison file are subjected to sequence removal to obtain an effective comparison file fitting reality in the enzyme digestion library building process; therefore, false positive artificial sequences caused by an enzyme digestion breaking method are effectively filtered out, false positive mutations in the enzyme digestion library building process are reduced, and the accuracy of a final somatic mutation detection result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gene sequencing technology, and in particular to a method for filtering enzyme-cut artificial sequences, a gene sequencing device, and a readable storage medium. Background Art

[0002] With the continuous development of bioinformatics, next-generation sequencing (NGS) technology is often used to analyze somatic mutations in cancer because it can sequence hundreds of thousands to millions of DNA / RNA molecules simultaneously, enabling in-depth, detailed, and comprehensive analysis of a species' genome and transcriptome. It also meets the high sensitivity, ease of use, and accurate data quality required for mutation detection. In the actual use of second-generation sequencing technology, library construction is an important step in ensuring data quality. Genome fragmentation is the first step in library construction, which improves the synergy of gene cluster replication by shortening sequence read lengths to avoid a decline in sequencing quality.

[0003] At present, the main genome fragmentation methods include ultrasonic shearing and enzymatic shearing. Among them, the ultrasonic shearing method uses the principle of ultrasonic expansion and contraction resonance to shear the genome, achieving an unbiased shearing effect on the genome and producing a large number of DNA fragments of uniform size. However, there are limitations such as high cost of instrument consumables, the need to explore different shearing times for samples of different quality and degradation degree, and excessively long shearing times that can easily lead to nucleic acid damage. The enzymatic shearing method uses nucleases to randomly shear the genome, which can better preserve the integrity of the nucleic acid. It does not require special instruments and the shearing operation is simple, with a high overall shearing efficiency. Therefore, the enzymatic shearing method has been gradually and widely used to realize manual or automated library construction.

[0004] However, it is worth noting that the enzyme cleavage method will produce single-stranded DNA sticky ends during the process of enzyme cleavage of double-stranded DNA. If there are two nucleotide sequences that are reverse complementary to each other on this sticky end, the two sequences will combine due to base complementary pairing. After DNA repair and PCR (Polymerase Chain Reaction) amplification, a DNA fragment different from the template DNA fragment will eventually be produced. This fragment is the artificial fragment produced during the enzyme cleavage process, and its sequence is called an artificial sequence. Although the proportion of these artificial fragments is not high, because their sequence belongs to the real molecular abnormal combination at the DNA fragment level, they are usually identified as mutations when they are inconsistent with the reference genome sequence. In other words, the generation of artificial sequences will introduce false positive low-frequency mutations. Since most somatic mutations are also low-frequency mutations, the artificial sequences produced by the enzyme cleavage method will have a significant impact on the results of somatic mutation detection, and the accuracy of the somatic mutation detection results based on the enzyme cleavage method cannot be guaranteed. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method for filtering artificial sequences caused by enzyme cleavage, a gene sequencing device and a readable storage medium, which can effectively filter out false positive artificial sequences caused by enzyme cleavage interruption method, reduce false positive mutations in the enzyme cleavage library construction process, and improve the accuracy of the final somatic mutation detection results.

[0006] In order to achieve the above objectives, the technical solutions adopted in the embodiments of the present application are as follows:

[0007] In a first aspect, the present application provides a method for filtering an artificial sequence by enzyme digestion, the method comprising:

[0008] Obtain the original alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome;

[0009] Based on the reference genome, reverse complementary artificial sequence feature detection is performed on all the sequences to be tested in the original comparison file to determine the target sequence with the reverse complementary artificial sequence feature;

[0010] All target sequences in the original alignment file are subjected to sequence removal processing to obtain a corresponding valid alignment file.

[0011] In an optional embodiment, the step of obtaining an original alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome includes:

[0012] Performing a sequencing quality assessment on the original enzyme digestion library construction sequencing data, and removing adapter sequences and low-quality sequences from the original enzyme digestion library construction sequencing data based on the sequencing quality assessment results to obtain initial sequencing data;

[0013] Performing gene sequence alignment on the initial sequencing data and the reference genome to obtain a corresponding initial alignment file;

[0014] According to the actual sequence position of each reference gene sequence in the reference genome, each alignment sequence in the initial alignment file is positionally mapped and sorted, and PCR duplicate sequences are removed from each sorted alignment sequence to obtain the original alignment file.

[0015] In an optional embodiment, the step of performing reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine a target sequence having the reverse complementary artificial sequence feature includes:

[0016] For each sequence to be tested, detecting whether there are mismatched bases in the sequence to be tested relative to a target reference gene sequence in the reference genome, wherein the target reference gene sequence is a reference gene sequence mapped to the sequence position of the sequence to be tested in the reference genome;

[0017] When a mismatched base is detected in the test sequence, extracting local test sequences corresponding to two ends of the test sequence, respectively, and generating reverse complementary paired sequences of the two local test sequences, wherein each local test sequence consists of a first preset number of base pairs of the test sequence in a continuously distributed state;

[0018] The two reverse complementary paired sequences associated with the sequence to be tested are respectively compared with the corresponding target reference gene sequence for base consistency, and a base consistency score of each of the two reverse complementary paired sequences relative to the corresponding target reference gene sequence is obtained;

[0019] According to the base consistency scores of the two reverse complementary paired sequences, it is determined whether the sequence to be tested belongs to the target sequence having the characteristics of the reverse complementary artificial sequence.

[0020] In an optional embodiment, the step of determining whether the test sequence belongs to a target sequence having reverse complementary artificial sequence characteristics based on the base identity scores of the two reverse complementary paired sequences includes:

[0021] Detecting whether the base consistency score of each of the two reverse complementary paired sequences is less than a first score threshold;

[0022] When it is detected that the base consistency scores of the two reverse complementary paired sequences are both less than the first score threshold, it is determined that the sequence to be tested does not belong to the target sequence having the reverse complementary artificial sequence feature;

[0023] When it is detected that the base consistency score of at least one reverse complementary paired sequence is greater than or equal to the first score threshold, it is determined that the sequence to be tested belongs to the target sequence with the reverse complementary artificial sequence feature.

[0024] In an optional embodiment, the step of performing reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine a target sequence having the reverse complementary artificial sequence feature further includes:

[0025] When it is detected that the sequence to be tested does not have a mismatched base, detecting whether the sequence to be tested has a soft clip read relative to the reference genome;

[0026] When a Soft Clip read is detected in the sequence to be tested, a local sequencing sequence of the sequence to be tested including the Soft Clip read is extracted, and a reverse complementary paired sequence of the local sequencing sequence is generated;

[0027] Determining a local gene sequence corresponding to a target reference gene sequence in the reference genome, wherein the local gene sequence consists of a second preset number of reference gene base pairs continuously distributed upstream and downstream of the corresponding target reference gene sequence, and the second preset number is greater than the first preset number;

[0028] Performing a base consistency comparison between the reverse complementary paired sequence of the local sequencing sequence and the corresponding local gene sequence to obtain a base consistency score of the reverse complementary paired sequence relative to the corresponding local gene sequence;

[0029] According to the base consistency score of the reverse complementary paired sequence, it is determined whether the sequence to be tested belongs to the target sequence with the characteristics of the reverse complementary artificial sequence.

[0030] In an optional embodiment, the step of extracting the local sequencing sequence of the sequence to be tested including the Soft Clip read segment includes:

[0031] detecting whether the total number of base pairs in the soft clip read segment is greater than a third preset number, wherein the third preset number is less than the first preset number;

[0032] When it is detected that the total number of base pairs of the Soft Clip read is greater than the third preset number, directly using the Soft Clip read as the local sequencing sequence of the sequence to be tested;

[0033] When it is detected that the total number of base pairs of the Soft Clip reads is less than or equal to the third preset number, a first preset number of base pairs including the Soft Clip reads and in a continuously distributed state are extracted from the sequence to be tested to obtain a local sequencing sequence of the sequence to be tested.

[0034] In an optional embodiment, the step of determining whether the test sequence belongs to a target sequence having the characteristics of a reverse complementary artificial sequence based on the base consistency score of the reverse complementary paired sequence includes:

[0035] Detecting whether the base consistency score of the reverse complementary paired sequence is less than a first score threshold;

[0036] When it is detected that the base consistency score of the reverse complementary paired sequence is less than the first score threshold, it is determined that the corresponding sequence to be tested does not belong to the target sequence having the reverse complementary artificial sequence feature;

[0037] When it is detected that the base consistency score of the reverse complementary paired sequence is greater than or equal to the first score threshold, it is determined that the corresponding sequence to be tested belongs to the target sequence having the reverse complementary artificial sequence feature.

[0038] In an optional embodiment, the step of performing reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine a target sequence having the reverse complementary artificial sequence feature further includes:

[0039] When it is detected that there are no mismatched bases in the sequence to be tested and when it is detected that there are no Soft Clip reads in the sequence to be tested, the sequence to be tested is directly regarded as a valid sequence of the non-target sequence.

[0040] In a second aspect, the present application provides a gene sequencing device comprising a processor and a memory, wherein the memory stores a computer program executable by the processor, and the processor can execute the computer program to implement the enzyme-cut artificial sequence filtering method described in any one of the aforementioned embodiments.

[0041] In a third aspect, the present application provides a readable storage medium having a computer program stored thereon. When the computer program is executed by a computer device, the method for filtering artificial sequences by enzyme digestion as described in any one of the aforementioned embodiments is implemented.

[0042] In this case, the beneficial effects of the embodiments of the present application may include the following:

[0043] After obtaining the original alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome, the present application performs reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome according to the sequence feature status of the reverse complementary sequence, so as to determine the target sequence with the reverse complementary artificial sequence feature in the original alignment file. Then, by removing the sequences of all target sequences existing in the original alignment file, an effective alignment file that fits the reality during the enzyme digestion library construction process is obtained, thereby effectively filtering out false positive artificial sequences caused by the enzyme digestion interruption method, reducing false positive mutations in the enzyme digestion library construction process, and improving the accuracy of the final somatic mutation detection results.

[0044] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0046] Figure 1 A schematic diagram of the composition of a gene sequencing device provided in an embodiment of the present application;

[0047] Figure 2 This is a schematic diagram of a process for filtering an artificial sequence by enzyme digestion provided in an embodiment of the present application;

[0048] Figure 3 for Figure 2 A schematic flow chart of the sub-steps included in step S210;

[0049] Figure 4 for Figure 2 One of the flowcharts of the sub-steps included in step S220;

[0050] Figure 5 for Figure 2 The second flowchart of the sub-steps included in step S220.

[0051] Icons: 10-gene sequencing equipment; 11-memory; 12-processor; 13-communication unit. DETAILED DESCRIPTION

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0053] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.

[0054] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0055] In the description of this application, it should be understood that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, or are the orientation or position relationship in which the product of the application is usually placed when in use, or are the orientation or position relationship commonly understood by those skilled in the art. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application.

[0056] It should also be noted that, in the description of this application, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.

[0057] In addition, in the description of the present application, it is understood that relational terms such as the terms "first" and "second" are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also include elements inherent to such process, method, article or equipment. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or equipment comprising the elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.

[0058] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0059] Please refer to Figure 1 , Figure 1Schematic diagram of the composition of a gene sequencing device 10 provided in an embodiment of the present application. In this embodiment of the present application, the gene sequencing device 10 can effectively filter out false-positive artificial sequences caused by the enzyme fragmentation method on the basis of the gene sequencing function of the sequencer for any nucleic acid sample to be tested, thereby reducing false-positive mutations in the corresponding constructed gene library and improving the accuracy of the final somatic mutation detection results. The gene sequencing device 10 can be a computer device that is communicatively connected to the sequencer. The computer device can be, but is not limited to, a personal computer, a laptop computer, a tablet computer, a server, etc.

[0060] In the embodiment of the present application, the gene sequencing device 10 may include a memory 11, a processor 12, and a communication unit 13. The memory 11, the processor 12, and the communication unit 13 are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, the memory 11, the processor 12, and the communication unit 13 may be electrically connected to each other via one or more communication buses or signal lines.

[0061] In this embodiment, the memory 11 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory 11 is used to store a computer program, and the processor 12 can execute the computer program accordingly after receiving an execution instruction. The memory 11 is also used to store the sequence feature status of the reverse complementary sequence, so as to effectively identify whether any nucleic acid sequence has the reverse complementary artificial sequence feature based on the corresponding sequence feature status.

[0062] In this embodiment, the processor 12 can be an integrated circuit chip with signal processing capabilities. The processor 12 can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, etc., which can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0063] In this embodiment, the communication unit 13 is used to establish a communication connection between the gene sequencing device 10 and other electronic devices via a network, and to send and receive data via the network, where the network includes both a wired communication network and a wireless communication network. For example, the gene sequencing device 10 can use the communication unit 13 to obtain raw enzyme digestion and library sequencing data obtained by the sequencer using the enzyme digestion and fragmentation method for any nucleic acid sample to be tested.

[0064] In this embodiment, the gene sequencing device 10 may pre-store a specific computer program related to the enzyme-cut artificial sequence filtering function in the memory 11, and by driving the processor 12 to execute the specific computer program accordingly, effectively filter out false-positive artificial sequences caused by the enzyme-cut interruption method, reduce false-positive mutations in the enzyme-cut library construction process, and improve the accuracy of the final somatic mutation detection results.

[0065] It is understandable that Figure 1 The block diagram shown is only a schematic diagram of the composition of the gene sequencing device 10. The gene sequencing device 10 may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0066] In the present application, in order to ensure that the gene sequencing device 10 can effectively filter out false positive artificial sequences caused by the enzyme cleavage interruption method, the embodiment of the present application provides an enzyme cleavage artificial sequence filtering method to achieve the aforementioned purpose. The enzyme cleavage artificial sequence filtering method provided in the present application is described in detail below.

[0067] Please refer to Figure 2 , Figure 2This is one of the flow charts of the method for filtering an artificial sequence by enzyme cleavage provided in the embodiment of the present application. In the embodiment of the present application, the method for filtering an artificial sequence by enzyme cleavage may include steps S210 to S230.

[0068] Step S210: Obtain an original alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome.

[0069] In this embodiment, the original enzyme digestion library construction sequencing data is the second-generation sequencing data obtained by the nucleic acid sample to be tested based on the enzyme digestion and fragmentation method; the reference genome and the nucleic acid sample to be tested belong to the same species; and the original alignment file is a SAM or BAM file.

[0070] Alternatively, see Figure 3 , Figure 3 yes Figure 2 Schematic diagram of the process of sub-steps included in step S210. In the embodiment of the present application, step S210 may include sub-steps S211 to S213 to ensure that each sequence to be tested in the original alignment file constructed based on the second-generation sequencing data has a high sequence quality, thereby ensuring the reliability of the final somatic mutation detection result.

[0071] Sub-step S211 , performing sequencing quality assessment on the original enzyme digestion library construction sequencing data, and removing adapter sequences and low-quality sequences from the original enzyme digestion library construction sequencing data based on the sequencing quality assessment results to obtain initial sequencing data.

[0072] In this embodiment, the sequencing quality assessment result may include information such as Raw Bases (raw base number), Raw Reads (raw read number), N Bases (the number of bases whose specific base type is not determined (for example, A (Adenine), T (Thymine), G (Guanine) or C (Cytosine))), Q20 / Q30 (sequencing quality score based on base recognition error rate), and GC content (the percentage of the total number of guanine (G) and cytosine (C) in the sequencing data to the total number of bases). The adapter sequence removal operation includes removing the original sequencing sequences in the original enzyme digestion library construction sequencing data that may have adapters at both ends. The low-quality sequence removal operation includes removing the original sequencing sequences in the original enzyme digestion library construction sequencing data whose corresponding N Bases do not meet the preset number condition, the original sequencing sequences whose total read length does not meet the preset read length condition, and the Poly (A) tail at the 3' end of each original sequencing sequence.

[0073] Optionally, in one implementation of this embodiment, after the gene sequencing device 10 determines the initial sequencing data by executing the above sub-step S211, the initial sequencing data can be subjected to sequencing quality assessment to obtain the initial sequencing quality results of the initial sequencing data (including the Clean bases (clean base number), Clean Reads (clean read number), N Bases (the number of bases whose specific base type is not determined (for example, A (Adenine), T (Thymine), G (Guanine) or C (Cytosine))), Q20 / Q30 (sequencing quality score based on base recognition error rate), GC content (the percentage of the total number of guanine (G) and cytosine (C) in the sequencing data to the total number of bases), sequence repetitiveness and average sequence length, etc.), and then, when the overall Q30 value of the initial sequencing data characterized by the corresponding initial sequencing quality result is greater than a preset score threshold (for example, 80), it is determined that the initial sequencing data is substantially high-quality sequencing data.

[0074] Sub-step S212, performing gene sequence alignment on the initial sequencing data and the reference genome to obtain a corresponding initial alignment file.

[0075] In this embodiment, the gene sequencing device 10 can obtain a corresponding initial alignment file by aligning the initial sequencing data back to the reference genome.

[0076] Sub-step S213, according to the actual sequence position of each reference gene sequence in the reference genome, position mapping and sorting are performed on each alignment sequence in the initial alignment file, and PCR duplicate sequences are removed from each aligned sequence after sorting to obtain the original alignment file.

[0077] In this embodiment, all the test sequences in the original comparison file are composed of the remaining comparison sequences after the PCR duplicate sequences are removed from the initial comparison file; the relative ranking position relationship between the various test sequences in the original comparison file is basically consistent with the relative sequence position relationship between the various reference gene sequences in the reference genome.

[0078] Optionally, in one implementation of this embodiment, after the gene sequencing device 10 determines the original alignment file by executing the above-mentioned sub-step S213, it can perform a quality status analysis on the original alignment file to obtain quality status data of the original alignment file (including information such as the alignment rate, uniformity, on-target rate, mismatch rate, effective depth, coverage, insert length and contamination rate between each sequence to be tested in the original alignment file), and then when the corresponding quality status data indicates that the overall contamination rate value of the original alignment file is less than a preset contamination rate threshold (for example, 5%), it can be determined that the original alignment file is essentially a high-quality sequence alignment file.

[0079] Therefore, the present application can ensure that each sequence to be tested in the original comparison file constructed based on the second-generation sequencing offline data has a high sequence quality by executing the above sub-steps S211 to S213, thereby ensuring the reliability of the final somatic mutation detection results.

[0080] Step S220 , performing reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome, and determining a target sequence having the reverse complementary artificial sequence feature.

[0081] In this embodiment, after the gene sequencing device 10 determines a high-quality original alignment file based on the original enzyme digestion library sequencing data and the reference genome, it can perform reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the sequence feature status of the pre-stored reverse complementary sequence and the reference gene sequences included in the reference genome to determine whether each test sequence has the reverse complementary artificial sequence feature, and use the test sequence with the reverse complementary artificial sequence feature as the target sequence that needs to be filtered out.

[0082] Alternatively, see Figure 4 , Figure 4 yes Figure 2 One of the flow charts of the sub-steps included in step S220 of the present application. In the embodiment of the present application, the step S220 may include sub-steps S221 to S224 to achieve a high-precision artificial sequence detection effect for the sequence to be tested containing mismatched bases.

[0083] Sub-step S221 , for each sequence to be tested, detecting whether the sequence to be tested has any mismatched bases relative to the target reference gene sequence in the reference genome.

[0084] In this embodiment, the gene sequencing device 10 can traverse each of the test sequences in the original comparison file in sequence according to the sequence arrangement order of each of the test sequences, and for the traversed test sequence, search for the target reference gene sequence corresponding to the position mapping of the test sequence in all the reference gene sequences included in the reference genome, and then detect whether there are mismatched bases in the test sequence compared with the target reference gene sequence found. In one implementation of this embodiment, it is possible to determine whether there are mismatched bases in the test sequence relative to the corresponding target reference gene sequence by observing the MD (Number of Mismatches / Edit Distance) label of any test sequence in the original enzyme digestion library sequencing data. Wherein, different test sequences each correspond to a different target reference gene sequence, and the target reference gene sequence is the reference gene sequence corresponding to the sequence position mapping of the test sequence at the reference genome.

[0085] In this embodiment, when the gene sequencing device 10 determines that a certain sequence to be tested has mismatched bases relative to the corresponding target reference gene sequence, it will execute sub-step S222 accordingly.

[0086] Sub-step S222: extracting local test sequences corresponding to the two ends of the sequence to be tested, and generating reverse complementary paired sequences of the two local test sequences.

[0087] In this embodiment, when the gene sequencing device 10 determines that any test sequence has mismatched bases, it extracts a plurality of consecutive base pairs at the two sequence ends of the test sequence according to a first preset number (for example, 8) to obtain local test sequences corresponding to the two sequence ends of the test sequence (which are composed of the first preset number of base pairs of the test sequence in a continuously distributed state), and then generates corresponding reverse complementary pairing sequences for the two test sequences extracted from the test sequence.

[0088] In sub-step S223 , the two reverse complementary paired sequences associated with the sequence to be tested are respectively compared with the corresponding target reference gene sequence for base consistency, to obtain a base consistency score of each of the two reverse complementary paired sequences relative to the corresponding target reference gene sequence.

[0089] Sub-step S224 , judging whether the sequence to be tested belongs to the target sequence having the characteristics of the reverse complementary artificial sequence according to the base consistency scores of the two reverse complementary paired sequences.

[0090] In this embodiment, when two reverse complementary paired sequences associated with any test sequence containing mismatched bases are determined, and the base consistency scores of each of the two reverse complementary paired sequences relative to the target reference gene sequence of the test sequence are determined, it can be determined whether the base consistency scores of each of the two reverse complementary paired sequences are less than a first score threshold (for example, 90%). If it is detected that the base consistency scores of each of the two reverse complementary paired sequences are less than the first score threshold, it is determined that the test sequence does not belong to the target sequence having the reverse complementary artificial sequence feature, or if it is detected that the base consistency score of at least one reverse complementary paired sequence is greater than or equal to the first score threshold, it is determined that the test sequence belongs to the target sequence having the reverse complementary artificial sequence feature.

[0091] Therefore, the present application can achieve a high-precision artificial sequence detection effect for the sequence to be tested containing mismatched bases by executing the above sub-steps S221 to S224.

[0092] Alternatively, see Figure 5 , Figure 5 yes Figure 2 The second flow chart of the sub-steps included in step S220 of the present application. Figure 4 Compared with the specific step process of step S220 shown in FIG. Figure 5 The specific step process of step S220 shown can also include sub-steps S225 to S229 to achieve a high-precision artificial sequence detection effect for the test sequence without mismatched bases but with Soft Clip reads, where Soft Clip reads are used to represent sequence reads that cannot be aligned to the reference genome but still exist in the SEQ (Segment SEQuence) field of the original alignment file.

[0093] Sub-step S225 , detecting whether there are soft clip reads in the sequence to be tested relative to the reference genome.

[0094] In this embodiment, upon determining that a test sequence has no mismatches relative to a corresponding target reference gene sequence, the gene sequencing device 10 will correspondingly execute sub-step S225 to detect whether a soft clip read exists in the test sequence without mismatches. In one implementation of this embodiment, the presence of a soft clip read in the test sequence without mismatches can be determined by detecting the presence of an S (Soft) symbol in the CIGAR information corresponding to the test sequence in the original alignment file.

[0095] In this embodiment, when the gene sequencing device 10 determines that a certain sequence to be tested has no mismatched bases relative to the corresponding target reference gene sequence, and that the sequence to be tested has Soft Clip reads relative to the reference genome, it will execute sub-step S226 accordingly.

[0096] Sub-step S226 , extracting the local sequencing sequence including the Soft Clip read of the sequence to be tested, and generating a reverse complementary paired sequence of the local sequencing sequence.

[0097] In this embodiment, for any test sequence without mismatched bases but with Soft Clip reads, it is possible to detect whether the total number of base pairs of the Soft Clip reads is greater than a third preset number (wherein the third preset number is less than the first preset number, and the third preset number may be 6). If it is detected that the total number of base pairs of the Soft Clip reads is greater than the third preset number, the Soft Clip reads are directly used as the local sequencing sequence of the test sequence, or if it is detected that the total number of base pairs of the Soft Clip reads is less than or equal to the third preset number, the first preset number of base pairs including the Soft Clip reads and in a continuously distributed state are extracted from the test sequence to obtain the local sequencing sequence of the test sequence, thereby ensuring that the obtained local sequencing sequence can effectively characterize the Soft Clip sequence of the corresponding test sequence.

[0098] Sub-step S227, determining the local gene sequence corresponding to the target reference gene sequence in the reference genome.

[0099] In this embodiment, for any test sequence with no mismatched bases but with Soft Clip reads, the target reference gene sequence corresponding to the position mapping of the test sequence can be searched among all the reference gene sequences included in the reference genome, and then a local gene sequence including the target reference gene sequence is extracted from the reference genome based on the relative sequence position relationship between all the reference gene sequences in the reference genome, wherein the local gene sequence is composed of a second preset number (for example, 300) of reference gene base pairs corresponding to the target reference gene sequence and its (i.e., the target reference gene sequence) continuously distributed upstream and downstream, and the second preset number is greater than the first preset number.

[0100] Sub-step S228 , performing a base consistency comparison between the reverse complementary paired sequence of the local sequencing sequence and the corresponding local gene sequence to obtain a base consistency score of the reverse complementary paired sequence relative to the corresponding local gene sequence.

[0101] Sub-step S229 , judging whether the sequence to be tested belongs to a target sequence having the characteristics of a reverse complementary artificial sequence according to the base consistency score of the reverse complementary paired sequence.

[0102] In this embodiment, when the reverse complementary paired sequence associated with any test sequence without mismatched bases but with soft clip reads is determined, as well as the base consistency score of the reverse complementary paired sequence relative to the corresponding local gene sequence, it is possible to detect whether the base consistency score of the reverse complementary paired sequence is less than a first score threshold (for example, 90%). If it is detected that the base consistency score of the reverse complementary paired sequence is less than the first score threshold, it is determined that the test sequence does not belong to the target sequence having the reverse complementary artificial sequence feature, or if it is detected that the base consistency score of the reverse complementary paired sequence is greater than or equal to the first score threshold, it is determined that the test sequence belongs to the target sequence having the reverse complementary artificial sequence feature.

[0103] Therefore, the present application can achieve a high-precision artificial sequence detection effect for the sequence to be tested that has no mismatched bases but has soft clip reads by executing the above sub-steps S225 to S229.

[0104] Optionally, in the embodiment of the present application, Figure 4 Compared with the specific step process of step S220 shown in FIG. Figure 5 The specific step flow of step S220 shown may further include sub-step S2210, in which the sequence to be tested without mismatched bases and soft clip reads is used as a valid gene sequence that is not an artificial sequence.

[0105] Sub-step S2210 : directly treating the sequence to be tested as a valid sequence of a non-target sequence.

[0106] In this embodiment, when the gene sequencing device 10 determines that a certain sequence to be tested does not have any mismatched bases relative to the corresponding target reference gene sequence, and that the sequence to be tested does not have any Soft Clip reads relative to the reference genome, it will determine that the sequence to be tested does not have the reverse complementary artificial sequence feature, and the sequence to be tested is not an artificial sequence. At this time, the gene sequencing device 10 will correspondingly execute sub-step S2210.

[0107] Therefore, the present application can use the sequence to be tested without mismatched bases and soft clip reads as a valid gene sequence that is not an artificial sequence by executing the above sub-step S2210.

[0108] Step S230 , performing sequence removal processing on all target sequences in the original alignment file to obtain a corresponding valid alignment file.

[0109] In this embodiment, when the gene sequencing device 10 determines all target sequences (i.e., artificial sequences) with reverse complementary artificial sequence characteristics existing in the original alignment file by executing the above sub-steps S221 to S2210, all the determined target sequences can be removed from the original alignment file, so that the original alignment file after sequence removal is essentially a valid alignment file that conforms to reality during the enzyme digestion library construction process, thereby effectively filtering out false positive artificial sequences caused by the enzyme digestion interruption method, reducing false positive mutations during the enzyme digestion library construction process, and effectively improving the accuracy of the final somatic mutation detection results.

[0110] Therefore, the present application can effectively filter out false-positive artificial sequences caused by enzyme cleavage interruption by executing the above steps S210 to S230 based on the sequence characteristics of the reverse complementary sequence, reduce false-positive mutations in the enzyme cleavage library construction process, and improve the accuracy of the final somatic mutation detection results.

[0111] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0112] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part. If the various functions provided by the present application are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device to perform all or part of the steps of the method described in each embodiment of the present application as a gene sequencing device 10. The aforementioned readable storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0113] The above are merely various embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for filtering an artificial sequence by enzyme digestion, characterized in that: The method comprises: Obtain the original alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome; Based on the reference genome, reverse complementary artificial sequence feature detection is performed on all the sequences to be tested in the original comparison file to determine the target sequence with the reverse complementary artificial sequence feature; All target sequences in the original alignment file are subjected to sequence removal processing to obtain a corresponding valid alignment file.

2. The method according to claim 1, characterized in that The step of obtaining the original alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome includes: Performing a sequencing quality assessment on the original enzyme digestion library construction sequencing data, and removing adapter sequences and low-quality sequences from the original enzyme digestion library construction sequencing data based on the sequencing quality assessment results to obtain initial sequencing data; Performing gene sequence alignment on the initial sequencing data and the reference genome to obtain a corresponding initial alignment file; According to the actual sequence position of each reference gene sequence in the reference genome, each alignment sequence in the initial alignment file is positionally mapped and sorted, and PCR duplicate sequences are removed from each sorted alignment sequence to obtain the original alignment file.

3. The method according to claim 1, characterized in that The step of performing reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine a target sequence having the reverse complementary artificial sequence feature comprises: For each sequence to be tested, detecting whether there are mismatched bases in the sequence to be tested relative to a target reference gene sequence in the reference genome, wherein the target reference gene sequence is a reference gene sequence mapped to the sequence position of the sequence to be tested in the reference genome; When a mismatched base is detected in the test sequence, extracting local test sequences corresponding to two ends of the test sequence, respectively, and generating reverse complementary paired sequences of the two local test sequences, wherein each local test sequence consists of a first preset number of base pairs of the test sequence in a continuously distributed state; The two reverse complementary paired sequences associated with the sequence to be tested are respectively compared with the corresponding target reference gene sequence for base consistency, and a base consistency score of each of the two reverse complementary paired sequences relative to the corresponding target reference gene sequence is obtained; According to the base consistency scores of the two reverse complementary paired sequences, it is determined whether the sequence to be tested belongs to the target sequence having the characteristics of the reverse complementary artificial sequence.

4. The method according to claim 3, characterized in that The step of determining whether the test sequence belongs to a target sequence having the characteristics of a reverse complementary artificial sequence based on the base consistency scores of the two reverse complementary paired sequences comprises: Detecting whether the base consistency score of each of the two reverse complementary paired sequences is less than a first score threshold; When it is detected that the base consistency scores of the two reverse complementary paired sequences are both less than the first score threshold, it is determined that the sequence to be tested does not belong to the target sequence having the reverse complementary artificial sequence feature; When it is detected that the base consistency score of at least one reverse complementary paired sequence is greater than or equal to the first score threshold, it is determined that the sequence to be tested belongs to the target sequence with the reverse complementary artificial sequence feature.

5. The method according to claim 3, characterized in that The step of performing reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine a target sequence having the reverse complementary artificial sequence feature further includes: When it is detected that the sequence to be tested does not have a mismatched base, detecting whether the sequence to be tested has a soft clip read relative to the reference genome; When a Soft Clip read is detected in the sequence to be tested, a local sequencing sequence of the sequence to be tested including the Soft Clip read is extracted, and a reverse complementary paired sequence of the local sequencing sequence is generated; Determining a local gene sequence corresponding to a target reference gene sequence in the reference genome, wherein the local gene sequence consists of a second preset number of reference gene base pairs continuously distributed upstream and downstream of the corresponding target reference gene sequence, and the second preset number is greater than the first preset number; Performing a base consistency comparison between the reverse complementary paired sequence of the local sequencing sequence and the corresponding local gene sequence to obtain a base consistency score of the reverse complementary paired sequence relative to the corresponding local gene sequence; According to the base consistency score of the reverse complementary paired sequence, it is determined whether the sequence to be tested belongs to the target sequence with the characteristics of the reverse complementary artificial sequence.

6. The method according to claim 5, characterized in that The step of extracting the local sequencing sequence of the sequence to be tested including the SoftClip read segment comprises: detecting whether the total number of base pairs in the soft clip read segment is greater than a third preset number, wherein the third preset number is less than the first preset number; When it is detected that the total number of base pairs of the Soft Clip read is greater than the third preset number, directly using the Soft Clip read as the local sequencing sequence of the sequence to be tested; When it is detected that the total number of base pairs of the Soft Clip reads is less than or equal to the third preset number, a first preset number of base pairs including the Soft Clip reads and in a continuously distributed state are extracted from the sequence to be tested to obtain a local sequencing sequence of the sequence to be tested.

7. The method according to claim 5, characterized in that The step of determining whether the sequence to be tested belongs to a target sequence having the characteristics of a reverse complementary artificial sequence based on the base consistency score of the reverse complementary paired sequence comprises: Detecting whether the base consistency score of the reverse complementary paired sequence is less than a first score threshold; When it is detected that the base consistency score of the reverse complementary paired sequence is less than the first score threshold, it is determined that the corresponding sequence to be tested does not belong to the target sequence having the reverse complementary artificial sequence feature; When it is detected that the base consistency score of the reverse complementary paired sequence is greater than or equal to the first score threshold, it is determined that the corresponding sequence to be tested belongs to the target sequence having the reverse complementary artificial sequence feature.

8. The method according to any one of claims 5 to 7, characterized in that: The step of performing reverse complementary artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine a target sequence having the reverse complementary artificial sequence feature further includes: When it is detected that there are no mismatched bases in the sequence to be tested and when it is detected that there are no Soft Clip reads in the sequence to be tested, the sequence to be tested is directly regarded as a valid sequence of the non-target sequence.

9. A gene sequencing device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program executable by the processor, and the processor can execute the computer program to implement the enzyme-cut artificial sequence filtering method according to any one of claims 1 to 8.

10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a computer device, the method for filtering enzyme-cut artificial sequences according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • High-throughput detection method for gene rare mutation

    CN111073961A

  • Restriction-site associated DNA sequencing library construction method, restriction-site associated DNA sequencing data analysis method, detection equipment and storage medium

    CN111524552A

  • Method for accurately detecting mutation in DNA single molecule

    CN116685692A

  • Filtering method for breaking false positive mutation generated by artificial fragments in library building through enzyme digestion method

    CN116895332A

  • Calculating device and method for identifying second-generation sequencing data artificial chimeric read segment of FFPE sample

    CN118471332A