Enzymatic sequence filtration methods, gene sequencing devices, and readable storage media

CN120613005BActive Publication Date: 2026-09-25SHANGHAI BAIZHEN BIOTECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510702439.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2026-09-25
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

虽然这些人工片段比例不高,但因其序列属于DNA片段层面的真实分子异常组合,通常会在与参考基因组序列不一致时被认定为突变,也就是说人工序列的产生会引入假阳性低频突变

Benefits of technology

[0043]本申请在获取到原始酶切建库测序数据相对于参考基因组的原始比对文件后,根据反向互补序列的序列特征状况,基于参考基因组对原始比对文件中的所有待测序列分别进行反向互补人工序列特征检测,以便在原始比对文件中确定出具备反向互补人工序列特征的目标序列,而后通过对原始比对文件中存在的所有目标序列进行序列去除,得到在酶切建库过程中贴合现实的有效比对文件,从而完成对因酶切打断法引起的假阳性人工序列的有效滤除,减少酶切建库过程中的假阳性突变,提高最终体细胞突变检测结果的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613005B_ABST
    Figure CN120613005B_ABST
Patent Text Reader

Abstract

The application provides an enzyme digestion artificial sequence filtering method, a gene sequencing device and a readable storage medium, and relates to the technical field of gene sequencing. After an original alignment file of original enzyme digestion library sequencing data relative to a reference genome is acquired, reverse complementary artificial sequence characteristics of all to-be-sequenced sequences in the original alignment file are detected based on the sequence characteristics of the reverse complementary sequences, so that target sequences with the reverse complementary artificial sequence characteristics are determined in the original alignment file, and then all the target sequences existing in the original alignment file are removed, so that an effective alignment file that is consistent with reality in the enzyme digestion library process is obtained, false positive artificial sequences caused by the enzyme digestion breaking method are effectively filtered out, false positive mutations in the enzyme digestion library process are reduced, and the accuracy of the final somatic mutation detection result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene sequencing technology, and more specifically, to an enzyme digestion artificial sequence filtering method, a gene sequencing device, and a readable storage medium. Background Technology

[0002] With the continuous development of bioinformatics, next-generation sequencing (NGS) technology has become increasingly popular. It can simultaneously sequence hundreds of thousands to millions of DNA / RNA molecules, enabling in-depth, detailed, and comprehensive analysis of a species' genome and transcriptome. Furthermore, it provides the high sensitivity, ease of use, and accurate data quality required for mutation detection, and is commonly used to analyze somatic mutations in cancer. In the practical application of NGS technology, library construction is a crucial step in ensuring data quality. Genome fragmentation is the first step in library construction, which improves gene cluster replication coordination by shortening read lengths, thus preventing a decline in sequencing quality.

[0003] Currently, genome fragmentation methods mainly include ultrasonic fragmentation and enzyme digestion fragmentation. Ultrasonic fragmentation utilizes the stretching resonance principle of ultrasound to fragment the genome, achieving unbiased cleavage and producing a large number of uniformly sized DNA fragments. However, it has limitations such as high cost of instruments and consumables, the need to explore different fragmentation times for samples of different quality and degradation levels, and the potential for nucleic acid damage due to excessively long fragmentation times. Enzyme digestion, on the other hand, uses endonucleases to randomly fragment the genome, which better preserves nucleic acid integrity. It does not require specialized instruments, is simple to operate, and has high overall fragmentation efficiency. Therefore, enzyme digestion fragmentation is gradually being widely used to achieve manual or automated library construction.

[0004] However, it's important to note that during the enzyme digestion process of double-stranded DNA, sticky ends are generated on single-stranded DNA. If two nucleotide sequences with opposite complementarity exist at these sticky ends, they will bind due to base pairing. After DNA repair and PCR (Polymerase Chain Reaction) amplification, a DNA fragment different from the template DNA fragment will be generated. This fragment is called an artificial fragment generated during the enzyme digestion process, and its sequence is called an artificial sequence. Although the proportion of these artificial fragments is low, because their sequences represent real molecular abnormal combinations at the DNA fragment level, they are usually identified as mutations when they are inconsistent with the reference genome sequence. In other words, the generation of artificial sequences introduces low-frequency false positive mutations. Since most somatic mutations are also low-frequency mutations, artificial sequences generated by the enzyme digestion method can significantly affect the results of somatic mutation detection, making it impossible to guarantee the accuracy of somatic mutation detection results based on the enzyme digestion method. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method for filtering artificial sequences by enzyme digestion, a gene sequencing device, and a readable storage medium, which can effectively filter out false positive artificial sequences caused by enzyme digestion and fragmentation, reduce false positive mutations in the enzyme digestion library construction process, and improve the accuracy of the final somatic mutation detection results.

[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0007] In a first aspect, this application provides a method for filtering artificial sequences via enzyme digestion, the method comprising:

[0008] Obtain the raw alignment file of the original enzyme digestion, library construction, and sequencing data relative to the reference genome;

[0009] Based on the reference genome, reverse complementary artificial sequence feature detection is performed on all test sequences in the original alignment file to determine the target sequences with reverse complementary artificial sequence features;

[0010] Sequence removal processing is performed on all target sequences present in the original alignment file to obtain the corresponding valid alignment file.

[0011] In an optional implementation, the step of obtaining the original alignment file of the raw enzyme digestion library construction sequencing data relative to the reference genome includes:

[0012] The original enzyme digestion library construction sequencing data were subjected to sequencing quality assessment, and based on the sequencing quality assessment results, adapter sequences and low-quality sequences were removed from the original enzyme digestion library construction sequencing data to obtain initial sequencing data;

[0013] Gene sequence alignment is performed between the initial sequencing data and the reference genome to obtain the corresponding initial alignment file;

[0014] Based on the actual sequence positions of each reference gene sequence in the reference genome, the alignment sequences in the initial alignment file are sorted by position mapping, and PCR repeat sequences are removed from the sorted alignment sequences to obtain the original alignment file.

[0015] In an optional implementation, the step of performing reverse complementarity artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine the target sequence possessing reverse complementarity artificial sequence features includes:

[0016] For each test sequence, it is detected whether there are mismatched bases between the test sequence and the target reference gene sequence in the reference genome, wherein the target reference gene sequence is the reference gene sequence that maps the sequence position of the corresponding test sequence in the reference genome;

[0017] If a mismatched base is detected in the sequence to be tested, the local test sequences corresponding to the two ends of the sequence to be tested are extracted, and the reverse complementary pairing sequences of the two local test sequences are generated respectively. Each local test sequence consists of a first preset number of base pairs of the sequence to be tested that are continuously distributed.

[0018] The two reverse complementary pairs associated with the sequence to be tested are compared with the corresponding target reference gene sequence for base identity, and the base identity scores of the two reverse complementary pairs relative to the corresponding target reference gene sequence are obtained.

[0019] Based on the base consistency scores of the two reverse complementary pairing sequences, it is determined whether the sequence to be tested belongs to the target sequence with the characteristics of reverse complementary artificial sequences.

[0020] In an optional implementation, the step of determining whether the test sequence belongs to the target sequence with the characteristics of an artificially paired reverse complementarity sequence based on the base consistency scores of the two reverse complementary sequences includes:

[0021] Detect whether the base consistency score of each of the two reverse complementary pairing sequences is less than a first score threshold;

[0022] If the base consistency scores of the two reverse complementary pairs are both less than the first score threshold, it is determined that the sequence to be tested does not belong to the target sequence with the characteristics of reverse complementary artificial sequences.

[0023] If the base consistency score of at least one reverse complementary pairing sequence is detected to be greater than or equal to the first score threshold, the sequence to be tested is determined to be a target sequence with reverse complementary artificial sequence characteristics.

[0024] In an optional implementation, the step of performing reverse complementarity artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine the target sequence possessing reverse complementarity artificial sequence features further includes:

[0025] If no mismatched bases are detected in the test sequence, the presence of a soft clip read relative to the reference genome is detected.

[0026] If a soft clip read is detected in the sequence to be tested, a local sequencing sequence including the soft clip read is extracted from the sequence to be tested, and a reverse complementary pairing sequence of the local sequencing sequence is generated.

[0027] In the reference genome, a local gene sequence is determined where the corresponding target reference gene sequence is located. The local gene sequence is composed of the corresponding target reference gene sequence and a second preset number of reference gene base pairs that are continuously distributed upstream and downstream of it. The second preset number is greater than the first preset number.

[0028] The reverse complementary pair sequence of the local sequencing sequence is compared with the corresponding local gene sequence for base consistency to obtain the base consistency score of the reverse complementary pair sequence relative to the corresponding local gene sequence.

[0029] Based on the base consistency score of the reverse complementary pairing sequence, it is determined whether the sequence to be tested belongs to the target sequence with the characteristics of reverse complementary artificial sequences.

[0030] In an optional implementation, the step of extracting the local sequencing sequence including the Soft Clip read from the sequence to be tested includes:

[0031] Detect whether the total number of base pairs in the Soft Clip read segment is greater than a third preset number, wherein the third preset number is less than the first preset number;

[0032] If the total number of base pairs in the Soft Clip read is greater than the third preset number, the Soft Clip read is directly used as the local sequencing sequence of the sequence to be tested.

[0033] If the total number of base pairs in the Soft Clip read is less than or equal to the third preset number, a first preset number of base pairs including the Soft Clip read and in a continuous distribution are extracted from the sequence to be tested to obtain the local sequencing sequence of the sequence to be tested.

[0034] In an optional implementation, the step of determining whether the test sequence belongs to a target sequence with reverse complementary artificial sequence characteristics based on the base consistency score of the reverse complementary pairing sequence includes:

[0035] The base consistency score of the reverse complementary pairing sequence is checked to see if it is less than the first score threshold.

[0036] If the base consistency score of the detected reverse complementary pairing sequence is less than the first score threshold, it is determined that the corresponding test sequence does not belong to the target sequence with the characteristics of reverse complementary artificial sequence.

[0037] If the base consistency score of the reverse complementary pairing sequence is detected to be greater than or equal to the first score threshold, the corresponding test sequence is determined to be a target sequence with reverse complementary artificial sequence characteristics.

[0038] In an optional implementation, the step of performing reverse complementarity artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine the target sequence possessing reverse complementarity artificial sequence features further includes:

[0039] If no mismatched bases are detected in the test sequence and no soft clip reads are detected in the test sequence, the test sequence is directly regarded as a valid sequence that is not the target sequence.

[0040] Secondly, this application provides a gene sequencing device, including a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the enzyme digestion artificial sequence filtering method described in any of the foregoing embodiments.

[0041] Thirdly, this application provides a readable storage medium storing a computer program thereon, which, when executed by a computer device, implements the enzyme digestion artificial sequence filtering method described in any of the foregoing embodiments.

[0042] In this case, the beneficial effects of the embodiments of this application may include the following:

[0043] After obtaining the original alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome, this application performs reverse complementation artificial sequence feature detection on all test sequences in the original alignment file based on the sequence characteristics of the reverse complementation sequence and the reference genome. This is to identify target sequences with reverse complementation artificial sequence features in the original alignment file. Then, by removing all target sequences in the original alignment file, an effective alignment file that closely matches reality is obtained during the enzyme digestion library construction process. This effectively filters out false positive artificial sequences caused by enzyme digestion and reduces false positive mutations during the enzyme digestion library construction process, thereby improving the accuracy of the final somatic mutation detection results.

[0044] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0045] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a schematic diagram of the composition of the gene sequencing equipment provided in the embodiments of this application;

[0047] Figure 2 This is one of the flowcharts illustrating the enzyme digestion and artificial sequence filtering method provided in the embodiments of this application;

[0048] Figure 3 for Figure 2 A flowchart illustrating the sub-steps included in step S210;

[0049] Figure 4 for Figure 2 One of the flowcharts for the sub-steps included in step S220;

[0050] Figure 5 for Figure 2 The second flowchart illustrates the sub-steps included in step S220.

[0051] Icons: 10-Gene sequencing equipment; 11-Memory; 12-Processor; 13-Communication unit. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0053] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0054] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0055] In the description of this application, it should be understood that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are used only for the convenience of describing this application and simplifying the description, and are not intended to indicate or imply that the equipment or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0056] In the description of this application, it should also be noted that, unless otherwise expressly specified and limited, the terms "set up," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0057] Furthermore, it is understood in the description of this application that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Those skilled in the art will understand the specific meaning of the above terms in this application based on the specific circumstances.

[0058] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0059] Please refer to Figure 1 , Figure 1This is a schematic diagram of the composition of the gene sequencing device 10 provided in this application embodiment. In this application embodiment, the gene sequencing device 10, based on the gene sequencing function achieved by the sequencer using enzyme digestion to break down any nucleic acid sample to be tested, can effectively filter out false positive artificial sequences caused by enzyme digestion, thereby reducing false positive mutations in the corresponding constructed gene library and improving the accuracy of the final somatic mutation detection results. The gene sequencing device 10 can be a computer device communicatively connected to the sequencer, and the computer device can be, but is not limited to, a personal computer, laptop, tablet computer, server, etc.

[0060] In this embodiment, the gene sequencing device 10 may include a memory 11, a processor 12, and a communication unit 13. The memory 11, the processor 12, and the communication unit 13 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected via one or more communication buses or signal lines.

[0061] In this embodiment, the memory 11 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The memory 11 is used to store computer programs, and the processor 12 can execute the computer programs accordingly after receiving execution instructions. Simultaneously, the memory 11 is also used to store the sequence characteristic status of inverse complementary sequences, so as to effectively identify whether any nucleic acid sequence possesses the characteristics of an artificial inverse complementary sequence based on the corresponding sequence characteristic status.

[0062] In this embodiment, the processor 12 can be an integrated circuit chip with signal processing capabilities. The processor 12 can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0063] In this embodiment, the communication unit 13 is used to establish a communication connection between the gene sequencing device 10 and other electronic devices via a network, and to send and receive data via the network, wherein the network includes wired communication networks and wireless communication networks. For example, the gene sequencing device 10 can obtain the raw enzyme digestion library construction sequencing data obtained by the sequencer using the enzyme digestion method for any nucleic acid sample to be tested through the communication unit 13.

[0064] In this embodiment, the gene sequencing device 10 may pre-store a specific computer program related to the enzyme digestion artificial sequence filtering function in the memory 11, and by driving the processor 12 to execute the specific computer program, the false positive artificial sequences caused by the enzyme digestion method can be effectively filtered out, thereby reducing false positive mutations in the enzyme digestion library construction process and improving the accuracy of the final somatic mutation detection results.

[0065] Understandable Figure 1 The block diagram shown is only a schematic diagram of one configuration of the gene sequencing device 10. The gene sequencing device 10 may also include components such as... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.

[0066] In this application, to ensure that the gene sequencing device 10 can effectively filter out false positive artificial sequences caused by enzyme digestion, this application provides an enzyme digestion artificial sequence filtering method to achieve the aforementioned objective. The enzyme digestion artificial sequence filtering method provided in this application will be described in detail below.

[0067] Please refer to Figure 2 , Figure 2This is one of the flowcharts illustrating the enzyme digestion artificial sequence filtering method provided in this application embodiment. In this application embodiment, the enzyme digestion artificial sequence filtering method may include steps S210 to S230.

[0068] Step S210: Obtain the original alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome.

[0069] In this embodiment, the original enzyme digestion library construction and sequencing data is the next-generation sequencing data obtained from the nucleic acid sample to be tested based on the enzyme digestion method; the reference genome and the nucleic acid sample to be tested belong to the same species; the original alignment file is a SAM or BAM file.

[0070] Alternatively, please refer to Figure 3 , Figure 3 yes Figure 2 The flowchart of step S210 includes the sub-steps. In the embodiments of this application, step S210 may include sub-steps S211 to S213 to ensure that each test sequence in the original alignment file constructed based on the next-generation sequencing data has high sequence quality, and to ensure the reliability of the final somatic mutation detection results.

[0071] Sub-step S211 involves performing sequencing quality assessment on the original enzyme digestion library construction sequencing data, and based on the sequencing quality assessment results, removing adapter sequences and low-quality sequences from the original enzyme digestion library construction sequencing data to obtain initial sequencing data.

[0072] In this embodiment, the sequencing quality assessment results may include information such as Raw Bases (number of raw bases), Raw Reads (number of raw reads), N Bases (number of bases with undetermined specific base types (e.g., A (adenine), T (thymine), G (guanine), or C (cytosine))), Q20 / Q30 (sequencing quality score based on base identification error rate), and GC content (percentage of the total number of guanine (G) and cytosine (C) bases in the sequencing data) in the corresponding raw enzyme digestion library sequencing data; adapter sequence removal operation includes removing raw sequencing sequences in the raw enzyme digestion library sequencing data whose ends may be adapters; low-quality sequence removal operation includes removing raw sequencing sequences in the raw enzyme digestion library sequencing data whose corresponding N Bases do not meet the preset number condition, raw sequencing sequences whose total read length does not meet the preset read length condition, and Poly(A) tails present at the 3' end of each raw sequencing sequence.

[0073] Optionally, in one embodiment of this example, after the gene sequencing device 10 determines the initial sequencing data by executing the above sub-step S211, it can perform sequencing quality assessment on the initial sequencing data to obtain the initial sequencing quality result of the initial sequencing data (including the Clean bases, Clean Reads, and N of each initial sequencing sequence in the initial sequencing data). Bases (number of bases whose specific base type is not determined (e.g., A (adenine), T (thymine), G (guanine), or C (cytosine))), Q20 / Q30 (sequencing quality score based on base identification error rate), GC content (percentage of guanine (G) and cytosine (C) in the sequencing data), sequence repetition, and average sequence length, etc.), and then when the overall Q30 value of the initial sequencing data, which corresponds to the initial sequencing quality results, is greater than a preset score threshold (e.g., 80), it can be determined that the initial sequencing data is essentially high-quality sequencing data.

[0074] Sub-step S212 involves aligning the initial sequencing data with the reference genome to obtain the corresponding initial alignment file.

[0075] In this embodiment, the gene sequencing device 10 can obtain the corresponding initial alignment file by pasting the initial sequencing data back onto the reference genome.

[0076] Sub-step S213: Based on the actual sequence positions of each reference gene sequence in the reference genome, position mapping and sorting are performed on each alignment sequence in the initial alignment file, and PCR repetitive sequences are removed from each sorted alignment sequence to obtain the original alignment file.

[0077] In this embodiment, all the test sequences in the original alignment file are the remaining alignment sequences after PCR repetitive sequence removal from the initial alignment file; the relative ordering position relationship between the test sequences in the original alignment file is basically consistent with the relative sequence position relationship between the reference gene sequences in the reference genome.

[0078] Optionally, in one embodiment of this example, after the gene sequencing device 10 determines the original alignment file by executing the above sub-step S213, it can perform quality status analysis on the original alignment file to obtain quality status data of the original alignment file (including information such as alignment rate, uniformity, target rate, mismatch rate, effective depth, coverage, insert length, and contamination rate among the various sequences to be tested in the original alignment file). Then, when the overall contamination rate value of the original alignment file, which is characterized by the corresponding quality status data, is less than a preset contamination rate threshold (e.g., 5%), it can be determined that the original alignment file is essentially a high-quality sequence alignment file.

[0079] Therefore, by executing the above sub-steps S211 to S213, this application can ensure that each sequence to be tested in the original alignment file constructed based on the second-generation sequencing data has high sequence quality, thus ensuring the reliability of the final somatic mutation detection results.

[0080] Step S220: Based on the reference genome, perform reverse complementarity artificial sequence feature detection on all test sequences in the original alignment file to determine the target sequence with reverse complementarity artificial sequence features.

[0081] In this embodiment, after the gene sequencing device 10 determines a high-quality original alignment file based on the original enzyme digestion library construction sequencing data and the reference genome, it can perform reverse complement artificial sequence feature detection on all test sequences in the original alignment file according to the sequence feature status of the pre-stored reverse complement sequences and based on the reference gene sequences included in the reference genome, to determine whether each test sequence has reverse complement artificial sequence features, and the test sequences with reverse complement artificial sequence features are selected as target sequences to be filtered out.

[0082] Alternatively, please refer to Figure 4 , Figure 4 yes Figure 2 One of the flowcharts for step S220 is shown below. In this embodiment, step S220 may include sub-steps S221 to S224 to achieve high-precision artificial sequence detection for sequences containing mismatched bases.

[0083] Sub-step S221: For each test sequence, detect whether there are mismatched bases in the test sequence relative to the target reference gene sequence in the reference genome.

[0084] In this embodiment, the gene sequencing device 10 can sequentially traverse each of the test sequences according to the sequence arrangement order of all test sequences in the original alignment file. For each traversed test sequence, it searches for a target reference gene sequence corresponding to the position mapping of the test sequence in all reference gene sequences included in the reference genome. Then, it detects whether there are mismatched bases between the test sequence and the found target reference gene sequence. In one embodiment, the presence of mismatched bases between the test sequence and the corresponding target reference gene sequence can be determined by observing the MD (Number of Mismatches / edit distance) tags of any test sequence in the original enzyme digestion library sequencing data. Different test sequences correspond to different target reference gene sequences, which are the reference gene sequences whose sequence positions at the corresponding test sequence are mapped in the reference genome.

[0085] In this embodiment, when the gene sequencing device 10 determines that a certain sequence to be tested has a mismatched base relative to the corresponding target reference gene sequence, it will execute the corresponding sub-step S222.

[0086] Sub-step S222: Extract the local test sequences corresponding to the ends of the two sequences of the sequence to be tested, and generate the reverse complementary pairing sequences of the two local test sequences respectively.

[0087] In this embodiment, when the gene sequencing device 10 determines any test sequence with mismatched bases, it extracts a number of consecutive base pairs from the two ends of the test sequence according to a first preset number (e.g., 8) to obtain local test sequences corresponding to the two ends of the test sequence (which are composed of a first preset number of base pairs in a continuous distribution of the test sequence). Then, it generates corresponding reverse complementary pairing sequences for the two test sequences extracted from the test sequence.

[0088] Sub-step S223 involves performing base identity comparisons between the two reverse complementary pairs associated with the sequence to be tested and their corresponding target reference gene sequences, thereby obtaining the base identity scores of each of the two reverse complementary pairs relative to their respective target reference gene sequences.

[0089] Sub-step S224: Based on the base consistency scores of the two reverse complementary pairing sequences, determine whether the sequence to be tested belongs to the target sequence with the characteristics of reverse complementary artificial sequences.

[0090] In this embodiment, when two inverse complementary pairs are identified as associated with any test sequence containing mismatched bases, and the base consistency scores of each of these two inverse complementary pairs relative to the target reference gene sequence of the test sequence are determined, the test sequence is determined not to belong to the target sequence with inverse complementary artificial sequence characteristics if the base consistency scores of the two inverse complementary pairs are less than a first score threshold (e.g., 90%), and if the base consistency scores of the two inverse complementary pairs are both less than the first score threshold, the test sequence is determined to belong to the target sequence with inverse complementary artificial sequence characteristics if the base consistency scores of at least one inverse complementary pair are greater than or equal to the first score threshold.

[0091] Therefore, by executing the above sub-steps S221 to S224, this application can achieve a high-precision artificial sequence detection effect for the test sequence containing mismatched bases.

[0092] Alternatively, please refer to Figure 5 , Figure 5 yes Figure 2 The flowchart of step S220 includes the second sub-step. In the embodiments of this application, with Figure 4 Compared to the specific steps and flow of step S220 shown, Figure 5 The specific steps of step S220 shown may also include sub-steps S225 to S229 to achieve high-precision manual sequence detection for test sequences without mismatched bases but with soft clip reads. The soft clip reads are used to represent sequence reads that cannot be aligned to the reference genome but still exist in the SEQ (Segment SEQuence) field of the original alignment file.

[0093] Sub-step S225: Detect whether the sequence to be tested has a soft clip read relative to the reference genome.

[0094] In this embodiment, when the gene sequencing device 10 determines that a test sequence does not have mismatched bases relative to the corresponding target reference gene sequence, it will execute sub-step S225 to detect whether the test sequence without mismatched bases has a Soft Clip read. In one implementation of this embodiment, the presence of a Soft Clip read can be determined by detecting the presence of an S (Soft) symbol at the CIGAR information corresponding to the test sequence in the original alignment file.

[0095] In this embodiment, when the gene sequencing device 10 determines that a certain sequence to be tested does not have mismatched bases relative to the corresponding target reference gene sequence, and the sequence to be tested has a soft clip read relative to the reference genome, it will execute sub-step S226.

[0096] Sub-step S226: Extract the local sequencing sequence including the Soft Clip read from the sequence to be tested, and generate the reverse complementary pairing sequence of the local sequencing sequence.

[0097] In this embodiment, for any test sequence without mismatched bases but containing a Soft Clip read, the total number of base pairs in the Soft Clip read can be detected as greater than a third preset number (wherein, the third preset number is less than the first preset number, and the third preset number can be 6). If the total number of base pairs in the Soft Clip read is detected as greater than the third preset number, the Soft Clip read can be directly used as the local sequencing sequence of the test sequence. Alternatively, if the total number of base pairs in the Soft Clip read is detected as less than or equal to the third preset number, a first preset number of base pairs including the Soft Clip read and in a continuous distribution can be extracted from the test sequence to obtain the local sequencing sequence of the test sequence. This ensures that the obtained local sequencing sequence can effectively characterize the Soft Clip sequence of the corresponding test sequence.

[0098] Sub-step S227: Determine the local gene sequence in the reference genome where the corresponding target reference gene sequence is located.

[0099] In this embodiment, for any test sequence without mismatched bases but with a soft clip read, a target reference gene sequence corresponding to the positional mapping of the test sequence can be found among all reference gene sequences included in the reference genome. Then, based on the relative sequence positional relationship between all reference gene sequences in the reference genome, a local gene sequence including the target reference gene sequence is extracted from the reference genome. The local gene sequence consists of a second preset number (e.g., 300) of reference gene base pairs that are continuously distributed upstream and downstream of the corresponding target reference gene sequence (i.e., the target reference gene sequence). The second preset number is greater than the first preset number.

[0100] Sub-step S228 involves performing a base consistency comparison between the reverse complementary pairing sequence of the local sequencing sequence and the corresponding local gene sequence to obtain the base consistency score of the reverse complementary pairing sequence relative to the corresponding local gene sequence.

[0101] Sub-step S229: Based on the base consistency score of the reverse complementary pairing sequence, determine whether the sequence to be tested belongs to the target sequence with the characteristics of reverse complementary artificial sequence.

[0102] In this embodiment, when an inverse complementary pairing sequence associated with any test sequence that has no mismatched bases but contains a Soft Clip read is determined, and the base consistency score of the inverse complementary pairing sequence relative to the corresponding local gene sequence is determined, the test sequence is determined not to belong to the target sequence with the characteristics of an inverse complementary artificial sequence if the base consistency score of the inverse complementary pairing sequence is less than a first score threshold (e.g., 90%), and the test sequence is determined to belong to the target sequence with the characteristics of an inverse complementary artificial sequence if the base consistency score of the inverse complementary pairing sequence is greater than or equal to the first score threshold.

[0103] Therefore, by executing the above sub-steps S225 to S229, this application can achieve a high-precision manual sequence detection effect for test sequences without mismatched bases but with soft clip reads.

[0104] Optionally, in the embodiments of this application, with Figure 4 Compared to the specific steps and flow of step S220 shown, Figure 5 The specific steps of step S220 shown may also include sub-step S2210, which uses the test sequence without mismatched bases and without soft clip reads as a valid gene sequence that is not an artificial sequence.

[0105] Sub-step S2210: The sequence to be tested is directly treated as a valid sequence that is not the target sequence.

[0106] In this embodiment, when the gene sequencing device 10 determines that a certain sequence to be tested does not have mismatched bases relative to the corresponding target reference gene sequence, and the sequence to be tested does not have a soft clip read relative to the reference genome, it will determine that the sequence to be tested does not have the characteristics of an inverse complementary artificial sequence, and the sequence to be tested is not an artificial sequence. At this time, the gene sequencing device 10 will execute sub-step S2210.

[0107] Therefore, by executing the above sub-step S2210, this application can use the test sequence without mismatched bases and without soft clip reads as an effective gene sequence that is not an artificial sequence.

[0108] Step S230: Perform sequence removal processing on all target sequences present in the original alignment file to obtain the corresponding valid alignment file.

[0109] In this embodiment, when the gene sequencing device 10 determines all target sequences (i.e., artificial sequences) with reverse complementary artificial sequence characteristics in the original alignment file by executing the above sub-steps S221 to S2210, the original alignment file after sequence removal is essentially a valid alignment file that conforms to reality during the enzyme digestion library construction process. This effectively filters out false positive artificial sequences caused by the enzyme digestion method, reduces false positive mutations during the enzyme digestion library construction process, and can effectively improve the accuracy of the final somatic mutation detection results.

[0110] Therefore, by executing the above steps S210 to S230, this application can effectively filter out false positive artificial sequences caused by enzyme digestion and fragmentation based on the sequence characteristics of the reverse complementary sequence, reduce false positive mutations in the enzyme digestion library construction process, and improve the accuracy of the final somatic mutation detection results.

[0111] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0112] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the various functions provided in this application are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause the computer device, as a gene sequencing device 10, to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned readable storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code.

[0113] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for filtering artificial sequences via enzyme digestion, characterized in that, The method includes: Obtain the raw alignment file of the original enzyme digestion, library construction, and sequencing data relative to the reference genome; Based on the reference genome, reverse complementary artificial sequence feature detection is performed on all test sequences in the original alignment file to determine the target sequence with reverse complementary artificial sequence features; All target sequences present in the original alignment file are subjected to sequence removal processing to obtain the corresponding valid alignment file; The step of performing reverse complementarity artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine the target sequence with reverse complementarity artificial sequence features includes: For each test sequence, it is detected whether there are mismatched bases between the test sequence and the target reference gene sequence in the reference genome, wherein the target reference gene sequence is the reference gene sequence that maps the sequence position of the corresponding test sequence in the reference genome; If a mismatched base is detected in the sequence to be tested, the local test sequences corresponding to the two ends of the sequence to be tested are extracted, and the reverse complementary pairing sequences of the two local test sequences are generated respectively. Each local test sequence consists of a first preset number of base pairs of the sequence to be tested that are continuously distributed. The two reverse complementary pairs associated with the sequence to be tested are compared with the corresponding target reference gene sequence for base identity, and the base identity scores of the two reverse complementary pairs relative to the corresponding target reference gene sequence are obtained. Based on the base consistency scores of the two reverse complementary pairing sequences, it is determined whether the sequence to be tested belongs to the target sequence with the characteristics of reverse complementary artificial sequences.

2. The method according to claim 1, characterized in that, The step of obtaining the raw alignment file of the original enzyme digestion library construction sequencing data relative to the reference genome includes: The original enzyme digestion library construction sequencing data were subjected to sequencing quality assessment, and based on the sequencing quality assessment results, adapter sequences and low-quality sequences were removed from the original enzyme digestion library construction sequencing data to obtain initial sequencing data; Gene sequence alignment is performed between the initial sequencing data and the reference genome to obtain the corresponding initial alignment file; Based on the actual sequence positions of each reference gene sequence in the reference genome, the alignment sequences in the initial alignment file are sorted by position mapping, and PCR repeat sequences are removed from the sorted alignment sequences to obtain the original alignment file.

3. The method according to claim 1, characterized in that, The step of determining whether the test sequence belongs to the target sequence with the characteristics of an artificially paired reverse complementarity sequence based on the base consistency scores of the two reverse complementary sequences includes: Detect whether the base consistency score of each of the two reverse complementary pairing sequences is less than a first score threshold; If the base consistency scores of the two reverse complementary pairs are both less than the first score threshold, it is determined that the sequence to be tested does not belong to the target sequence with the characteristics of reverse complementary artificial sequences. If the base consistency score of at least one reverse complementary pairing sequence is detected to be greater than or equal to the first score threshold, the sequence to be tested is determined to be a target sequence with reverse complementary artificial sequence characteristics.

4. The method according to claim 1, characterized in that, The step of performing reverse complementarity artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine the target sequence with reverse complementarity artificial sequence features further includes: If no mismatched bases are detected in the test sequence, the presence of a soft clip read relative to the reference genome is detected. If a soft clip read is detected in the sequence to be tested, a local sequencing sequence including the soft clip read is extracted from the sequence to be tested, and a reverse complementary pairing sequence of the local sequencing sequence is generated. In the reference genome, a local gene sequence is determined where the corresponding target reference gene sequence is located. The local gene sequence is composed of the corresponding target reference gene sequence and a second preset number of reference gene base pairs that are continuously distributed upstream and downstream of it. The second preset number is greater than the first preset number. The reverse complementary pair sequence of the local sequencing sequence is compared with the corresponding local gene sequence for base consistency to obtain the base consistency score of the reverse complementary pair sequence relative to the corresponding local gene sequence. Based on the base consistency score of the reverse complementary pairing sequence, it is determined whether the sequence to be tested belongs to the target sequence with the characteristics of reverse complementary artificial sequences.

5. The method according to claim 4, characterized in that, The step of extracting the local sequencing sequence including the SoftClip read from the sequence to be tested includes: Detect whether the total number of base pairs in the Soft Clip read segment is greater than a third preset number, wherein the third preset number is less than the first preset number; If the total number of base pairs in the Soft Clip read is greater than the third preset number, the Soft Clip read is directly used as the local sequencing sequence of the sequence to be tested. If the total number of base pairs in the Soft Clip read is less than or equal to the third preset number, a first preset number of base pairs including the Soft Clip read and in a continuous distribution are extracted from the sequence to be tested to obtain the local sequencing sequence of the sequence to be tested.

6. The method according to claim 4, characterized in that, The step of determining whether the test sequence belongs to the target sequence with reverse complementary artificial sequence characteristics based on the base consistency score of the reverse complementary pairing sequence includes: The base consistency score of the reverse complementary pairing sequence is checked to see if it is less than the first score threshold. If the base consistency score of the detected reverse complementary pairing sequence is less than the first score threshold, it is determined that the corresponding test sequence does not belong to the target sequence with the characteristics of reverse complementary artificial sequence. If the base consistency score of the reverse complementary pairing sequence is detected to be greater than or equal to the first score threshold, the corresponding test sequence is determined to be a target sequence with reverse complementary artificial sequence characteristics.

7. The method according to any one of claims 4-6, characterized in that, The step of performing reverse complementarity artificial sequence feature detection on all test sequences in the original alignment file based on the reference genome to determine the target sequence with reverse complementarity artificial sequence features further includes: If no mismatched bases are detected in the test sequence and no soft clip reads are detected in the test sequence, the test sequence is directly regarded as a valid sequence that is not the target sequence.

8. A gene sequencing device, characterized in that, It includes a processor and a memory, the memory storing a computer program that can be executed by the processor, the processor executing the computer program to implement the enzyme digestion artificial sequence filtering method according to any one of claims 1-7.

9. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a computer device, it implements the enzyme digestion artificial sequence filtering method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Filtering method for breaking false positive mutation generated by artificial fragments in library building through enzyme digestion method

    CN116895332A