An apparatus, method, and computer readable storage medium for filtering low-complexity high-throughput sequencing data

By calculating the READ complexity of high-throughput sequencing data and comparing it with the complexity distribution of a reference genome sequence, low-complexity sequences were identified and deleted. This solved the problem of low-complexity base sequences introduced during PCR affecting sequencing accuracy and improved the accuracy of sequencing data analysis.

CN116312769BActive Publication Date: 2025-11-28SHENZHEN HAPLOX BIOTECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310071957.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-12
Publication Date
2025-11-28
Estimated Expiration
2043-01-12

AI Technical Summary

Technical Problem

Existing next-generation sequencing technologies introduce low-complexity base sequences during the PCR process, which affects the accuracy of sequencing data, especially in fields such as whole-genome sequencing, transcriptome sequencing, exon sequencing, and targeted capture sequencing, thus reducing the accuracy of bioinformatics analysis.

Method used

By calculating the complexity of each READ in the high-throughput sequencing data, using the complexity distribution of the reference genome sequence to identify and delete low-complexity sequences, and using the complexity threshold corresponding to the 95th percentile of the reference genome sequence, the true sequencing information is preserved.

Benefits of technology

It effectively removes low-complexity sequences, improves the accuracy of sequencing data, and retains the true sequencing READ information to the maximum extent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312769B_ABST
    Figure CN116312769B_ABST
Patent Text Reader

Abstract

The present application relates to gene detection and bioinformatics field, especially to a kind of filtering low-complexity high-throughput sequencing data device, method and computer readable storage medium, the method of the present application: the complexity of each READ in high-throughput sequencing data is compared with complexity threshold, and READ with lower complexity than complexity threshold is deleted, that is, the filtering of low-complexity high-throughput sequencing data is realized;The complexity threshold is the complexity corresponding to the 95% quantile in the frequency distribution of the complexity of all reference sequences in the reference genome sequence.The method of the present application is based on the information of existing NGS detection data and reference genome, accurately distinguishes and identifies the low-complexity repetitive base sequence generated in PCR process from the repetitive base sequence in reference genome, removes the low-complexity sequence generated in sequencing experiment process, maximizes the retention of real sequencing READ information, and improves the accuracy of sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of gene detection and bioinformatics, and particularly relates to a device and method for filtering low-complexity high-throughput sequencing data and a computer readable storage medium. BACKGROUND

[0002] Next-generation sequencing (NGS) is a DNA sequencing technology based on PCR and gene chip development, which is an important research method and tool for genomics and gene detection. Existing technical platforms mainly include Roche's 454FLX, Illumina's Miseq / Hiseq, etc.

[0003] In the application of next-generation sequencing (NGS), PCR amplification method is applied in the process of extracting library and sequencing. Since the PCR process will amplify the sequence sample, errors in base pairs will inevitably be introduced when the template chain is amplified. With the increasing instability of the polymerase related to PCR reaction, the probability of base error is also gradually increasing, resulting in a large number of low-complexity base sequences, which further leads to alignment errors and affects subsequent bioinformatics analysis. In cancer-related gene detection such as whole genome sequencing, transcriptome sequencing, exon sequencing, targeted capture sequencing and amplicon sequencing, the existence of low-complexity base sequences in high-throughput sequencing data will affect the bioinformatics analysis of downstream sequencing data, thereby reducing the accuracy of sequencing data analysis results. SUMMARY

[0004] The main purpose of the present application is to provide a device, method and computer readable storage medium for filtering low-complexity high-throughput sequencing data, which aims to identify and remove low-complexity sequences in high-throughput sequencing data, maximize the retention of true sequencing READ information, and improve the accuracy of sequencing.

[0005] To achieve the above purpose, the present application provides a device for filtering low-complexity high-throughput sequencing data, which comprises:

[0006] a. READ complexity calculation module in high-throughput sequencing data

[0007] This module calculates the complexity of each READ in high-throughput sequencing data by using the complexity calculation method, which is: according to the order of READ from 5' end to 3' end, the sum of the number of adjacent two bases is calculated, the sum is divided by the base length of the READ, then 1 is subtracted from the value, and multiplied by 100%;

[0008] b. Reference genome sequence complexity distribution determination module

[0009] The module is used for determining complexity of a reference genome sequence, and counting frequency distribution of the complexity.

[0010] In the probe design range or detection interval range, all reference sequences in the reference genome are obtained, and a sequence with a read length set by a sequencing machine is intercepted from a first base at a 5' end as a starting point; then, the starting point is moved by one base towards a 3' end, and a sequence with the read length set by the sequencing machine is intercepted again, the process is cycled, the starting point is sequentially moved by one base, and all sequences are intercepted; complexity of all sequences is calculated according to the complexity calculation method in the READ complexity calculation module in the high-throughput sequencing data, and frequency distribution of the complexity is counted;

[0011] c. A filtering module of low-complexity sequences

[0012] The module deletes the determined low-complexity sequences from the sequencing data, and realizes filtering of the low-complexity high-throughput sequencing data.

[0013] The low-complexity sequences are sequences with a READ complexity lower than a complexity threshold value calculated in the READ complexity calculation module in the high-throughput sequencing data, and the low-complexity sequence data is deleted, that is, filtering of the low-complexity high-throughput sequencing data is realized.

[0014] The complexity threshold value is a complexity corresponding to a 95% quantile in the frequency distribution determined by the complexity distribution determination module of the reference genome sequence.

[0015] d. An output module

[0016] The module is used for outputting the sequencing data after filtering of the low-complexity sequences.

[0017] In addition, in order to achieve the above object, the application further provides a method for filtering low-complexity high-throughput sequencing data, wherein the complexity of each READ in the high-throughput sequencing data is compared with a complexity threshold value, and the READ with a complexity lower than the complexity threshold value is deleted, that is, filtering of the low-complexity high-throughput sequencing data is realized.

[0018] The complexity calculation method is as follows: according to the order of the READ from a 5' end to a 3' end, the sum of the number of times of adjacent two bases being the same is counted, the sum is divided by the base length of the READ, then 1 is subtracted from the number, and the result is multiplied by 100%.

[0019] The complexity threshold value is a complexity corresponding to a 95% quantile in the frequency distribution of all reference sequence complexities in the reference genome sequence.

[0020] Further, the method comprises the following steps:

[0021] S1. calculating the complexity of the READ in the high-throughput sequencing data

[0022] The complexity of each READ in the high-throughput sequencing data is calculated, and the complexity calculation method is as follows: according to the order of the READ from the 5' end to the 3' end, the sum of the number of adjacent two bases being the same is counted, the sum is divided by the base length of the READ, then 1 is subtracted from the value, and the result is multiplied by 100%;

[0023] S2. Determining the complexity distribution of the reference genome sequence

[0024] In the probe design range or detection interval range, all reference sequences in the reference genome are obtained, and the sequence with a read length set by a sequencing machine is intercepted from the first base at the 5' end as a starting point; then, the starting point is moved one base towards the 3' end, and the sequence with a read length set by a sequencing machine is intercepted again, the process is repeated, and the starting point is sequentially moved one base until all sequences are intercepted; the complexity of all sequences is calculated according to the complexity calculation method described in step S1, and the frequency distribution of the complexity is counted;

[0025] S3. Determining and filtering low-complexity sequences

[0026] The sequence with a READ complexity lower than the complexity threshold value calculated in step S1 is a low-complexity sequence, and the low-complexity sequence is deleted from the sequencing data, thereby realizing filtering of the low-complexity high-throughput sequencing data;

[0027] The complexity threshold value is the complexity corresponding to the 95th percentile in the frequency distribution determined in step S2;

[0028] S4. Output

[0029] The sequencing data after filtering the low-complexity sequence is output.

[0030] Optionally, the reference genome is a human reference genome.

[0031] Optionally, the complexity threshold value is 30%.

[0032] Optionally, the high-throughput sequencing data includes high-throughput sequencing data of any detection platform, including DNA, RNA, and other sequencing data.

[0033] In addition, to achieve the above-mentioned purpose, the application further provides a computer readable storage medium, wherein the computer readable storage medium stores a program for filtering low-complexity high-throughput sequencing data, and the program for filtering low-complexity high-throughput sequencing data is executed by a processor to realize the steps of the above-mentioned method for filtering low-complexity high-throughput sequencing data.

[0034] The method for filtering low-complexity high-throughput sequencing data can accurately distinguish and identify low-complexity repetitive base sequences generated in a PCR process from repetitive base sequences in a reference genome based on information of detection data and the reference genome of an existing high-throughput sequencing technology (NGS), remove low-complexity sequences and error base sequences generated in a sequencing experiment process, and maximize the retention of true sequencing read information, thereby improving the accuracy of sequencing. The method is suitable for processing of DNA, RNA and other sequencing data. BRIEF DESCRIPTION OF DRAWINGS

[0035] Figure 1 A functional module schematic diagram of a device for filtering low-complexity high-throughput sequencing data according to an embodiment of the present application is shown in the figure.

[0036] Figure 2 A flowchart of a method for filtering low-complexity high-throughput sequencing data according to an embodiment of the present application is shown in the figure.

[0037] Figure 3 A structural schematic diagram of a computer readable storage medium according to an embodiment of the present application is shown in the figure.

[0038] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0039] It should be understood that the specific embodiments described herein are merely intended to explain the present application and not to limit the present application.

[0040] An apparatus for filtering low-complexity high-throughput sequencing data is provided in an embodiment of the present application, as shown in the figure. Figure 1 Figure 1 A functional module schematic diagram of a device for filtering low-complexity high-throughput sequencing data according to an embodiment of the present application is shown in the figure.

[0041] The apparatus for filtering low-complexity high-throughput sequencing data according to an embodiment of the present application comprises:

[0042] a. READ complexity calculation module 1 in high-throughput sequencing data

[0043] This module calculates the complexity of each READ in high-throughput sequencing data by using a complexity calculation method. The complexity calculation method is as follows: according to the order of READ from 5' end to 3' end, the sum of the number of adjacent two bases being the same is calculated, the sum is divided by the base length of the READ, then 1 is subtracted from the value, and the result is multiplied by 100%.

[0044] b. Reference genome sequence complexity distribution determination module 2

[0045] ​The module is used for determining complexity of a reference genome sequence, and counting frequency distribution of the complexity.

[0046] In the probe design range or detection interval range, all reference sequences in the reference genome are obtained, and a sequence with a read length set by a sequencing machine is intercepted from a first base at a 5' end as a starting point; then, the starting point is moved by one base towards a 3' end, and a sequence with the read length set by the sequencing machine is intercepted again, the process is cycled, the starting point is sequentially moved by one base, and all sequences are intercepted; complexity of all sequences is calculated according to the complexity calculation method in the READ complexity calculation module in the high-throughput sequencing data, and frequency distribution of the complexity is counted.

[0047] c. Low-complexity sequence filtering module 3

[0048] The module deletes the determined low-complexity sequence from the sequencing data, and realizes filtering of the low-complexity high-throughput sequencing data.

[0049] The sequence with a READ complexity calculated in the READ complexity calculation module in the high-throughput sequencing data and lower than the complexity threshold is a low-complexity sequence, and the low-complexity sequence data is deleted, that is, filtering of the low-complexity high-throughput sequencing data is realized.

[0050] The complexity threshold is a complexity corresponding to a 95% quantile in the frequency distribution determined by the complexity distribution determination module of the reference genome sequence, and is 30%.

[0051] d. Output module 4

[0052] The module is used for outputting the sequencing data after filtering of the low-complexity sequence.

[0053] In the filtering of the low-complexity high-throughput sequencing data, each functional module is respectively realized when running, and the steps of the filtering of the low-complexity high-throughput sequencing data are realized.

[0054] The embodiment of the application provides a method for filtering low-complexity high-throughput sequencing data. Figure 2 As shown in the figure, Figure 2 It is a flowchart of the method for filtering low-complexity high-throughput sequencing data.

[0055] The embodiment of the application provides a method for filtering low-complexity high-throughput sequencing data, and it should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from the order shown.

[0056] In the embodiment, the method for filtering low-complexity high-throughput sequencing data comprises:

[0057] S1. Calculate the complexity of each READ in high-throughput sequencing data

[0058] The complexity of each READ in high-throughput sequencing data is calculated as follows: according to the order of the READ from 5' end to 3' end, the sum of the number of adjacent two bases being the same is counted, the sum is divided by the length of the READ, then 1 is subtracted from the number, and the result is multiplied by 100%;

[0059] S2. Determine the complexity distribution of the reference genome sequence

[0060] In the probe design range or detection interval range, all reference sequences in the reference genome are obtained, and the sequence with a read length set by a sequencing machine is intercepted from the first base at the 5' end as a starting point; then, the starting point is moved one base towards the 3' end, and the sequence with a read length set by a sequencing machine is intercepted again, the process is repeated, and the starting point is moved one base at a time until all sequences are intercepted; the complexity of all sequences is calculated according to the complexity calculation method described in step S1, and the frequency distribution of the complexity is counted;

[0061] S3. Determination and filtering of low-complexity sequences

[0062] The sequence with a READ complexity calculated in step S1 lower than the complexity threshold is a low-complexity sequence, and the low-complexity sequence is deleted from the sequencing data, thereby realizing filtering of low-complexity high-throughput sequencing data;

[0063] The complexity threshold is 30% corresponding to the 95% quantile in the frequency distribution determined in step S2;

[0064] S4. Output

[0065] The sequencing data after filtering the low-complexity sequence is output.

[0066] In addition, the present application also provides a computer readable storage medium. Refer to Figure 3 , Figure 3 The structure of the computer readable storage medium involved in the embodiment of the present application is shown in the figure. The computer readable storage medium stores a program for filtering low-complexity high-throughput sequencing data, and the program for filtering low-complexity high-throughput sequencing data is executed by a processor to realize the steps of the method for filtering low-complexity high-throughput sequencing data as described above.

[0067] It should be noted that, in this document, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more limitations, foreclose the existence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0068] The above-mentioned embodiment numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0069] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and the necessary general hardware platform, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a computer readable storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a number of instructions for causing a device (which can be a mobile phone, a computer, a server, or a network device) to perform the methods described in the various embodiments of the present application.

[0070] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent flow transformation made by using the content of the specification and drawings, or directly or indirectly applied to other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An apparatus for filtering low-complexity high-throughput sequencing data, the apparatus comprising: The device comprises: a. READ complexity calculation module in high-throughput sequencing data This module calculates the complexity of each READ in high-throughput sequencing data by using the complexity calculation method, which is: according to the order of READ from 5' end to 3' end, counting the sum of the number of adjacent two bases being the same, dividing the sum by the base length of the READ, then subtracting the value from 1 and multiplying by 100%; b. Reference genome sequence complexity distribution determination module This module is used to determine the complexity of the reference genome sequence and to count the frequency distribution of the complexity; In the probe design range or detection interval range, all reference sequences in the reference genome are obtained; the complexity of all reference sequences is calculated according to the complexity calculation method in the READ complexity calculation module in high-throughput sequencing data, and the frequency distribution of the complexity is counted; c. Low complexity sequence filtering module This module deletes the determined low complexity sequence from the sequencing data, realizing the filtering of low complexity high-throughput sequencing data; The low complexity sequence is the sequence whose READ complexity calculated in the READ complexity calculation module in high-throughput sequencing data is lower than the complexity threshold, and the low complexity sequence data is deleted, which realizes the filtering of low complexity high-throughput sequencing data; The complexity threshold is the complexity corresponding to the 95% quantile in the frequency distribution determined by the reference genome sequence complexity distribution determination module; d. Output module This module is used to output the sequencing data after filtering the low complexity sequence.

2. A method of filtering low-complexity high-throughput sequencing data, characterized in that, The complexity of each READ in high-throughput sequencing data is compared with the complexity threshold, and the READ whose complexity is lower than the complexity threshold is deleted, which realizes the filtering of low complexity high-throughput sequencing data; The complexity calculation method is the complexity calculation method as claimed in claim 1; The complexity threshold is the complexity threshold as claimed in claim 1.

3. The method of claim 2, wherein, The method comprises the following steps: S1. Calculating the complexity of READ in high-throughput sequencing data The complexity of each READ in high-throughput sequencing data is calculated, and the complexity calculation method is the complexity calculation method as claimed in claim 1; S2. Determining the complexity distribution of reference genome sequence In the probe design range or detection interval range, all reference sequences in the reference genome are obtained; the complexity of all reference sequences is calculated according to the complexity calculation method as claimed in claim 1, and the frequency distribution of the complexity is counted; S3. Determination and filtering of low complexity sequence The sequence whose READ complexity calculated in step S1 is lower than the complexity threshold is a low complexity sequence, and the low complexity sequence is deleted from the sequencing data, which realizes the filtering of low complexity high-throughput sequencing data; The complexity threshold is the complexity corresponding to the 95% quantile in the frequency distribution determined in step S2; S4. Output Output the sequencing data after filtering the low complexity sequence.

4. The method of claim 3, wherein, The method for obtaining all reference sequences in the reference genome is: taking the first base at the 5' end as the starting point, and intercepting a sequence with a length of the read length set by the sequencing machine; then, moving the starting point by one base towards the 3' end, intercepting a sequence with a length of the read length set by the sequencing machine again, and repeating the process, moving the starting point by one base in turn until all sequences are intercepted.

5. The method of claim 2 or 3, wherein, The reference genome is a human reference genome.

6. The method of claim 2 or 3, wherein, The complexity threshold is 30%.

7. The method of claim 2 or 3, wherein, The high-throughput sequencing data comprises DNA or RNA sequencing data.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a program for filtering low-complexity high-throughput sequencing data, and the program for filtering low-complexity high-throughput sequencing data, when executed by the processor, implements the steps of the method for filtering low-complexity high-throughput sequencing data according to any one of claims 2 to 7.

Citation Information

Patent Citations

  • Method and device for correcting high-flux sequencing data

    CN109920480A

  • Analysis method for detecting microorganisms by using metagenome or metatranscriptome

    CN110349629A