A method for quality assessment of nucleic acid sequencing data

CN116246703BActive Publication Date: 2026-09-15CYGNUS BIOSCIENCES GUANGZHOU CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310295466.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2026-09-15
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

而现有的给每个碱基赋予一个质量值的做法,是Sanger测序时代的孑余,只是恰好适配Illumina的测序化学,所以大行于世,但实际上并不适合3’端开放的测序技术,存在诸多缺陷

Benefits of technology

[0035] This method uses polymers as the basic unit for sequencing quality assessment, assigning a quality value to the polymers that make up the nucleic acid sequence, rather than assigning a quality value to the bases. It is particularly suitable for sequencing reactions with open 3' ends, which helps to obtain more accurate bioinformatics analysis results, including the identification of variants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246703B_ABST
    Figure CN116246703B_ABST
Patent Text Reader

Abstract

The application discloses a nucleic acid sequencing data quality evaluation method, which takes a polymer in a nucleic acid sequence as a basic unit for quality evaluation instead of taking a base as a basic unit for quality evaluation in a prior art method, and is more suitable for a sequence obtained by a 3' end open sequencing method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed in this invention relates to a method for quality assessment of nucleic acid sequencing data, which belongs to the field of gene sequencing. Background Technology

[0002] Gene sequencing technology can determine the sequence of genetic material and is widely used in clinical tumor typing, microbial identification, and genetic disease diagnosis. Modern mainstream nucleic acid sequencing technologies, in addition to producing the sequence of the nucleic acid sample, also assign a quality value to each base detected to assess the accuracy of the measurement. This quality value is generally expressed in Phreed form.

[0003] q = -10log 10 (1-a)

[0004] In the formula, 'a' represents the accuracy of the base, and 'q' represents the Phred value. For example, the Phred values ​​corresponding to accuracy rates of 99%, 99.9%, and 99.99% are 20, 30, and 40, respectively.

[0005] In bioinformatics analysis of nucleic acid sequencing data, quality values ​​play a crucial role. For example, when identifying gene mutations, if a base in the sequenced sequence differs from its corresponding base in the reference sequence, a high quality value for that base indicates a gene mutation; conversely, a low quality value suggests a sequencing error and the absence of a gene mutation.

[0006] For 3' open sequencing technologies such as 454, Ion Torrent, and fluorescence sequencing, the reaction proceeds one polymer at a time, with each polymer being a single unit from a sequencing chemistry perspective. The existing practice of assigning a quality value to each base is a remnant of the Sanger sequencing era; it just happens to fit Illumina's sequencing chemistry, hence its widespread use, but it is actually unsuitable for 3' open sequencing technologies and has many drawbacks. For example, many sequencing technologies are prone to insertion or deletion errors when sequencing long homologous polymers, such as sequencing AAAA as AAAAA or AAA. This error occurs at the homologous polymer level, making it difficult to accurately assess the quality value of each base within that homologous polymer. In some cases, some bases are not prone to substitution errors and should therefore be retained when identifying single-base substitution mutations. However, these bases often have lower quality values ​​due to their susceptibility to insertion or deletion errors and are easily discarded during mutation identification, resulting in false negatives. Therefore, a quality assessment method more suitable for 3' open sequencing sequences is needed. Summary of the Invention

[0007] This invention discloses a quality assessment method for nucleic acid sequencing data, which uses polymers in the nucleic acid sequence as the basic unit for quality assessment, instead of using bases as the basic unit in existing methods. This method is more suitable for sequences obtained by sequencing methods with open 3' ends.

[0008] Specifically, the present invention provides a method for quality assessment of nucleic acid sequencing data, characterized by comprising:

[0009] Provide the nucleic acid sequence to be tested, and calculate the sequencing signal characteristics of the polymer as the basic unit;

[0010] The quality score of the polymer is predicted using a training-calibrated quantization scheme and based on the sequencing signal features.

[0011] The quantization scheme for training calibration includes:

[0012] For the provided standard nucleic acid sequence, the sequencing signal characteristics of the polymer are calculated using the polymer as the basic unit. Based on the alignment results of the standard nucleic acid sequence, the polymer is marked as correctly or incorrectly sequenced. A classifier is trained to fit the relationship between the sequencing signal characteristics of the polymer and its marking.

[0013] According to a preferred embodiment, the polymer includes homopolymers, binary copolymers, terpolymers, etc.

[0014] According to a preferred embodiment, the sequencing signal characteristics of the polymer refer to the characteristics of the signal generated when the polymer undergoes a sequencing chemical reaction during the sequencing process, including but not limited to the types of bases that make up the polymer, the length of the polymer, the number of rounds of the sequencing chemical reaction, the signal intensity, the degree to which the signal intensity (and its neighboring signal intensities) are close to integers, the parameters of the sequencing signal (unit signal, background signal, lead coefficient, lag coefficient, attenuation coefficient), the degree of phase loss when the polymer is detected, etc.

[0015] According to a preferred embodiment, fitting the relationship between the sequencing signal features of the polymer and its label includes converting the fitting result of the classifier into a quality score.

[0016] According to the preferred embodiment, a standard nucleic acid sample refers to a nucleic acid sample whose source and sequence have been determined and which is highly homozygous at almost all sites in the genome, including λ phage DNA, Escherichia coli DNA, Saccharomyces cerevisiae DNA, etc.

[0017] According to a preferred embodiment, the quality score refers to a numerical value that characterizes the accuracy of polymer sequencing, and is selected from accuracy, error rate, Phred value, etc.

[0018] According to a preferred embodiment, the quality score is logarithmically based on the polymer detection error probability, and the quality score includes Q10, Q15, Q20, Q25, Q30, Q35, Q40, Q45, Q50, Q55, Q60, etc.

[0019] According to preferred embodiments, the classifiers include, but are not limited to, linear regression, multinomial regression, logistic regression, support vector machines, artificial neural networks, random forests, Phred algorithms, ensemble learning, etc.

[0020] According to a preferred embodiment, training a classifier includes dividing polymers into several classes based on the sequencing signal characteristics of the polymers and calculating the sequencing accuracy of each class of polymers.

[0021] According to the preferred implementation, the classifier is trained based on the maximum likelihood probability distribution model; the probability distribution refers to a probability distribution with a unimodal shape feature, including but not limited to two-point distribution, binomial distribution, negative binomial distribution, Poisson distribution, geometric distribution, exponential distribution, normal distribution, Γ distribution, chi-square distribution, t distribution, F distribution, β distribution, log-normal distribution, and high-dimensional extensions of the above distributions.

[0022] According to a preferred embodiment, the method further includes performing bioinformatics analysis on the nucleic acid sequence to be tested based on the quality score of the polymer.

[0023] According to a preferred embodiment, bioinformatics analysis includes screening high-quality nucleic acid sequences based on assigned quality values. Screening methods include, but are not limited to, screening nucleic acid sequences where all quality values ​​are higher than or lower than a certain threshold, screening nucleic acid sequences where the average quality value is higher than or lower than a certain threshold, screening regions within nucleic acid sequences where all quality values ​​are higher than or lower than a certain threshold, and screening regions where the average quality value is higher than or lower than a certain threshold, etc.

[0024] According to a preferred embodiment, bioinformatics analysis includes aligning a nucleic acid sequence to a reference sequence based on an assigned quality value.

[0025] According to a preferred embodiment, bioinformatics analysis includes identifying gene variations based on the alignment results and the quality value assigned to the aligned sequence.

[0026] According to a preferred embodiment, when identifying gene variations, certain features of the comparison results can be used to remove potential false positive or false negative results.

[0027] According to a preferred embodiment, bioinformatics analysis includes assembling nucleic acid sequences into longer nucleic acid sequences based on assigned quality values.

[0028] According to a preferred embodiment, the bioinformatics analysis includes performing at least two orthogonal degenerate sequencing operations to obtain a quality value of the degenerate polymer length, and then using the quality value for error correction.

[0029] The present invention also provides a method for quality assessment of nucleic acid sequencing data, characterized in that it includes: performing fuzzy sequencing or deletion sequencing on the nucleic acid sample to be tested to obtain input data, generating degenerate polymer length information of the input data, and calculating sequencing signal characteristics of the degenerate polymer length;

[0030] The quality score of the polymer is predicted using a quantization scheme for training calibration and based on the sequencing signal features.

[0031] The quantization scheme for training calibration includes:

[0032] Nucleic acid sequences are obtained by sequencing standard nucleic acid samples. The sequencing signal characteristics of the nucleic acid sequences are calculated. The nucleic acid sequences are aligned to reference sequences, and the polymers of the nucleic acid sequences are marked as correctly or incorrectly sequenced. A classifier is trained to fit the relationship between the sequencing signal characteristics of the polymers and their markings.

[0033] Advantages of the present invention

[0034] Compared with the prior art, the method of the present invention has the following advantages:

[0035] This method uses polymers as the basic unit for sequencing quality assessment, assigning a quality value to the polymers that make up the nucleic acid sequence, rather than assigning a quality value to the bases. It is particularly suitable for sequencing reactions with open 3' ends, which helps to obtain more accurate bioinformatics analysis results, including the identification of variants. Attached Figure Description

[0036] The novel features of the invention are set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description and accompanying drawings, which describe illustrative embodiments utilizing the principles of the invention, in which:

[0037] Figure 1 Examples illustrating the characteristics of sequencing signals are provided. Detailed Implementation

[0038] To further illustrate the core content of this invention, the following examples are provided. These examples are intended to further explain the invention and do not constitute a limitation on it.

[0039] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. To better disclose the methods and content of this invention, detailed explanations are provided for the key terms used in this invention.

[0040] Terminology Explanation

[0041] Each

[0042] The term "each" is intended to identify an individual item in the collection, but does not necessarily refer to every item in the collection. Exceptions may occur if explicitly stated or otherwise specified in the context.

[0043] include

[0044] The term "includes" is intended to be open-ended in this document, encompassing not only the listed elements but also any additional elements.

[0045] sample

[0046] The term "sample" in this invention refers to a sample containing nucleic acids or mixtures of nucleic acids, typically derived from biological fluids, cells, tissues, organs, or organisms, wherein the nucleic acids or mixtures contain at least one nucleic acid sequence to be sequenced and / or phased. Such samples include, but are not limited to, blood, blood fractions, sputum / oral fluid, amniotic fluid, fine-needle biopsy samples (e.g., surgical biopsy, fine-needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, tissue explants, organ cultures, and any other tissue or cell preparations, fractions or derivatives thereof, or fractions or derivatives isolated therefrom. While samples are typically taken from human subjects (e.g., patients), samples can be taken from any organism possessing chromosomes, including but not limited to bacteria, viruses, fungi, birds, mammals, etc. Samples can be used directly as is from their biological source or after pretreatment to alter their properties. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may also include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, etc.

[0047] Standard nucleic acid samples

[0048] This refers to nucleic acid samples whose origin and sequence are known, and which are highly homozygous at almost all sites in the genome. The advantage of sequencing using standard nucleic acid samples is that it allows for a more accurate differentiation between sequences in the sequencing results that differ from the reference genome and whether these are variations or sequencing errors. For example, standard nucleic acid samples can be nucleic acids from *E. coli*, *Saccharomyces cerevisiae*, or λ phage DNA produced by New England Biolabs.

[0049] Error correction code Correcting Code (ECC) sequencing and error correction

[0050] In this application, error-correcting code sequencing has the following characteristics:

[0051] This sequencing method requires multiple sequencing runs, each run yielding incomplete information, while the total information from multiple runs is redundant. This redundancy is utilized to detect and correct potential sequencing errors, resulting in highly accurate sequences. For example, in 2+2 sequencing, sequencing reagents are divided into three pairs of matched bases (e.g., MK, RY, and WS), and the DNA sequence to be tested is sequenced three times independently, generating three degenerate sequence codes. These three codes can cross-check each other, allowing not only the derivation of the actual base sequence information through decoding but also the ability to correct errors from single-run sequencing. This correction process is called error correction.

[0052] sequence

[0053] In this invention, a "sequence" is a nucleic acid sequence, comprising or representing chains of nucleotides coupled together. It can be a defined nucleotide sequence or a degenerate base sequence, and can be based on DNA or RNA. It should be understood that a sequence can include multiple sub-sequences. For example, a single sequence (e.g., the sequence of a PCR amplicon) can have 350 nucleotides. A sample read can include multiple sub-sequences within these 350 nucleotides. The sequence is not divided into basic units of single bases, but rather into basic units of polymers, which can be 1 bp or longer.

[0054] Degenerate bases

[0055] According to the IUPAC symbol naming rules (Nucleic acid notation), the letters in Table 1 below are used to represent degenerate bases. For example, the letter M represents A and / or C; the degenerate polymer MMKKK has a length of 5, that is, DPL is 5.

[0056] Table 1

[0057] M A / C K G / T R A / G Y C / T W A / T S C / G B C / G / T D A / G / T H A / C / T V A / C / G

[0058] Reference sequence

[0059] A reference sequence is any specific known genome sequence of any organism, whether partial or complete, that can be used to reference an identified sequence from a subject. For example, reference genomes for human subjects and many other organisms can be found at the National Center for Biotechnology Information (NCBI) at ncbi.nlm.nih.gov. A “genome” refers to the complete genetic information of an organism or virus expressed as a nucleic acid sequence. A genome includes both genes and the non-coding sequence of DNA. Reference sequences can be larger than the reads they are aligned to. For example, a reference sequence can be at least approximately 100 times larger, or at least approximately 1000 times larger, or at least approximately 10,000 times larger, or at least approximately 10... 5 Larger, or at least about 10 6 Larger, or at least about 10 7 The reference genome sequence is, in one example, a sequence of the full-length human genome. In another example, the reference genome sequence is limited to a specific human chromosome, such as chromosome 13. In some implementations, the reference chromosome is a chromosome sequence from human genome version hg19. Such sequences may be referred to as chromosome reference sequences, but the term reference genome is intended to encompass such sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, subchromosome regions (such as chains), etc., of any species. In various implementations, the reference genome is a shared sequence or other combination derived from multiple individuals. However, in some applications, the reference sequence may be taken from a specific individual. In other implementations, “genome” also encompasses the so-called “graphic genome,” which uses a specific storage format and representation of the genome sequence. In one implementation, a graphical genome stores data in a linear file.

[0060] Alignment (or alignment)

[0061] Alignment is a common concept in bioinformatics, where it is frequently used to compare the similarity between different nucleic acids or different proteins. In this invention, alignment refers to the process of comparing a nucleic acid sequence with a reference sequence to determine whether the reference sequence contains a nucleic acid sequence. Commonly used sequence alignment algorithms and software include, but are not limited to, Smith-Waterman algorithm, Bowtie, BWA, SOAP, Needleman-Wunch algorithm, Bowtie2, BLAST, ELAND, TMAP, MAQ, minimap2, SHRiMP, etc.

[0062] Quality Score

[0063] The quality value, also known as the sequencing accuracy, is a numerical value that characterizes the sequencing accuracy. Quality values ​​can be expressed in different mathematical ways, such as accuracy, error rate, and Phred value. For example, accuracy rates of 99%, 99.9%, and 99.99% correspond to error rates of 1%, 0.1%, and 0.01%, respectively, and corresponding Phred values ​​of 20, 30, and 40. In some implementations, for ease of recording and storage, the Phred value is converted to ASCII code by adding 33; for example, Phred values ​​of 20, 30, and 40 are converted to the characters '5', '?', and 'I', respectively. The different forms of quality value expression do not affect the essence of this invention.

[0064] Fuzzy sequencing

[0065] In this invention, fuzzy sequencing refers to 2+2 sequencing in a single round, 1+3 sequencing in a single round, or 3x4 sequencing in a single round. Specifically, 2+2 sequencing refers to sequencing that includes two different sequencing reagents: a first sequencing reagent and a second sequencing reagent; the two sequencing reagents are added cyclically; wherein the first sequencing reagent contains two different nucleotide monomers with detectable labels; the second sequencing reagent contains two different nucleotide monomers with detectable labels, and the nucleotide monomers are different from those present in the first sequencing reagent; and wherein the second sequencing reagent is provided after the first sequencing reagent is provided, and the signal generated by the detectable label, such as a fluorescence signal, is detected after the nucleotide monomers are incorporated into the nucleic acid to be tested; in 2+2 sequencing, the combination of fluorescently labeled nucleotides can be used to obtain a fluorescence signal value associated with the target DNA sequence. Possible combinations are as follows: M / K mode: dA4P and dC4P are presented in odd-numbered rounds, and dG4P and dT4P are presented in even-numbered rounds; or vice versa; R / Y mode: dA4P and dG4P are presented in odd-numbered rounds, and dC4P and dT4P are presented in even-numbered rounds; or vice versa; and W / S mode: dA4P and dT4P are presented in odd-numbered rounds, and dC4P and dG4P are presented in even-numbered rounds; or vice versa. Sequencing the nucleic acid to be tested using one of the above modes (e.g., M / K mode) for multiple cycles can be called a round (or a single round). 1+3 sequencing, similar to 2+2 sequencing, refers to sequencing that includes two different sequencing reagents: a first sequencing reagent and a second sequencing reagent; these reagents are added cyclically; the first sequencing reagent contains three different nucleotide monomers with detectable labels; the second sequencing reagent contains one nucleotide monomer with a detectable label, and this nucleotide monomer is different from the nucleotide monomer present in the first sequencing reagent; the second sequencing reagent is provided after the first sequencing reagent, and the nucleotide monomer is incorporated into the nucleic acid to be tested, followed by detection of the signal generated by the detectable label, such as a fluorescence signal. 3x4 sequencing refers to sequencing that includes four different sequencing reagents: a first sequencing reagent, a second sequencing reagent, a third sequencing reagent, a fourth sequencing reagent, a fifth sequencing reagent, a sixth sequencing reagent, a seventh sequencing reagent, a seventh sequencing reagent, a thief ... The sequencing reagents consist of a second sequencing reagent, a third sequencing reagent, and a fourth sequencing reagent; these four sequencing reagents are added cyclically; each sequencing reagent contains three different nucleotide monomers with detectable labels; the second sequencing reagent is provided after the first sequencing reagent is provided, and the signal generated by the detectable label is detected after the nucleotide monomers are incorporated into the nucleic acid to be tested; the third sequencing reagent is provided after the second sequencing reagent is provided, and the signal generated by the detectable label is detected after the nucleotide monomers are incorporated into the nucleic acid to be tested; the fourth sequencing reagent is provided after the third sequencing reagent is provided, and the signal generated by the detectable label is detected after the nucleotide monomers are incorporated into the nucleic acid to be tested; for example, the four sequencing reagents mentioned above are B, D, H, and V.

[0066] The fuzzy sequence information obtained from the signal refers to the base sequence information from which the nucleotide sequence cannot be definitively determined. Here, the definitively determined base sequence information refers to nucleic acid sequences encoded by A, G, T, C, or by A, G, U, C; the bases can be methylated bases. Fuzzy base sequences are a common concept in scientific research; for example, the letter W can represent bases A and / or T. A related definition can also be found on Wikipedia (https: / / en.wikipedia.org / wiki / Nucleotide).

[0067] Deletion sequencing and deleted sequences

[0068] Multiple sequencing chemical reaction cycles are performed on the nucleic acid molecule to be tested on a gene sequencing chip; at least one cycle is pre-selected, and signal acquisition is performed only in the aforementioned cycle; in other cycles, only sequencing chemical reactions are performed, and no signal acquisition is performed; the acquired signals are encoded into sequences to obtain the missing nucleic acid sequence, and this process is called deletion sequencing.

[0069] A missing sequence is a sequence with missing information. Here, missing information means that some sequence information in the obtained sequence is missing. For example, the sequence ATTCGNNTTT is a missing sequence. N indicates that the sequence information is unknown, that is, the sequence information is missing here, so this sequence is a missing sequence.

[0070] Multibase sequencing

[0071] Multibase sequencing, also known as 3' open sequencing, involves sequencing reactions where the 3' end of the substrate nucleotide molecule is a hydroxyl group, allowing for free elongation. Theoretically, a single sequencing reaction can elongate one or more nucleotide molecules. Common multibase sequencing methods include Ion Torrent's semiconductor sequencing, 454 pyrosequencing, and fuzzy sequencing from Seine Biotech.

[0072] Unit signal

[0073] The unit signal is the increase in signal detected by the sequencer for each base extension of DNA. It is related to the number of DNA molecules involved in the extension reaction, camera exposure time, excitation light intensity, and camera photosensitivity. The unit signal refers to the signal intensity at each sequencing site where one base is extended; it is a physical quantity proportional to the number of template DNA molecules at that sequencing site. The intensity of the sequencing signal is equal to the unit signal multiplied by the number of bases involved in the extension reaction. For any sequencing technology, the measurement accuracy of the unit signal directly affects the accuracy of the sequencing results.

[0074] Background signal

[0075] Background signal refers to the baseline signal detected by the sequencer when there is no base extension, and it is related to factors such as chip material and spontaneous hydrolysis of the sequencing reaction substrate. Furthermore, the background signal can change as the sequencing read length increases. Background signal is a general definition.

[0076] Degenerate polymer length (DPL)

[0077] Degenerate sequencing is a type of multibase sequencing. Unlike single-base sequencing, which extends only one nucleotide molecule per reaction, multibase sequencing may extend multiple nucleotides per reaction. The intensity of the fluorescence signal released by the sequencing reaction is positively correlated with the number of fluorescent groups released. Under ideal conditions without attenuation or phase loss, the fluorescence signal of each reaction reflects the number of bases extended in that round, which is called the degenerate polymer length (DPL).

[0078] wheel (cycle)

[0079] A sequencing cycle, also known as a sequencing reaction round, refers to the process of using a sequencing reagent to incorporate a nucleotide with a detectable label into the nucleic acid to be tested, and then detecting the signal generated by the detectable label.

[0080] polymers

[0081] The polymers described in this invention include homopolymers, binary copolymers, and terpolymers, and the length of the polymers can be 1 bp or longer.

[0082] Homopolymers

[0083] The homopolymers or homopolymers in this invention refer to polymers composed of multiple homonucleotide monomers, such as AAAA or TTTTT, which are homopolymers.

[0084] binary copolymer

[0085] Binary copolymers are polymers with two different monomer units generated by the polymerization of two different monomers. In this invention, they specifically refer to polymers composed of two nucleotide monomers. For example, when performing MK degenerate sequencing, M is composed of two nucleotide monomers, A and C, and K is composed of two nucleotide monomers, G and T. ACAC and GGTT are both binary copolymers.

[0086] Terpolymer

[0087] Similar to binary copolymers, ternary copolymers are polymers composed of three different monomers. In this invention, they specifically refer to polymers composed of three nucleotide monomers. For example, when performing AB degenerate sequencing, B may be composed of three nucleotide monomers: C, G, and T. TCGG and GTGCC are both ternary copolymers.

[0088] Classifier

[0089] Classification is a very important method in data mining. In machine learning, the role of a classifier is to determine the category to which a new observation sample belongs based on the training data with labeled categories.

[0090] The construction and implementation of a classifier generally involves the following steps:

[0091] Select samples (including positive and negative samples), and divide all samples into two parts: training samples and test samples;

[0092] The classifier algorithm is executed on the training samples to generate a classification model;

[0093] The classification model is executed on the test samples to generate prediction results.

[0094] Preferably, based on the prediction results, the necessary evaluation indicators are calculated to evaluate the performance of the classification model.

[0095] Gene mutation

[0096] Somatic variants refer to nucleic acid sequences that differ from a reference sequence. Typical variants include, but are not limited to, single nucleotide variants (SNVs), short deletions and insertions (Indels), copy number variants (CNVs), epigenetic variants, microsatellite markers or short tandem repeats, and structural variants. Somatic variant detection is the work of identifying variants present at low frequencies in DNA samples. Somatic variant detection is particularly meaningful in the context of cancer treatment. Cancer is caused by the accumulation of mutations in DNA. DNA samples from tumors are often heterogeneous, including some normal cells, some cells in the early stages of cancer progression (with fewer mutations), and some late-stage cells (with more mutations). Due to this heterogeneity, when sequencing tumors (e.g., from FFPE samples), somatic variants will typically appear at low frequencies. For example, SNVs may be seen in only 10% of reads covering a given base. Variants to be classified as somatic or germline by a variant classifier are also referred to as “genetic variants” in this paper.

[0097] threshold

[0098] In this document, a threshold refers to a numeric or non-numeric value used as a cutoff value to characterize a sample, nucleic acid, or a portion thereof (e.g., a read). Thresholds may be varied based on empirical analysis. Thresholds can be compared to measurements or calculated values ​​to determine whether sources producing such values ​​should be classified in a particular manner. Thresholds can be identified based on experience or analysis. The choice of a threshold depends on the level of confidence the user expects to have to classify. Thresholds can be selected for specific purposes (e.g., to balance sensitivity and selectivity). As used herein, the term "threshold" indicates a point at which the analytical process can be altered and / or an action can be triggered. A threshold does not need to be a predetermined number. Instead, a threshold can be, for example, a function based on multiple factors. Thresholds can be adjusted as needed. Furthermore, a threshold can indicate an upper limit, a lower limit, or a range between limits.

[0099] The general meanings of the terms used in this invention have been explained above. These terms have conventional meanings in the art and are repeated here to avoid ambiguity. The terms themselves have no specific meaning. Invention Details

[0101] Quality assessment of nucleic acid sequencing data plays a crucial role in subsequent bioinformatics analysis. For example, in identifying gene mutations, if a base has a high quality value and differs from its corresponding base in the reference sequence, it is considered a gene mutation. Conversely, if the quality value of that base is low, the sequence is considered to have a sequencing error and no gene mutation exists. Current techniques that assign a quality value to each base as a single unit have several drawbacks. For instance, many sequencing technologies are prone to insertion or deletion errors when sequencing long homologous polymers, such as sequencing TTTT as TTT or TTTTT. Because these errors occur on homologous polymers, it is difficult to accurately assess the quality value of each base within that homologous polymer. Furthermore, some bases are not prone to substitution errors and should be retained when identifying single-base substitution mutations. However, these bases often have low quality values ​​due to their susceptibility to insertion or deletion errors and are easily discarded during mutation identification, leading to false negatives. To overcome the shortcomings of existing technologies, this invention discloses a new method for assessing the quality of sequencing data. Instead of using bases as the basic unit, it uses polymers as the basic unit for scoring, which is particularly suitable for sequencing reactions with open 3' ends and helps to obtain more accurate bioinformatics analysis results.

[0102] Specifically, the present invention provides a method for quality assessment of nucleic acid sequencing data, characterized by comprising:

[0103] Provide the nucleic acid sequence to be tested, and calculate the sequencing signal characteristics of the polymer as the basic unit;

[0104] The quality score of the polymer is predicted using a training-calibration quantization scheme based on sequencing signal features; the training-calibration quantization scheme includes:

[0105] For the provided standard nucleic acid sequence, the sequencing signal characteristics of the polymer are calculated using the polymer as the basic unit. Based on the alignment results of the standard nucleic acid sequence, the polymer is marked as correctly or incorrectly sequenced. A classifier is trained to fit the relationship between the sequencing signal characteristics of the polymer and its marking.

[0106] In this invention, sequencing methods for obtaining nucleic acid sequences include dideoxynucleotide termination (Sanger sequencing), chemical degradation (Gilbert method), pyrosequencing, semiconductor sequencing, cyclic reversible terminator, fluorogenic sequencing, error-correction code sequencing, fuzzy sequencing, deletion sequencing (patent CN 2022101040373), combined probe-anchor ligation, combined probe-anchor polymerization, sequencing by oligonucleotide ligation and detection, sequencing-by-binding, single-molecule fluorescence sequencing, single-molecule real-time sequencing, and nanopore sequencing.

[0107] According to the preferred embodiments, the sequencing method is preferably selected from sequencing reactions with an open 3' end, including but not limited to pyrosequencing, semiconductor sequencing, fluorescence sequencing, error-correcting code sequencing, fuzzy sequencing, deletion sequencing, nanopore sequencing, etc. In the above sequencing technologies, one or more nucleotide molecules can be extended in a sequencing chemical reaction cycle. That is to say, the reaction in each cycle occurs at the polymer level, rather than at the base level. Therefore, using polymers as the basic unit for quality assessment can undoubtedly more accurately reflect the actual sequencing process, and the obtained sequencing quality assessment results are more accurate.

[0108] In this invention, nucleic acid samples include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), peptide nucleic acid (PNA), xylonucleic acid (XNA), locked nucleic acid (LNA), etc.

[0109] In some implementations, the nucleic acid sample to be tested and the standard nucleic acid sample can be labeled separately and then sequenced simultaneously using the same sequencing method to obtain corresponding nucleic acid sequences. In some implementations, the standard nucleic acid sample can be sequenced first to obtain its corresponding nucleic acid sequence; then, the nucleic acid sample to be tested can be sequenced using the same sequencing method to obtain its corresponding nucleic acid sequence. In some implementations, the nucleic acid sample to be tested can also be sequenced first to obtain its corresponding nucleic acid sequence; then, the standard nucleic acid sample can be sequenced to obtain its corresponding nucleic acid sequence. The sequencing order of the sample to be tested and the standard nucleic acid sample can be interchanged; the important thing is that both must use the same sequencing method for sequencing and base identification.

[0110] According to a preferred embodiment, the input data is image data, which originates from sequencing images generated by the sequencer during sequencing operations. For example, the input data generated by sequencing methods such as pyrosequencing, fluorogenic sequencing, error-correction code sequencing, cyclic reversible terminator, and combinedatorial probe-anchor ligation are image data.

[0111] In some implementations, the input data is based on pH changes caused by the release of hydrogen ions during the elongation of the nucleotide substrate molecule. The pH changes are detected and converted into voltage changes proportional to the amount of incorporated nucleotides. For example, this is the input data generated by Ion Torrent's semiconductor sequencing method.

[0112] In some implementations, the input data is created based on nanopore sensing, which uses a biosensor to measure interruptions in current as an analyte passes through or near the opening of a nanopore, while simultaneously determining the type of base. For example, Oxford Nanopore Sequencing (ONT) is based on the concept of passing single-stranded DNA (or RNA) through a nanopore membrane with a voltage difference applied across the membrane. The nucleotides present in the pore affect the pore's resistance, and thus, current measurements over time indicate the sequence of DNA bases passing through the pore. This current signal (referred to as a "squiggle" due to its appearance when plotted) is raw data collected by the ONT sequencer. These measurements are stored as 16-bit integer data acquisition (DAC) values ​​acquired at, for example, a frequency of 4 kHz. At a DNA strand speed of approximately 450 base pairs / second, this yields an average of approximately nine raw observations per base. The signal is then processed to identify interruptions in the opening signal corresponding to each reading. Maximizing these uses of the raw signal involves base detection, which is the process of converting DAC values ​​into DNA base sequences. In some implementations, the input data includes normalized or scaled DAC values.

[0113] In this invention, polymers include homopolymers, binary copolymers, ternary copolymers, etc. The detected nucleic acid sequence is divided into several polymers. During this division, all polymers can be homopolymers, all binary copolymers, all ternary copolymers, or a combination of homopolymers, binary copolymers, and ternary copolymers. Each polymer can be 1 bp or longer, for example, 5 bp, 10 bp, or more than 10 bp.

[0114] The division of polymers should correspond to the sample introduction process during sequencing. For example, if the sequencing method is MK degenerate sequencing, the obtained nucleic acid sequence may be MMKKKMKMMM, which is a degenerate sequence. Therefore, when dividing the polymers, it should be divided according to the actual extension of MK sequencing, i.e., (MM)(KKK)(M)(K)(MMM). If the sequencing method is AB degenerate sequencing, the obtained sequence may be ABBAAABB, so when dividing the polymers, it should be divided according to the actual extension of AB sequencing, i.e., (A)(BBB)(AAA)(BB). If the sequencing method is 1x4, the obtained sequence may be ACCCTTGGATT, which is a defined nucleotide sequence. Therefore, when dividing the polymers, it should be divided according to the actual extension of 1x4 sequencing, with homopolymer as the basic unit, i.e., (A)(CCC)(TT)(GG)(A)(TT).

[0115] In a preferred embodiment, the nucleic acid sequence is a defined nucleotide sequence, i.e., a sequence represented by A, G, C, T, or a sequence represented by A, G, C, U.

[0116] In some implementations, the nucleic acid sequence is a degenerate sequence, or an ambiguous sequence, meaning that the sequence contains uncertain information, represented by degenerate bases such as M, K, R, Y, W, S, B, D, H, V, etc. Understandably, parts of the sequence may be definite.

[0117] In this invention, a standard nucleic acid sample refers to a nucleic acid sample whose source and sequence have been determined and which is highly homozygous at almost all sites in the genome. For example, it can be λ phage DNA produced by New England Biolabs. Since the reference sequence of the standard nucleic acid is known, the sequenced nucleic acid sequence can be aligned to its corresponding reference nucleic acid sequence, and each polymer can be marked as correct or incorrect. For example, if a sequenced nucleic acid sequence is ATTGGCCAAAT, it can be divided into three binary copolymers: (ATT)(GGCC)(AAAT). The reference sequence is ATTGGCCAAAA. The first and second polymers are marked as correct, and the third polymer is marked as incorrect.

[0118] In a specific implementation, the sequenced nucleic acid sequence is aligned to its corresponding reference nucleic acid sequence to obtain the alignment result. Based on the alignment result, polymers are marked as correctly sequenced or incorrectly sequenced. Preferably, high-quality aligned sequences are further screened from the alignment results, and polymers in the high-quality aligned sequences are marked as correctly sequenced or incorrectly sequenced, ignoring sequences that cannot be determined (i.e., bases that cannot be successfully aligned to the reference sequence or bases with low alignment quality). Based on the alignment result, polymers with an alignment result of "match" are marked as "correctly sequenced," and polymers with alignment results of "mismatch," "insertion," or "deletion" are marked as "incorrectly sequenced." The high-quality alignment described in this invention requires specific selection of the quality value range based on the alignment software or algorithm used. For example, when using BWA for sequence alignment, a high-quality aligned sequence refers to a base sequence with an alignment quality greater than 0, 10, 20, 30, 40, 50, or 60.

[0119] According to a preferred embodiment, the sequencing signal characteristics corresponding to the polymer refer to the characteristics of the signal generated when the polymer undergoes a sequencing chemical reaction during the sequencing process. Figure 1Examples of sequencing signal characteristics are given, including but not limited to: the type of the base, i.e., whether the base belongs to A, G, C, T (or U); the position of the base in the sequence, i.e., the position of the base in its nucleotide sequence; for example, for single-end sequencing, bases located earlier in the sequence generally have higher sequencing quality values ​​than bases located later in the sequence; the length of the polymer in which the base is located, i.e., the number of bases in the homopolymer or degenerate polymer in which the base is located; generally, shorter polymer lengths result in higher sequencing quality values; the position of the base within the polymer, i.e., the distance between the base and the nearest end of the homopolymer or degenerate polymer in which the base is located; and the number of rounds of sequencing chemical reactions in which the base is incorporated. The sequence number of nucleotides corresponds to the cycle number; generally, a smaller cycle number indicates a higher quality value. Signal intensity can be the intensity of the signal directly acquired by the sequencer, including brightness, voltage level, or current level; it can be a normalized signal or a phase-corrected signal. The degree to which the signal intensity (and its neighboring signal intensities) approaches an integer, i.e., the difference between the normalized signal, the phase-corrected signal, or the error-corrected signal and the nearest integer; generally, a smaller difference indicates higher accuracy. Sequencing signal parameters include unit signal, background signal, lead coefficient, hysteresis coefficient, and attenuation coefficient. The degree of phase loss when the base is detected; generally, a lower degree of phase loss indicates higher accuracy, and so on.

[0120] According to a preferred embodiment, one or more sequencing signal features of the polymer are calculated. For example, only one sequencing signal feature may be selected for calculation, or two or more sequencing signal features may be selected.

[0121] Furthermore, after labeling each polymer of the standard nucleic acid sequence, a classifier is trained to fit the relationship between the sequencing signal features of the polymer and its label. Classifiers are a common concept in pattern recognition, including but not limited to linear regression, multinomial regression, logistic regression, support vector machines, artificial neural networks, random forests, Phred algorithms, and ensemble learning. With the development of pattern recognition, many novel classifier algorithms have been proposed in recent years. Using novel classifier algorithms does not change the essence of this invention.

[0122] Specifically, a classifier is trained to fit the relationship between the sequencing signal features of each polymer and its label. This classifier can divide polymers into several classes based on their sequencing signal features and calculate the accuracy of each class. For example, polymers with lengths of 1, 2, 3, 4, 5, and greater than 5 can be grouped into one class, or polymers with unit signals in the ranges of 100-199, 200-299, 300-399, and >400 can be grouped into another class. When using multiple sequencing signal features, orthogonal partitioning can be performed. For example, polymers with a length of 1 and a unit signal within the range of 100-199 can be grouped into one class, polymers with a length of 1 and a unit signal within the range of 200-299 can be grouped into another class, and so on.

[0123] In a preferred embodiment, after fitting, the classifier's fitting result is converted into a quality score. Numerous studies report on how to convert classifier predictions into quality values. Taking the well-known softmax algorithm as an example, suppose the output of a classifier is (a,b), where (1,0) represents correct and (0,1) represents incorrect. Due to factors such as the accuracy of classifier training or computational errors during prediction, the classifier's output during prediction is not always exactly (1,0) or (0,1), but rather a value closer to (1,0) or (0,1), such as (0.9,0.05) or (0.1,0.99). In this case, the softmax algorithm uses the following formula to convert the output (a,b) into an accuracy score:

[0124]

[0125] With the development of the field of pattern recognition, a variety of novel transformation algorithms have been proposed in recent years, such as Sparse-softmax, log-softmax, Taylor softmax, log-Taylor softmax, soft-marginsoftmax, SM-Taylor softmax, etc. Using novel transformation algorithms does not change the essence of this invention.

[0126] For the nucleic acid sample to be tested, the method described above for converting the fitting results of the classifier into quality scores is used to predict the quality score of each polymer based on the sequencing signal characteristics of the nucleic acid to be tested.

[0127] In some implementations, the quality score refers to a numerical value characterizing sequencing accuracy, selected from base detection accuracy, base detection error rate, Phred value, etc. It is understood that the quality score is merely a representation of base sequencing accuracy; the representation itself is not important and does not affect the essence of the invention. The quality score can also be directly expressed as accuracy.

[0128] In some implementations, the quality scores are logarithmically based on the base detection error probability, and the plurality of quality scores include Q10, Q15, Q20, Q25, Q30, Q35, Q40, Q45, Q50, Q55, and Q60.

[0129] In some preferred embodiments, the quantification scheme for training calibration can be completed in advance, and the trained classifier can be stored in the system as a configuration file, which can be retrieved when performing quality scoring of the nucleic acid sequence to be tested.

[0130] In some preferred embodiments, standard nucleic acid samples and test nucleic acid samples can carry different molecular markers and be mixed together for simultaneous sequencing. After sequencing, the two types of samples are first separated using molecular markers to complete the training and calibration quantification scheme, resulting in a trained classifier, which is then applied to the test nucleic acid samples.

[0131] The present invention also discloses a method according to any of the foregoing embodiments, wherein a reaction solution containing a nucleotide substrate molecule is used for sequencing. The nucleotide substrate molecule refers to any one, two, or three of A, G, C, and T nucleotide substrate molecules; or any one, two, or three of A, G, C, and U nucleotide substrate molecules.

[0132] This document discloses a sequencing method using fluorophoretically labeled nucleotide substrate molecules according to any of the foregoing embodiments, wherein each sequencing run uses a set of reaction solutions, each set comprising at least two reaction solutions, each reaction solution containing at least one of A, G, C, and T nucleotide substrate molecules, or each reaction solution containing at least one of A, G, C, and U nucleotide substrate molecules. In one aspect, the method includes immobilizing the nucleotide sequence fragment to be tested, introducing one reaction solution from the set, and recording fluorescence information. In another aspect, the method includes introducing one reaction solution at a time, followed by the introduction of another reaction solution from the same set. In yet another aspect, the reaction solution set contains at least one reaction solution containing two or three nucleotide molecules.

[0133] This paper discloses a sequencing method using nucleotide substrate molecules labeled with fluorescent markers. Each sequencing run uses a set of reaction solutions, each set comprising two portions of the reaction solution, each containing two nucleotides with different bases. On one hand, the nucleotides in one portion of the reaction solution are complementary to two bases on the target nucleotide sequence, and the nucleotides in the other portion are complementary to two other bases on the target nucleic acid sequence. The method includes immobilizing the target nucleotide sequence fragment and introducing a first portion of the reaction solution into the set. Then, a second portion of the reaction solution from the same set is added. The two portions of the reaction solution can be added alternately to obtain the coding information of the target nucleotide substrate through fluorescence information. When the two nucleotides in each reaction solution are labeled with different fluorescent markers (i.e., two-color 2+2 sequencing), two orthogonal two-color 2+2 sequencing runs are performed; when the two nucleotides in each reaction solution have the same fluorescent marker (i.e., single-color 2+2 sequencing), preferably, three orthogonal single-color 2+2 sequencing runs are performed. Optionally, one of the reaction solutions in the sequencing reaction may contain three nucleotides, while the other reaction solution may contain a different nucleotide. In this case, four orthogonal sequencing reactions (1+3 sequencing) can be performed alternately.

[0134] According to a preferred embodiment, the method of the present invention further includes performing bioinformatics analysis on the measured nucleic acid sequence based on the quality score of the polymer.

[0135] According to a preferred embodiment, the bioinformatics analysis refers to performing two or more orthogonal degenerate sequencing cycles to obtain the quality value of each degenerate polymer length, and then using this quality value for error correction code decoding (or error correction). For example, if the nucleic acid sequence to be tested is subjected to three orthogonal degenerate sequencing cycles (MK, RY, WS), the resulting nucleic acid sequence is ACCGTTTGC. Taking polymer CC as an example, in the MK sequencing cycle, its binary polymer is ACC; in the RY sequencing cycle, its degenerate polymer is CC; and in the WS sequencing cycle, its degenerate polymer is CCG. The characteristics of polymer CC in each degenerate polymer can be calculated, such as the degenerate polymer length, which are 3, 2, and 3 respectively, and its quality score can be calculated for each. Error correction (or error correction code decoding) is then performed based on the quality score to determine the final nucleotide sequence.

[0136] Taking 2+2 degenerate sequencing as an example to illustrate the ECC decoding principle, ECC sequencing consists of three degenerate sequencing passes: MK, RY, and WS. Each pass yields a degenerate sequence containing half the information of the DNA to be tested. By taking the intersection of three degenerate bases at the same position in the three degenerate sequences, the accurate base composition of the DNA to be tested can be obtained. There are a total of 8 possible intersection cases:

[0137]

[0138] Of these eight intersection scenarios, four are valid, yielding four different base pairs. The remaining four are invalid, resulting in an empty intersection. Ideally, without sequencing errors, the intersection of three degenerate sequences should all be valid, yielding the sequence of the DNA to be tested. However, when the DPL calculated by the signal processing algorithm contains errors, invalid scenarios occur near the sequencing errors. Therefore, invalid scenarios in the intersection of degenerate sequences indicate the presence of sequencing errors. To correct these errors, the probabilities of different sequencing error patterns were first statistically determined using sequencing data from standard samples. Then, based on the maximum likelihood principle, the ECC decoding algorithm attempted to correct the DPL calculated by the signal processing algorithm, achieving two objectives: 1. The corrected DPL ensures that the intersection of the three degenerate sequences is entirely valid; 2. Under the constraint of the first objective, this correction method has the highest probability of occurrence according to the probabilities of different sequencing error patterns. An optimization method can be used to achieve these two objectives.

[0139] In some embodiments, the bioinformatics analysis refers to screening high-quality nucleic acid sequences based on assigned quality values. Screening methods include, but are not limited to, screening nucleic acid sequences where all quality values ​​are above or below a certain threshold, screening nucleic acid sequences where the average quality value is above or below a certain threshold, screening regions within nucleic acid sequences where all quality values ​​are above or below a certain threshold, and screening regions within nucleic acid sequences where the average quality value is above or below a certain threshold, etc. The high-quality nucleic acid sequences obtained after screening can be used to detect gene variations, detect gene expression levels, detect RNA alternative splicing status, detect gene modification status, identify the species or individual from which the nucleic acid originates, detect the three-dimensional structure of the genome, detect interactions between nucleic acids, detect interactions between nucleic acids and proteins, detect chromatin accessibility, and analyze RNA structure, etc.

[0140] In some implementations, the bioinformatics analysis refers to aligning a nucleic acid sequence to a reference sequence based on an assigned quality value. Alignment is a common concept in bioinformatics and can be performed using software or algorithms such as BWA, Smith-Waterman algorithm, Bowtie, SOAP, Needleman-Wunch algorithm, Bowtie2, BLAST, ELAND, TMAP, MAQ, minimap2, and SHRiMP. Methods for utilizing quality values ​​in alignment include, but are not limited to:

[0141] 1. Select a subsequence with high-quality values ​​as the seed for the initial positioning sequence;

[0142] 2. When there are multiple possible alignment methods, polymers / bases with lower quality values ​​are given priority as the mismatched parts in the alignment.

[0143] In some embodiments, the bioinformatics analysis refers to identifying gene variations based on alignment results and quality values ​​assigned to the aligned sequences. Gene variation is a conventional concept in biology, including but not limited to single nucleotide polymorphisms, copy number variations, epigenetic variations, and large-scale structural variations. Methods using quality values ​​in identifying gene variations include, but are not limited to:

[0144] 1. For the genomic site of the variant to be identified, among all the bases in the sequence aligned to that site, select the bases with higher quality values ​​of the polymer / base for identification.

[0145] 2. State the null hypothesis: There is no gene variation at this locus. Based on the quality value and the alignment results, calculate the probability that the null hypothesis is true. If the probability is greater than the given significance level, accept the null hypothesis; otherwise, reject the null hypothesis and consider that there is a gene variation at this locus.

[0146] In some implementations, when identifying gene variations, certain features of the comparison results can be used to remove potential false positives or false negatives. These are standard practices in bioinformatics, and their addition does not affect the essence of the present invention. Such features include, but are not limited to:

[0147] 1. This gene variation is concentrated in sequences aligned forward or backward, and less frequently appears in sequences aligned backward or forward.

[0148] 2. This gene variation is concentrated at both ends of the sequence, and less frequently occurs in the center of the sequence;

[0149] 3. When using pair-end sequencing, read1 detects that the locus is mainly G changing to T, while read2 detects that the locus is mainly C changing to A, or read1 detects that the locus is mainly C changing to T, while read2 detects that the locus is mainly G changing to A;

[0150] 4. Other different gene variations frequently appear near this gene variation.

[0151] In some implementations, bioinformatics analysis refers to assembling nucleic acid sequences into longer nucleic acid sequences based on assigned quality values.

[0152] According to a preferred embodiment, the classifier can be trained using a probability distribution model based on maximum likelihood. A probability distribution refers to a probability distribution with a unimodal shape, including but not limited to two-point distributions, binomial distributions, negative binomial distributions, Poisson distributions, geometric distributions, exponential distributions, normal distributions, Γ distributions, chi-square distributions, t distributions, F distributions, β distributions, log-normal distributions, and higher-dimensional extensions of the above distributions. In the aforementioned probability distribution model, the expected value or peak value of the probability distribution is related to the sequencing signal characteristics of the polymer, while the variance or skewness is related to the quality value of the polymer. As is known from basic statistics, different probability distributions can be completely determined by a set of parameters; for example, a normal distribution can be completely determined by the mean and standard deviation. After aligning the sequenced data to a reference genome, the correspondence between the sequencing signal characteristics of each polymer and the polymer length (denoted as n) can be determined based on the alignment results. After being completely determined by the given parameters, the integral area of ​​the probability distribution between n-0.5 and n+0.5 represents the probability that the sequencing signal characteristics correspond to the polymer length n. The likelihood function refers to the probability calculated by multiplying the probabilities of a set of polymer sequencing signal features corresponding to a polymer length n after determining a probability distribution with given parameters. Maximum likelihood refers to finding a set of parameters that maximizes the likelihood function obtained after determining the probability distribution with these parameters.

[0153] The maximum likelihood-based probability distribution model can divide the features of multi-polymer sequencing signals into several groups, and the maximum likelihood-based probability distribution model is applied to each group separately.

[0154] The likelihood function can be mathematically transformed for the sake of computational simplicity. For example, taking the logarithm can transform the multiplication of probabilities into the addition of probabilities.

[0155] The present invention also provides a method for quality assessment of nucleic acid sequencing data, characterized in that it includes: performing fuzzy sequencing or deletion sequencing on the nucleic acid sample to be tested to obtain input data, generating degenerate polymer length information of the input data, and calculating sequencing signal characteristics of the degenerate polymer length;

[0156] The quality score of the polymer is predicted using a quantization scheme for training calibration and based on the sequencing signal features.

[0157] The quantization scheme for training calibration includes:

[0158] Nucleic acid sequences are obtained by sequencing standard nucleic acid samples. The sequencing signal characteristics of the nucleic acid sequences are calculated. The nucleic acid sequences are aligned to reference sequences, and the polymers of the nucleic acid sequences are marked as correctly or incorrectly sequenced. A classifier is trained to fit the relationship between the sequencing signal characteristics of the polymers and their markings.

[0159] According to a preferred embodiment, the sequencing signal characteristics of degenerate polymer length include: the length of the polymer in which the base is located, i.e., the number of bases in the degenerate polymer in which the base is located; generally, a shorter polymer length results in a higher sequencing quality value; the position of the base in the polymer in which it is located, i.e., the distance between the base and the nearest end of the degenerate polymer in which it is located, etc.

[0160] Specifically, a classifier is trained to fit the relationship between the sequencing signal features of each polymer and its label. The classifier can divide polymers into several classes based on their sequencing signal features, and the accuracy of each class is calculated. For example, polymers with lengths of 1, 2, 3, 4, 5, and above can be grouped into separate classes. When using multiple sequencing signal features, orthogonal partitioning can be performed.

[0161] In a preferred embodiment, after fitting, the classifier's fitting result is converted into a quality score. Numerous studies report on how to convert classifier predictions into quality values.

[0162] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for quality assessment of nucleic acid sequencing data, characterized in that, include: The input data is obtained by performing fuzzy sequencing or deletion sequencing on the nucleic acid sample to be tested, generating degenerate polymer length information of the input data, and calculating the sequencing signal characteristics of the degenerate polymer length; The quality score of the polymer is predicted using a quantization scheme for training calibration and based on the sequencing signal features. Based on the quality score of the polymer, bioinformatics analysis was performed on the nucleic acid sequence to be tested. The quantization scheme for training calibration includes: Nucleic acid sequences are obtained by sequencing standard nucleic acid samples. The sequencing signal characteristics of the nucleic acid sequences are calculated. The nucleic acid sequences are aligned to reference sequences, and the polymers of the nucleic acid sequences are marked as correctly or incorrectly sequenced. A classifier is trained to fit the relationship between the sequencing signal characteristics of the polymers and their markings. The nucleic acid sequencing mentioned refers to sequencing methods with open ends; The training classifier, based on a maximum likelihood probability distribution model, includes classifying polymers into several classes according to their sequencing signal characteristics and calculating the sequencing accuracy of each class of polymers; wherein, the bioinformatics analysis includes performing at least two orthogonal degenerate sequencing operations, obtaining a quality value of the degenerate polymer length, and then using the quality value for error correction.

2. The method according to claim 1, characterized in that, The sequencing signal characteristics of the polymer refer to the characteristics of the signal generated when the polymer undergoes a sequencing chemical reaction during the sequencing process. These characteristics include the types of bases that make up the polymer, the length of the polymer, the number of rounds of the sequencing chemical reaction, the signal intensity, the degree to which the signal intensity is close to an integer, the parameters of the sequencing signal, and the degree of phase loss when the polymer is detected.

3. The method according to claim 1 or 2, characterized in that, The polymers include homopolymers, binary copolymers, and terpolymers.

4. The method according to claim 3, characterized in that, The classifiers include, but are not limited to, linear regression, multinomial regression, logistic regression, support vector machine, artificial neural network, random forest, Phred algorithm, and ensemble learning.

5. The method according to claim 1 or 4, characterized in that, The standard nucleic acid sequence is the sequence obtained by sequencing a standard nucleic acid sample; the standard nucleic acid sample refers to a nucleic acid sample whose source and sequence have been determined and which is highly homozygous at almost all sites in the genome, including λ phage DNA, Escherichia coli DNA, and Saccharomyces cerevisiae DNA.

6. The method according to claim 1, characterized in that, The bioinformatics analysis includes identifying gene variations based on the alignment results and the quality values ​​assigned to the aligned sequences.