Methods and systems for enhancing nucleic acid sequencing quality in high throughput sequencing processes using machine learning

By adopting a system with a hierarchical network structure in the next-generation sequencing technology, the basic network and sequence filters are used to improve the sequencing quality indicator, the problem of low base interpretation accuracy in the next-generation sequencing technology is solved, and higher sequencing accuracy and efficiency are achieved.

CN119948569APending Publication Date: 2025-05-06SHANGHAI XINXIANG BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280097919.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-07-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Next-generation sequencing technology generates a large amount of data for base interpretation, but the prior art is difficult to process these data effectively, resulting in low sequencing accuracy, especially in the presence of fluorescence signal quality attenuation and crosstalk.

Method used

A system with a hierarchical network structure is used to generate sequencing quality indicators through the basic network, and filter based on these indicators using sequence filters to obtain higher quality sequence groups to improve the accuracy of base interpretation.

Benefits of technology

By improving the quality of sequencing data, the system can significantly improve the accuracy of base interpretation and overall sequencing quality, reduce error rates and improve processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948569A_ABST
    Figure CN119948569A_ABST
Patent Text Reader

Abstract

A filter structure is provided. The filtering structure uses a hierarchical network structure comprising one or more network frames for obtaining a high quality sequence. Each network frame includes a base network (450) module and a sequence filter (460). The base network (450) generates one or more sequencing quality indicators. The sequencing quality indicator may represent a quality of accuracy of a base interpretation for each base in the sequence, an individual quality of one or more sequences, or an overall quality of a set of sequences. A sequence filter (460) generates filtering results based on the various filtering policies based on the one or more sequencing quality indicators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to nucleic acid sequencing, and more particularly to systems, devices, and methods for utilizing machine learning to enhance the quality of base call results obtained by high-throughput sequencing processes. Background Art

[0002] Sequencing while synthesis is a method for identifying the sequence of a fragment (also referred to as a chain) of a nucleic acid (e.g., DNA) molecule. Sanger sequencing is the first generation sequencing technology using sequencing while synthesis method. Historically, Sanger sequencing has higher accuracy, but sequencing throughput is lower. The second generation sequencing technology (also referred to as next generation sequencing or NGS technology) increases the throughput of synthesis process in a large number by parallelizing many reactions similar to Sanger sequencing. The third generation sequencing technology allows the direct sequencing of a single nucleic acid molecule. In any sequencing technology, base interpretation is a necessary process, and the order of the nucleotide bases in the template chain is inferred by the base interpretation during sequencing readout or after sequencing readout.

[0003] When performing nucleic acid sequencing (e.g., DNA sequencing), fluorescently labeled dNTPs and polymerases are often used. Under the reaction of the polymerase, the dNTPs complement the template strand according to the principle of base complementation to form a new strand. The added fluorescent dye is excited by absorbing the light energy from the laser. The fluorescent signal is collected and analyzed to predict the nucleic acid sequence. Summary of the invention

[0004] Next generation sequencing technology (and other future generation sequencing technology) increases the throughput of synthesis process in a large amount, and therefore generates a large amount of data for base calling. Processing large amounts of data is still challenging. For example, they may be time-consuming, computationally complex and require a large amount of computing resources. In addition, current data processing technology may not provide satisfactory base calling accuracy under various conditions. For example, in the NGS process, the fluorescence signal quality decays over time, which may have a negative impact on the accuracy of the data processing results. In addition, during the sequencing process, there may be crosstalk between different fluorescence signal channels, and there may be synchronization loss (also referred to as cluster phasing and pre-phasing) in the cluster molecules. The synchronization loss in the cluster molecules is caused by the random nature of the chemical reaction and other factors, wherein some molecules may not be able to incorporate the nucleotides of the label, and some other molecules may incorporate more than one nucleotide. This causes the signal intensity leakage between the cycles. Crosstalk and synchronization loss in turn lead to the difficulty of predicting nucleotide bases.

[0005] Recently, machine learning models for base calling have been developed. Machine learning techniques provide a method for self-learning by a computing device. Some existing machine learning models use a combination of, for example, convolutional neural networks (CNNs) and recurrent neural networks (RNNs) networks. CNNs are configured to perform image analysis to detect fluorescent signal clusters, and RNNs are configured to process sequence data. Other machine learning-based models have also been used.

[0006] Compared with traditional base calling methods, base calling methods based on deep learning can provide matching or improved performance while achieving a high-throughput base calling process. During the base calling process, it is desirable to achieve high sequencing quality, such as higher overall sequencing accuracy. Existing techniques for improving sequencing quality can filter sequences based on purity filtering techniques, which calculate the intensity ratio between the DNA signal channel with the highest intensity and the channel with the second highest intensity.

[0007] The present disclosure provides an improved filtering technology, which can provide higher sequencing accuracy for processing high-throughput sequencing data. The filtering structure uses a hierarchical network structure, which includes one or more network frames for obtaining high-quality sequences. Each network frame includes a basic network module (or simply referred to as a basic network) and a sequence filter. The basic network generates one or more sequencing quality indicators. The sequencing quality indicator can represent the quality of the accuracy of the base calls of each base in the sequence, the individual quality of one or more sequences, or the overall quality of the sequence group. The sequence filter generates filtering results based on various filtering strategies based on one or more sequencing quality indicators.

[0008] Embodiments of the present invention provide a computer-implemented method for enhancing the quality of base call results obtained by a high-throughput nucleic acid molecule sequencing process. The method includes obtaining input data including a first nucleic acid sequence group; determining one or more sequencing quality indicators and / or base call predictions based on the input data and one or more neural network-based models. The method further includes filtering the first sequence group using a sequence filter based on one or more of the multiple sequencing quality indicators. The method further includes obtaining a second sequence group based on the filtering results. The second sequence group has a higher data quality than the first sequence group. The method further includes using the second sequence group to provide base call predictions by at least one of the one or more neural network-based models.

[0009] Embodiments of the present invention further provide a system for enhancing the quality of base call results obtained by a high-throughput nucleic acid molecule sequencing process. The system includes one or more processors of at least one computing device; and a memory storing one or more instructions, which, when executed by one or more processors, causes one or more processors to perform steps, the steps including obtaining input data for base calling; and determining one or more sequencing quality indicators and a first nucleic acid sequence group based on the input data and one or more neural network-based models trained for base calling. The instructions further cause one or more processors to perform steps, the steps including filtering the first sequence group using a sequence filter based on one or more of the multiple sequencing quality indicators; and obtaining a second sequence group based on the filtering results. The second sequence group has higher data quality than the first sequence group. The instructions further cause one or more processors to perform steps, the steps including providing base call predictions using the second sequence group by at least one of the one or more neural network-based models.

[0010] Embodiments of the present invention further provide a non-transitory computer-readable medium including a memory storing one or more instructions that, when executed by one or more processors of at least one computing device, cause the at least one computing device to perform a method for enhancing the quality of base call results obtained by a high-throughput nucleic acid molecule sequencing process. The method includes obtaining input data including a first nucleic acid sequence group; determining one or more sequencing quality indicators and / or base call predictions based on the input data and one or more neural network-based models. The method further includes filtering the first sequence group using a sequence filter based on one or more of the plurality of sequencing quality indicators. The method further includes obtaining a second sequence group based on the filtering result. The second sequence group has a higher data quality than the first sequence group. The method further includes providing base call predictions using the second sequence group by at least one of the one or more neural network-based models.

[0011] These and other embodiments are described more fully below. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 An exemplary next generation sequencing (NGS) system according to an embodiment of the present invention is shown;

[0013] Figure 2 An exemplary sequencing-by-synthesis process using an NGS system according to an embodiment of the present invention is shown;

[0014] Figure 3 is a flow chart showing a method for enhancing the quality of base calling results obtained by a high-throughput nucleic acid molecule sequencing process according to an embodiment of the present invention;

[0015] Figure 4 is a block diagram illustrating a hierarchical processing network structure for base calling using one or more network blocks according to an embodiment of the present invention;

[0016] Figure 5A is a flow chart showing a method for performing base call quality filtering (BCQF) according to an embodiment of the present invention;

[0017] Figure 5B is a flow chart showing a method for performing sequence quality filtering according to an embodiment of the present invention;

[0018] Figure 6 is a flow chart illustrating a method for training data set label filtering according to an embodiment of the present invention;

[0019] Figure 7 is a flow chart illustrating a method for cleaning a training data set according to an embodiment of the present invention;

[0020] Fig. 8A is a block diagram showing a hierarchical processing network structure for base calling using one network frame according to an embodiment of the present invention;

[0021] Figure 8B is a block diagram showing a hierarchical processing network structure for base calling using two network frames according to another embodiment of the present invention;

[0022] Figure 8C is a block diagram showing a hierarchical processing network structure for base calling using one network frame according to another embodiment of the present invention;

[0023] Fig.8D is a block diagram showing a hierarchical processing network structure for base calling using one network frame according to another embodiment of the present invention;

[0024] Fig. 8E is a block diagram showing a hierarchical processing network structure for base calling using two network frames according to another embodiment of the present invention;

[0025] FIG. 9A to FIG. 9G is a block diagram illustrating a basic network according to different embodiments of the present invention;

[0026] Fig. 10A is a block diagram illustrating a backbone network model using layers of a one-dimensional (1D) CNN according to one embodiment of the present invention;

[0027] Fig. 10B is a block diagram illustrating a backbone network model of layers using a transformer encoder according to one embodiment of the present invention; and

[0028] Fig.11 A block diagram of an exemplary computing device is shown that may incorporate embodiments of the present invention.

[0029] Although embodiments of the present invention are described with reference to the above drawings, the drawings are intended to be illustrative and other embodiments are consistent with the spirit of the invention and are within the scope of the invention. DETAILED DESCRIPTION

[0030] Various embodiments will now be described more fully below with reference to the accompanying drawings forming a part of this document, and specific examples of practical embodiments are shown by way of illustration. However, this specification may be embodied in many different forms and should not be construed as being limited to the embodiments illustrated herein; on the contrary, these embodiments are provided so that this specification will be thorough and complete, and the scope of the invention will be fully communicated to those skilled in the art. In addition, this specification may be embodied as a method or apparatus. Therefore, any of the various embodiments herein may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Therefore, the following description should not be understood in a limiting sense.

[0031] As described above, next generation sequencing technology (and other future generation sequencing technologies) greatly increases the throughput of the synthesis process, and therefore generates a large amount of data for base calling. Accurate analysis and processing of large amounts of fluorescent signals remains challenging, at least in part because the quality of the fluorescent signal decays over time. For example, when predicting DNA bases in a sequence, it may be difficult to mitigate or eliminate crosstalk and / or cluster phasing.

[0032] Machine learning provides techniques that allow computing devices to make predictions without being explicitly programmed, and enables computing devices to study the characteristics of data. Traditional machine learning systems are developed so that users need to manually design features and select classifiers. With the rapid development of deep learning, the emergence of deep neural networks makes end-to-end learning possible, thereby reducing or eliminating the workload of manual programming and feature selection by users. Deep learning technology is developing rapidly. Convolutional neural networks (ConvNet / CNN) are deep learning algorithms often used for image analysis. Recurrent neural networks (RNN) are often used to process sequence data. Due to their wide applicability and enhanced predictive power, CNN and RNN have great potential in bioinformatics research.

[0033] In DNA sequencing, statistical models such as AYB (All Your Base) models have been proposed to produce more accurate base calling results. Recently, neural network-based base calling models have been proposed, such as RNN-based models, CNN-based models, and models based on the combination of transformers and CNNs. Compared with traditional base calling methods, base calling methods based on deep learning can provide matching or improved performance while realizing high-throughput base calling processes. During the base calling process, it is desirable to achieve high sequencing quality, such as high overall sequencing accuracy. DNA sequencing accuracy can be improved from several aspects, including, for example, improving the efficiency of biochemical reagents, improving the quality of optical systems, and / or improving base calling algorithms.

[0034] The present disclosure provides methods and systems for improving base calling algorithms by filtering an initial sequence group to obtain a high-quality sequence group. The prior art for filtering sequences to exclude low-quality sequencing data uses purity filtering, which calculates the intensity ratio between the DNA signal channel with the highest intensity and the channel with the second highest intensity. The intensity ratio represents the clustering quality. The purity filtering method requires correction of data crosstalk and cluster phasing. Therefore, the quality of crosstalk and cluster phasing corrections limits the filtering performance. In addition, there are usually limitations in using a single deep learning network to obtain higher quality sequencing results. Therefore, a more effective method for improving sequence quality without the limitations described above is desired.

[0035] Embodiments of the present invention are discussed herein. In some embodiments, a hierarchical processing network structure for base calling is provided. The network structure uses one or more network frames for obtaining high-quality sequences. Each network frame includes a base network and a sequence filter. The base network generates one or more sequencing quality indicators, such as sequence quality filtering (SQFN) through the network through index, sequence data set quality index, base call confidence score, etc. The sequencing quality indicator can represent the accuracy of base calling of each base in the sequence, the individual quality of one or more sequences, and the overall quality of the sequence group. The sequence filter generates filtering results based on a preconfigured filtering strategy. This filtering strategy can include base call quality filtering (BCQF) and through index synthesis (e.g., logical AND operation). In some embodiments, the filtering results obtained by a network frame can be provided to another network frame for further processing. For example, multiple network frames can be used in the hierarchical processing network structure to measure the DNA sequence signal quality to exclude low-quality signals that are easily misidentified, thereby improving the final DNA sequencing quality. The details of an embodiment of the present invention are described below. Next-Generation (and Future-Generation) Sequencing Systems

[0036] Figure 1 is a block diagram illustrating an exemplary analysis system 110. Figure 1 As shown in , the analysis system 110 includes an optical subsystem 120, an imaging subsystem 118, a fluid subsystem 112, a control subsystem 114, a sensor 116, and a power subsystem 122. The analysis system 110 can be used to perform a next generation sequencing (NGS) reaction and generate fluorescent images 140 captured during multiple synthesis cycles. These images 140 are provided to a computer 103 for base calling.

[0037] refer to Figure 1 , one or more flow cells 132 are provided to the analysis system 110. A flow cell is a glass slide with fluid channels or lanes in which sequencing reactions occur. In some embodiments, each fluid channel of the flow cell includes an array of cells. Each cell can have several clusters generated on the surface and form a logical unit for imaging and data processing. Figure 2 A flow cell 132 having a plurality of cells is shown, and also shown is an exemplary cell 208. The synthesis process occurs in the flow cell 132, and is described in more detail below.

[0038] refer to Figure 1 , the optical subsystem 120, the imaging subsystem 118, and the sensor 116 are configured to perform various functions, including providing excitation light, directing or guiding the excitation light (e.g., using an optical waveguide), detecting light emitted from the sample due to the excitation light, and converting the photons of the detected light into electrical signals. For example, the optical subsystem 120 includes an excitation optical module and one or more light sources, an optical waveguide, and / or one or more filters. In some embodiments, the excitation optical module and the light source include a laser and / or a light source based on a light emitting diode (LED) that generates and emits the excitation light. The excitation light can have a single wavelength, multiple wavelengths, or a range of wavelengths (e.g., a wavelength between 200nm and 1600nm). For example, if the system 110 has a four-fluorescence channel configuration, the optical subsystem 120 uses four different fluorescences with different wavelengths to excite four different corresponding fluorescent dyes (one for each base A, G, T, C).

[0039] In some embodiments, the excitation optical module may include additional optical components, such as beam shaping optical elements, to form uniform collimated light. The excitation optical module may be optically coupled to the optical waveguide. For example, one or more gratings, mirrors, prisms, diffusers, and other optical coupling devices may be used to direct the excitation light from the excitation optical module to the optical waveguide.

[0040] In some embodiments, the optical waveguide may include three parts or three layers: a first light-conducting layer, a fluid reaction channel, and a second light-conducting layer. The fluid reaction channel may be defined by the first light-conducting layer on one side (e.g., the top side) and by the second light-conducting layer on the other side (e.g., the bottom side). The fluid reaction channel may be used to place a flow cell 132 carrying a biological sample. The fluid reaction channel may be coupled to a fluid conduit, such as in the fluid subsystem 112, to receive and / or exchange liquid reagents. The fluid reaction channel may be further coupled to other fluid conduits to transfer the liquid reagents to the next fluid reaction channel or pump / waste container.

[0041] In some embodiments, fluorescence is delivered to flow cell 132 without the use of optical waveguides. For example, fluorescence can be directed from the excitation optics module to flow cell 132 using free space optical components such as lenses, gratings, mirrors, prisms, diffusers, and other optical coupling devices.

[0042] As described above, the fluid subsystem 112 uses a fluid conduit to deliver reagents to the flow cell 132 directly or through a fluid reaction channel. The fluid subsystem 112 performs reagent exchange or mixing, and places waste generated from the liquid photon system. An embodiment of the fluid subsystem 112 is a microfluidic subsystem that can process a small amount of fluid using channels measured from tens to hundreds of microns. The microfluidic subsystem allows for accelerated PCR processes, reduced reagent consumption, high throughput assays, and integrated pre-PCR assays or post-PCR assays on a chip. In some embodiments, the fluid subsystem 112 may include one or more reagents, one or more multi-port rotary valves, one or more pumps, and one or more waste containers.

[0043] One or more reagents can be sequencing reagents in which sequencing samples are placed. Different reagents can include the same or different chemicals or solutions (e.g., nucleic acid primers) for analyzing different samples. Biological samples that can be analyzed using the system described in the present application include, for example, fluorescent or fluorescently labeled biomolecules, such as nucleic acids, nucleotides, deoxyribonucleic acid (DNA), ribonucleic acid (RNA), peptides or proteins. In certain embodiments, fluorescent or fluorescently labeled biomolecules include fluorescent markers that can emit light in one, two, three or four wavelengths (e.g., emitting red and yellow light) when the biomolecule is provided with excitation light. The emitted light can be further processed (e.g., filtered) before they reach the image sensor.

[0044] refer to Figure 1, the analysis system 110 further includes a control subsystem 114 and a power subsystem 122. The control subsystem 114 can be configured (e.g., via software) to control various aspects of the analysis system 110. For example, the control subsystem 114 can include controlling the optical subsystem 120 (e.g., controlling excitation light generation), the fluid subsystem 112 (e.g., controlling a multi-port rotary valve and a pump), and the power subsystem 122 (e.g., controlling Figure 1 It should be understood that the hardware and software for the operation of the various systems shown in the Figure 1 The various subsystems of the analysis system 110 shown in FIG. are for illustration only. The analysis system 110 may include Figure 1 Furthermore, one or more of the subsystems included in analysis system 110 may be combined, integrated, or divided in any manner desired.

[0045] refer to Figure 1 , the analysis system 110 includes a sensor 116 and an imaging subsystem 118. The sensor 116 detects photons of light emitted from the biological sample and converts the photons into electrical signals. The sensor 116 is also referred to as an image sensor. The image sensor can be a semiconductor-based image sensor (e.g., a silicon-based CMOS sensor) or a charge-coupled device (CCD) image sensor. The semiconductor-based image sensor can be an image sensor based on back side illumination (BSI) or an image sensor based on front side illumination (FSI). In some embodiments, the sensor 116 may include one or more filters to remove scattered light or leakage light while allowing most of the light emitted from the biological sample to pass. Therefore, the filter can improve the signal-to-noise ratio of the image sensor.

[0046] The photons detected by the sensor 116 are processed by the imaging subsystem 118. The imaging subsystem 118 includes a signal processing circuit that is electrically coupled to the sensor 116 to receive the electrical signal generated by the sensor 116. In some embodiments, the signal processing circuit may include one or more charge storage elements, analog signal readout circuits, and digital control circuits. In some embodiments, the charge storage element receives or reads out electrical signals generated in parallel based on substantially all photosensitive elements of the image sensor 116 (e.g., using a global shutter); and transmits the electrical signals to the analog signal readout circuit. The analog signal readout circuit may include, for example, an analog-to-digital converter (ADC) that converts the analog electrical signal into a digital signal.

[0047] In some embodiments, after the signal processing circuit of the imaging subsystem 118 converts the analog electrical signal into a digital signal, it can transmit the digital signal to a data processing system to generate a digital image, such as a fluorescent image 140. For example, the data processing system can perform various digital signal processing (DSP) algorithms (e.g., compression) for high-speed data processing. In some embodiments, at least a portion of the data processing system can be integrated with the signal processing circuit on the same semiconductor die or chip. In some embodiments, at least a portion of the data processing system can be implemented separately from the signal processing circuit (e.g., using a separate DSP chip or cloud computing resources). Therefore, data can be efficiently processed and shared to improve the performance of the sample analysis system 110. It should be understood that the signal processing circuit in the imaging subsystem 118 and at least a portion of the data processing system can be implemented using, for example, a CMOS-based application specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete IC technology, and / or any other desired circuit technology.

[0048] It should be further understood that the power subsystem 122, the optical subsystem 120, the imaging subsystem 118, the sensor 116, the control subsystem 114, and the fluid subsystem 112 can be separate systems or components, or can be integrated with each other. The combination of at least a portion of the optical subsystem 120, the imaging subsystem 118, and the sensor 116 is sometimes referred to as a liquid photonic system.

[0049] refer to Figure 1 , the analysis system 110 provides the fluorescence image 140 and / or other data to the computing device 103 for further processing, including image preprocessing, cluster detection, feature extraction, and base calling. A computer program product 104 for implementing one or more deep learning neural networks 102 and storing them in storage 105 resides on the computing device 103, and those instructions are executable by the processor 106. One or more deep learning neural networks 102 can be used to perform various processes described below. When the processor 106 is executing the instructions of the computer program product 104, the instructions or a portion thereof are typically loaded into the working memory 109, and the processor 106 easily accesses the instructions or a portion thereof from the working memory 109. In one embodiment, the computer program product 104 is stored in the storage 105 or another non-transitory computer-readable medium (which may include media distributed on different devices and different locations). In an alternative embodiment, the storage medium is temporary.

[0050] In one embodiment, processor 106 actually includes multiple processors, which may include additional working memory (additional processors and memory are not shown separately), and the additional working memory includes a graphics processing unit (GPU), which includes at least thousands of arithmetic logic units that support large-scale parallel computing. Other embodiments include one or more special-purpose processing units, which include a systolic array and / or other hardware arrangements that support effective parallel processing. In some embodiments, such special-purpose hardware works in conjunction with the CPU and / or GPU to perform the various processes described herein. In some embodiments, such special-purpose hardware includes application-specific integrated circuits, etc. (which may refer to an application-specific part of an integrated circuit), field programmable gate arrays, etc., or a combination thereof. However, in some embodiments, a processor such as processor 106 may be implemented as one or more general-purpose processors (preferably with multiple cores) without departing from the spirit and scope of the present invention.

[0051] User device 107 includes a display 108 for displaying the results of processing performed by one or more deep learning neural networks 102. In alternative embodiments, a neural network such as neural network 102, or a portion thereof, may be stored in a storage device and executed by one or more processors resident on analysis system 110 and / or user device 107. Such alternatives do not depart from the scope of the present invention. Sequencing by Synthesis

[0052] Figure 2 An exemplary sequencing-by-synthesis process 200 using an analysis system (e.g., system 110) according to an embodiment of the present invention is shown. In step 1 of process 200, the analysis system heats a biological sample to break the two strands of a DNA molecule. One of the single strands will be used as a DNA template strand. Figure 2 Such a DNA template strand 202 is shown, which can be genomic DNA. Template strand 202 can be a strand comprising a nucleotide base sequence (e.g., a long sequence having hundreds or thousands of bases). It should be understood that many such template strands can be generated using polymerase chain reaction (PCR) technology. It should be further understood that other separation and purification processes can also be applied to biological samples to obtain DNA template strands.

[0053] In step 2 of process 200, the analysis system generates a plurality of DNA fragments from a DNA template strand 202. These DNA fragments, such as Figure 2The fragments 204A-D shown in the figure are smaller fragments containing a smaller number of nucleotide bases. Therefore, these DNA fragments can be sequenced in a large number of parallel ways to increase the throughput of the sequencing process. Step 3 of process 200 performs adapter connection. The adapter is an oligonucleotide having a sequence complementary to the priming oligomer placed on the flow cell. The ends of the nucleic acid fragments are connected to the adapter to obtain connected DNA fragments (e.g., 206A-D) so that the subsequent sequencing process can be performed.

[0054] DNA fragmentation and joint connection steps prepare nucleic acid to be sequenced. These prepared, ready-to-sequence samples are referred to as "libraries" because they represent the set of molecules that can be sequenced. After DNA fragmentation and joint connection steps, the analysis system generates a sequencing library representing the set of DNA fragments with joints attached to its ends. In certain embodiments, the prepared library is also quantitatively (and normalized if necessary) so that the optimal concentration of the molecule to be sequenced is loaded into the system. In certain embodiments, other processes can also be carried out in the library preparation process. This process can include size selection, by pcr amplification library and / or targeted enrichment.

[0055] After library preparation, process 200 proceeds to step 4 for clonal amplification to generate clusters of DNA fragment chains (also called template chains). In this step, each DNA fragment is amplified or cloned to generate thousands of identical copies. These copies form clusters so that the fluorescent signal of the clusters in subsequent sequencing reactions is strong enough to be detected by the analysis system. One such amplification process is known as bridge amplification. In a bridge amplification process, cells (e.g., Figure 2 The cell 208 in the library is located, and the primer oligomer is placed on the cell. Each DNA fragment in the library is annealed with the primer oligomer placed on the cell via the adapter attached to the DNA fragment. Then, the complementary strand of the connected DNA fragment is synthesized. The complementary strand folds and anneals with other types of primer oligomers placed on the cell. Therefore, a double-stranded bridge is formed after the complementary strand is synthesized.

[0056] The double-stranded bridge is denatured, forming two single strands attached to the cell. This process of bridge amplification is repeated many times. The double-stranded clone bridge is denatured, the reverse strand is removed, and the forward strand is retained as a cluster for subsequent sequencing. Two such strand clusters are shown as Figure 2Cluster 214 and cluster 216 in the cell. Many clusters with different DNA fragments can be attached to the cell. For example, cluster 214 can be a cluster of connected fragmented DNA 206A placed on cell 208; and cluster 216 can be a cluster of connected fragmented DNA 206B also placed on cell 208. Subsequent sequencing can be performed in parallel with some or all of these different clusters placed on the cell, and in turn, some or all of the clusters are placed on many cells of the flow cell. Therefore, the sequencing process can be massively parallel.

[0057] refer to Figure 2 After clonal amplification in step 4, process 200 proceeds to step 5, where the cluster is sequenced by synthesis (SBS). In the SBS step, nucleotides are incorporated into the complementary DNA strand of the clonal cluster of DNA fragments one base at a time in each synthesis cycle by a DNA polymerase. For example, Figure 2 As shown in , if cycle 1 is the starting cycle, the first complementary nucleotide base is incorporated into the complementary DNA strand of each strand in cluster 214. For simplicity, Figure 2 Only one chain in cluster 214 is shown. But it should be understood that, for some or all other chains of cluster 214, some or all other clusters on community 208, some or all other communities and some or all other flow cells, similar processes can occur. This building-up process is repeated in cycle 2, wherein, the second complementary nucleotide base is incorporated into the complementary DNA strand. Then, in cycles 3, 4, etc., this building-up process is repeated until all bases in template strand 206A are incorporated into complementary nucleotide bases, or until a predetermined number of cycles are reached. Therefore, if template strand 206A has "n" nucleotide bases, then for the entire sequencing process while synthesizing, there can be "n" cycles or a predetermined number of cycles (less than "n"). Complementary strand 207A is at least partially completed after all synthesis cycles. In certain embodiments, this building-up process can be carried out in parallel for some or all chains, clusters, communities and flow cells.

[0058] Step 6 of process 200 is an imaging step that can be performed after step 5 or in parallel with step 5. As an example, the flow cell can be imaged after the sequencing-by-synthesis process of the flow cell is completed. As another example, the flow cell can be imaged while the sequencing-by-synthesis process is being performed on another flow cell, thereby increasing throughput. Figure 2, in each cycle, the analysis system captures one or more images (e.g., images 228A-D) of the cell of the flow cell. The image represents the fluorescent signals of all clusters placed on the cell detected in a specific cycle. In some embodiments, the analysis system can have a four-channel configuration, wherein four different fluorescent dyes are used to identify four nucleotide bases. For example, four fluorescent channels use different types of dyes to generate fluorescent signals with different spectral wavelengths. Different dyes can each be combined with different targets and produce signals with different fluorescent colors or spectra. Examples of different dyes can include dyes based on carboxyfluorescein (FAM) that produce signals with blue fluorescent colors, dyes based on hexachlorofluorescein (HEX) that produce signals with green fluorescent colors, dyes based on 6-carboxyl-X-rhodamine (ROX) that produce signals with red fluorescent colors, and dyes based on tetramethylrhodamine (TAMRA) that produce signals with yellow fluorescent colors.

[0059] In a four-channel configuration, the analysis system captures images of the same cell for each channel. Therefore, for each cell, the analysis system generates four images in each cycle. The imaging process can be performed relative to some or all of the cells and flow cells, and a large number of images are generated in each cycle. These images represent the fluorescent signals of all clusters placed on the cell detected in this particular cycle. The images captured in all cycles can be used for base calling to determine the sequence of the DNA fragment. The sequence of the DNA fragment includes an ordered combination of nucleotide bases with four different types, that is, adenine (A), thymine (T), cytosine (C) and guanine (G). The sequences of multiple DNA fragments can be integrated or combined to generate the sequence of the original genomic DNA chain. The embodiments of the present invention described below can process a large number of images in an effective manner, and base calling is performed using the improved architecture of the deep learning neural network. Therefore, the base calling process according to the embodiment of the present invention has a faster speed and a lower error rate. Although the above description uses DNA as an example, it should be understood that the same or similar process can be used for other nucleic acids, such as RNA and artificial nucleic acids.

[0060] Figure 3300 is a flowchart showing a method 300 for enhancing the quality of base call results obtained by a high-throughput nucleic acid molecule sequencing process according to an embodiment of the present invention. Method 300 can be performed by one or more computing devices (such as device 103). Method 300 can start at step 302, which obtains input data of a hierarchical processing network structure. The input data includes an initial group (e.g., a first group) of nucleic acid sequences (e.g., DNA sequence signals). The initial sequence group can represent an unknown nucleic acid sequence and is sometimes represented by an input embedding vector. The initial sequence group can be obtained based on image preprocessing and cluster detection or by any other desired sequencing detection method. In some embodiments, image preprocessing processes fluorescent signal images captured by an analysis system in multiple synthesis cycles. Before performing subsequent cluster detection steps and base call processes, the image is processed to improve the accuracy of cluster detection, reduce signal interference between close-range clusters, and improve the accuracy of base calls. Image preprocessing can include, for example, light correction, image registration, image normalization, image enhancement, etc.

[0061] Cluster detection uses preprocessed images to detect the center position of fluorescent signal clusters (or simply cluster detection). In some embodiments, cluster detection can use a trained CNN to generate an output feature map, and use a local maximum algorithm to determine the center position of the cluster. The extracted cluster information can be represented in an embedding vector for base calling. The embedding vector represents an unknown nucleic acid sequence. An embodiment of the image preprocessing and cluster detection method is described in more detail in the international application PCT / CN2021 / 141269 entitled "Nucleic Acid Sequencing Method and System Based on Deep Learning" filed on December 24, 2021, and the contents of the application are fully incorporated herein by reference for all purposes.

[0062] Return to reference Figure 3 , step 304 uses one or more neural network-based models to determine one or more sequencing quality indicators and / or one or more base call assertions. In some embodiments, after filtering, the base call assertion is output at step 310. One or more neural network-based models may include a deep learning model based on RNN, a deep learning module based on a transformer, a deep learning model based on one-dimensional convolution, and / or any other desired machine learning or deep learning-based model for base calling. Some models are described in more detail in the international application PCT / CN2021 / 141269. Base call prediction represents the prediction of bases in a sequence by using one or more neural network-based models. The base call result includes a predicted nucleic acid sequence (e.g., a DNA fragment). In some embodiments, the neural network-based model can generate base call predictions for all fluorescent signal clusters captured in the image in "n" cycles in parallel. This greatly reduces the processing time for base calling.

[0063] Using these deep learning models for base calling, in addition to base call predictions, one or more sequencing quality indicators can be determined. The sequencing quality indicators include at least one of a sequence quality filter network (SQFN) pass index, a dataset quality index, and a confidence score. These sequencing quality indicators are described in more detail below.

[0064] Still reference Figure 3 In step 306, the sequence filter filters the first sequence group based on one or more sequencing quality indicators. In step 308, based on the filtering result, a second sequence group is obtained. The second sequence group has a higher data quality than the first sequence group. For example, compared with the first sequence group, the sequences in the second sequence group may generally have a higher base prediction accuracy. In the second sequence group, the number of sequences with higher base call accuracy may be greater than the number in the first sequence group. Therefore, the overall data set quality of the second sequence group may be higher than the first sequence group. Therefore, the second sequence group has a higher data quality than the first sequence group alone or as a group. High-throughput next-generation sequencing technology generally expects higher quality. In some embodiments, using filtering, the quality of the sequence obtained by using a high-throughput sequencing process can match or even exceed the quality of the sequence obtained by using a low-throughput traditional sequencing technology. In some embodiments, step 310 outputs the base call prediction of the second sequence group. The second sequence group has a higher quality than the first sequence group. Therefore, the second sequence group can be used for base calling, and the base call prediction is output.

[0065] Figure 4 4 is a block diagram illustrating a hierarchical processing network structure 400 for base calling using one or more network frames according to an embodiment of the present invention. The network structure 400 receives input data 402. As described above, the input data 402 includes a first nucleic acid sequence (e.g., DNA sequence) group. The first sequence group may represent an unknown nucleic acid sequence, whose bases have not yet been predicted. The first sequence group is sometimes represented by an input embedding vector.

[0066] like Figure 4 As shown in , the network structure 400 includes one or more network boxes, including, for example, a first network box 410 and an optional second network box 420, a third network box (not shown), etc. In general, the network structure can have N network boxes (e.g., up to the Nth network box 430), where N is a number greater than 1. Figure 4Also shown are embodiments of a single network block (e.g., network block 410, 420, or 430). The network block includes a base network 450 and a sequence filter 460. The base network 450 may include one or more network models, such as a first network model 452A and an optional second network model 452B, a third network model (not shown), etc. In general, the base network 450 may include M network models, where M is a number greater than 1. Each of the network models 452A-M included in the base network 450 may include, for example, a deep learning network based on MLP, CNN, RNN, a deep learning network based on a transformer, a one-dimensional CNN, and / or any other desired neural network model that can process sequence data.

[0067] If multiple network models are included in the basic network 450, parallel processing of the input data 402 can be performed, thereby increasing the performance of the basic network 450. In addition, in some embodiments, multiple basic networks can also be used to further improve the quality of the results. The network models used in the basic network 450 can be the same or different. For example, the first network models 452A-M can use the same type of network model (for example, all use RNN-based deep learning network models) or use different types of network models (for example, the first network model 452A uses an RNN-based deep learning network model, the second network model 452B uses a transformer-based deep learning network model, etc.).

[0068] One or more network models 452A-M in the base network 450 can generate one or more sequencing quality indicators as intermediate outputs. Sequencing quality indicators include, for example, SQFN pass index 454, data set quality index 456 and / or confidence score 458. In some embodiments, the base network 450 also provides base call prediction 459, which includes base prediction. Confidence score 458 includes multiple confidence scores associated with base call prediction. The confidence score of the base prediction indicates the degree of certainty of the base call prediction by the base network 450. For example, the confidence score can represent 99%, 90%, 80% and other confidences. Typically, high-quality base signals have high confidence scores (e.g., 95% or higher). In some embodiments, the confidence score can be used as a reference index for measuring the quality of base prediction. Based on the confidence score, it can be determined whether the base prediction is a high-quality prediction (or simply, whether the base is a high-quality base base). For example, for each base prediction, its confidence score can be compared with a confidence threshold. If the confidence score is greater than or equal to the confidence threshold, the base prediction can be classified as a high quality base prediction (or simply, the base can be classified as a high quality base). As described in more detail below, the confidence threshold can be a fixed threshold number or an adaptive threshold number.

[0069] As described in more detail below, based on the number of high-quality bases (or high-quality base predictions) in the sequence, it can be determined whether the sequence is a high-quality sequence. In one example, the number of high-quality bases is compared with the high-quality base threshold number. If the number of high-quality bases is greater than or equal to the high-quality base threshold number, the sequence is classified as a high-quality sequence. Similar to the confidence threshold, the high-quality base threshold number can be a fixed threshold number or an adaptive threshold number. The confidence threshold number and the high-quality base threshold number can be determined based on the data set quality index 456. For example, if the data set quality index is 85%, the threshold number is determined via the corresponding threshold number in the search threshold table. The table can be a fixed table generated according to the experiment, or it can be generated by real-time calculation. The real-time calculation is a variable threshold combination calculation through an exponential distribution table to find a distribution close to the data set quality index. In some embodiments, it can be determined whether the sequence is a high-quality sequence based on the SQFN directly generated by one or more network models in the network model 452A-M through the index 454. As described in more detail below, the SQFN can be used to directly classify the sequence into a high-quality sequence or a low-quality sequence through the index 454.

[0070] Still reference Figure 4 In some embodiments, when determining whether each sequence in a sequence group is a high-quality sequence (by counting the number of high-quality bases in each sequence or by direct classification using SQFN), a data set quality index 456 can be determined. The data set quality index 456 represents the overall sequencing quality of the sequence group. It can be determined based on the number of high-quality sequences. For example, a sequence group may include 100 sequences, of which 99 sequences are determined to be high-quality sequences, and therefore, the overall data set quality can be considered to be 99%.

[0071] like Figure 4 As shown in , base network 450 can also generate base call predictions 459. As described above, base call predictions identify the type of base in the sequence (e.g., A, T, C, G). Base call predictions can be generated by first network block 410, second network block 420, Nth network block 430, or any network block in structure 400. In some embodiments, base call predictions are generated by the last network block (e.g., Nth network block 430). As described below, using the last network block to generate base call predictions can reduce the amount of computational work required because the last network block only receives high-quality sequences provided by the previous network blocks through filtering.

[0072] Still reference Figure 4, one or more of the sequencing quality indicators (e.g., confidence score 458, SQFN pass index 454, data set quality index 456) can be provided to sequence filter 460 to filter out low-quality sequences. Sequence filter 460 can have different structures and perform various filtering methods to obtain high-quality sequences. These structures and filtering methods are described in more detail below.

[0073] In the hierarchical processing network structure 400, the subsequent network frame receives the sequence processed (e.g., filtered) by the previous network frame, and the received sequence can be further processed. As an example, the first network frame 410 processes the sequence received as its input data 402, and obtains a high-quality sequence as its output. These high-quality sequences are passed to the second network frame 420 as its input data. The second network frame 420 can further process the input data and generate its output data. The output data of the second network frame 420 may include the sequence filtered further, so the quality of the sequence can be further improved. In some embodiments, the second network frame 420 may have different filtering thresholds (e.g., a higher threshold of confidence score, a higher threshold of high-quality base number, etc.). As a result, the second network frame 420 can further improve the overall quality of the sequence in its output data. In some embodiments, even if the second network frame 420 has a higher filtering threshold, the sequence included in the output data of the second network frame 420 may also be the same as the sequence in the input data received by the second network frame 420. That is, filtering by the second network block 420 may not remove any further sequences (if the sequences at the input of the second network block 420 are already of sufficiently high quality).

[0074] In a similar manner, the sequence in the output data generated by the second network frame 420 can be passed to the next network frame as its input data, and the filtering process can be repeated. Therefore, the hierarchical processing network structure 400 can be used for multi-stage filtering to gradually identify high-quality sequences. The number of filtering levels (or the number of network frames) can be configured in any desired manner and based on sequencing quality requirements. Therefore, the final output 440 of the hierarchical processing network structure 400 can include a high-quality sequence with accurately identified bases. Therefore, nucleic acid sequencing accuracy can be significantly improved. In addition, as described above, multiple network structures 400 can be implemented, wherein each structure has multiple network frames suitable for running on GPU. As a result, this process can achieve or maintain high throughput, while significantly improving base interpretation accuracy.

[0075] As described above, a sequence filter (eg, filter 460) may perform different filtering methods. Figure 5A5 is a flow chart illustrating a method 500 for performing base call quality filtering (BCQF) according to an embodiment of the present invention. Method 500 can be performed on each sequence based on a confidence score received by a sequence filter (e.g., filter 460). In step 502, the sequence filter evaluates whether the bases included in the sequence are high-quality bases based on a confidence threshold. For example, for each base prediction, its confidence score can be compared with the confidence threshold. If the confidence score is greater than or equal to the confidence threshold, the base prediction can be classified as a high-quality base prediction (or simply, the base is classified as a high-quality base). The sequence filter can classify each base in the sequence (or at least some bases in the sequence) as a high-quality base or a low-quality base (or a non-high-quality base).

[0076] Next, in step 504, the sequence filter can count the high-quality base number in the sequence. In step 506, the sequence filter determines whether the sequence is a high-quality sequence based on the high-quality base number in the sequence. For example, if the high-quality base number is greater than or equal to the high-quality base threshold number, the sequence is classified as a high-quality sequence (step 508). If the sequence is classified as a high-quality sequence, it passes through the sequence filter, and can be provided to the next network frame or included in the output data of the entire hierarchical processing network structure. If the sequence is classified as a low-quality sequence (or is classified as a non-high-quality sequence), the sequence filter continues to evaluate the next sequence (step 510). If there are no more sequences to be evaluated, the method 500 ends. In some embodiments, the method 500 also generates a BCQF index for each classified sequence. For example, if the sequence is classified as a high-quality sequence, the method 500 can generate a BCQF index indicating "pass", and if the sequence is classified as a low-quality sequence (or is classified as a non-high-quality sequence), the method 500 can generate a BCQF index indicating "failure".

[0077] Figure 5B Another filtering method is shown in . Figure 5B is a flow chart showing a method 530 for performing sequence quality filtering according to an embodiment of the present invention. Figure 5BAs shown in , in step 532, the sequence filter (e.g., filter 460) obtains a sequence quality filter network (SQFN) pass index. As described above, the SQFN pass index can be directly generated by one or more neural network-based models of the base network. In step 532, the sequence filter obtains the SQFN pass index of the sequence group. In one embodiment, the SQFN pass index is derived based on the confidence score generated by the base network. For example, if binary classification is used by one or more neural network-based models in the base network, the confidence score can be expressed as (0.9, 0.1), (0.95, 0.05), (0.99, 0.01), (0.1, 0.9), (0.2, 0.8), etc. These confidence scores can be converted to Boolean values, such as "true" or "false", or binary values, such as "1" or "0". In some embodiments, the highest score of a sequence can be used to indicate whether the sequence is a "pass" sequence or a "fail" sequence. A "pass" sequence may be classified as a high-quality sequence, and a "fail" sequence may be classified as a low-quality sequence (or non-high-quality sequence).

[0078] exist Figure 5B In step 534, the sequence filter determines whether the SQFN pass index of the sequence indicates "pass" or "fail". If the SQFN pass index indicates "pass", the sequence filter classifies the sequence as a high-quality sequence (step 536). If the SQFN pass index indicates "fail", the sequence filter classifies the sequence as a low-quality sequence (or non-high-quality sequence) and checks whether there are more sequences to be classified (step 538). If so, the sequence filter repeats steps 534 and 536 to classify the next sequence. If not, the method 530 can end.

[0079] In some embodiments, the SQFN pass index of the sequence group can be used to determine (step 540) a data set quality index. For example, in the previous step, the percentage of sequences in the sequence group that passed is an index representing the overall data set quality. The data set quality index can be used to determine the filter threshold number via a search threshold table. The table can be a fixed table generated based on experiments, or generated by real-time calculations. The real-time calculation calculates a pass index distribution table for the varying threshold combinations to find a distribution that is close to the data set quality index.

[0080] It should be understood that the above description (e.g., BCQF or sequence quality filter) for filtering is an example process. These filtering processes can be changed in any desired manner. Steps and their order can be added, removed and changed while still achieving the same filtering results. As an example, although the above example shows that an index (e.g., BCQF index, SQFN index, data set quality index) corresponds to a binary classification (e.g., pass or fail, true or false, 1 or 0), the index can be configured to have any numerical value or Boolean value for indicating the quality of bases, sequences and / or data sets.

[0081] Figure 6 6 is a flow chart illustrating a method 600 for training data set label filtering according to an embodiment of the present invention. Neural network models (e.g., model 452 in base network 450) are typically trained using labeled training data before they can be used to make base call predictions. Therefore, the quality of the training data can affect the ability of the neural network model to make accurate predictions. Low-quality training data may mislead the fitting process during training. Therefore, filtering out low-quality training data can improve network performance. Processing training data sets to filter out low-quality training data can also use the above methods (e.g., methods 500 and 530).

[0082] In another embodiment, method 600 can be used to exclude low-quality training data from a training data set. Method 600 is also called a label filtering method, which is a sequencing filtering method based on machine learning. Method 600 can be applied to training data and non-training data (e.g., real-world application data). Figure 6, step 622 obtains labeled base call data, and step 624 obtains a sequence determined by a neural network model trained for classification (e.g., model 452 in base network 450). The labeled base call data and the sequence determined by the trained neural network model are cross-checked to remove incorrectly labeled base call data. In particular, using the data obtained in steps 622 and 624, step 626 determines whether the base is correctly classified in the sequence in the labeled base call data. Then, step 628 counts the number of correctly classified bases in the sequence. Step 630 determines whether the number of correctly classified bases in the sequence is greater than or equal to a threshold number. If so, step 632 determines that the sequence passes the filter, indicating that the labeled base call data for the particular sequence in the training data is of high quality. If not, step 633 determines that the sequence does not pass the filter, indicating that the labeled base call data for the particular sequence in the training data is of low quality (or non-high quality). The threshold number may be fixed or adaptive. For each sequence in the training data, steps 626, 628, 630, 632, and 633 above may be repeated. If a particular sequence passes the filter, method 600 proceeds to evaluate the next sequence (step 634). If there are no more sequences to evaluate, method 600 may end. In some embodiments, if one or more sequences do not pass the filter, indicating that the training data includes incorrectly labeled base call data, optional method 700 may be performed to clean up the training data.

[0083] Figure 7 700 is a flow chart illustrating a method for cleaning a training data set according to an embodiment of the present invention. As described above, an original (or uncleaned) training data set can be provided to train one or more neural network models in a base network (e.g., base network 450). The base network generates confidence scores and base call predictions. One or more filtering methods (e.g., methods 500 and 600 described above) can be used to filter sequences in the uncleaned training data set. Sequences that pass the filter are used as a new training data set (or retraining data set) to replace the previous training data set.

[0084] In particular, if Figure 7As shown in , the training data set cleaning method 700 can begin at step 702, which obtains a retraining data set that only includes sequences that have previously passed filtering (e.g., in method 500 or method 600). Step 704 uses the retraining data set to retrain one or more neural network models (e.g., model 452 in base network 450). In step 706, based on the retraining data set (which has been filtered to include only high-quality sequences), one or more neural network models in the base network re-determine confidence scores and base call predictions. In some embodiments, one or more neural network models in the base network do not generate SQFN indices during the retraining process. In step 708, the retraining data set is filtered using the re-determined confidence scores. The filtering process of the retraining data set can be the same or similar to the methods 500 and 600 described above, so it will not be repeated. Step 710 determines whether all sequences in the retraining data set pass the filter. For example, if the filtering process (e.g., method 500 or method 600) determines that all sequences in the retraining data set are not all high-quality sequences, then step 710 determines that not all sequences pass the filter. Method 700 can then repeat steps 702, 704, 706, 708, and 710. As described above, during the filtering process, a threshold number is used to compare with the number of correctly classified bases in the sequence. The threshold number can be fixed or adaptive. For example, the threshold number can be configured to increase step by step as the filtering and cleaning process is repeated. Ultimately, if step 710 determines that all sequences in the retraining data set pass the filter, which indicates that all sequences are considered to be high-quality sequences, then method 700 can continue to the end. High-quality sequences are used as new training data. Therefore, the new training data represents a cleaned training data set that only includes high-quality sequences. Using the new training data, the training of the neural network model can be improved. In turn, the trained neural network model can make more accurate base call predictions.

[0085] Typically, a base network trained using only a high-quality training data set (e.g., a cleaned training data set) can provide a higher prediction accuracy than a network trained with a data set (e.g., an uncleaned training data set) having an original training data set. In some embodiments, for a base network trained using only a cleaned training data set, its confidence score may not produce a strong deviation from data quality. Therefore, in some cases, base call predictions are generated by a base network trained using a cleaned data set, and the confidence score for the BCQF process is generated by a base network trained using an uncleaned training data set. In some embodiments, multiple neural network models are used in the base network so that, for example, a model can be trained using a cleaned training data set, and another model can be trained using an uncleaned training data set. As a result, processing speed and efficiency can be improved. In some embodiments, two independent base networks can be partially integrated to share the same backbone network, but have different decoders to provide different outputs (e.g., one for providing base call predictions, and one for providing confidence scores). Various network structures are described in more detail below.

[0086] Fig. 8A 4 is a block diagram illustrating an embodiment of a hierarchical processing network structure 400 for base calling using one network frame. Fig. 8A As shown in FIG. 4 , the network structure 400 in this embodiment includes a network box 810. The network box 810 receives input data 802 including a nucleic acid sequence. The sequence may be an unknown sequence from which base calls are performed to perform base prediction. The sequence may have a high-quality sequence and / or a low-quality sequence. The network box 810 includes a base network 804 and a sequence filter 807, which are the same or similar to the base network 450 and the sequence filter 460, respectively. Figure 4 . In some embodiments, the base network 804 generates a confidence score 806, and provides the score 806 to the sequence filter 807. The sequence filter 807 performs a base call quality filtering (BCQF) process based on the confidence score 806. The BCQF process is as described above. The sequence filter 807 generates a pass index 809 representing the quality of one or more sequences in the input data 802 based on the BCQF result. Therefore, the sequence filter 807 can filter the sequence in the input data 802 based on the BCQF pass index 809. The method described above (e.g., method 500) can be used for filtering. After filtering, a high-quality sequence is obtained, and the base network 804 can make a base call prediction using the high-quality sequence. The base call prediction and / or the BCQF pass index 809 can be provided as the output 803 of the network box 810.

[0087] Figure 8B8 is a block diagram illustrating another embodiment of a hierarchical processing network structure for base calling using two network blocks 820 and 830. Figure 8B , the first network box 820 receives input data 822 including a nucleic acid sequence group. The sequence may be an unknown sequence from which base calls are made to perform base prediction. The sequence may have a high-quality sequence and / or a low-quality sequence. The first network box 820 includes a base network 824, but without a sequence filter. The second network box 830 includes a base network 834 and a sequence filter 837. In some embodiments, the base network 824 uses one or more neural network models of the base network 824 to determine a first pass index (e.g., SQFN pass index 826). The first pass index is provided to the base network 834 of the second network box 830. As described above, the SQFN pass index indicates whether a particular sequence is a high-quality sequence. Based on the SQFN pass index, the base network 834 can only process high-quality sequences provided by the first network box 820.

[0088] like Figure 8B As shown in, the base network 834 of the second network block 830 determines the confidence score associated with the sequence received by the second network module 830. The sequence filter 837 of the second network block 830 performs the BCQF process 838 and generates a second pass index (e.g., BCQF pass index 839). Therefore, the sequence filter 837 can filter the sequence received by the second network block to obtain a high-quality sequence. The BCQF process described above can be used to filter. The base network 834 can use the high-quality sequence obtained as a result of the filtering to perform base call prediction. Base call prediction, high-quality sequence and / or BCQF pass index 839 can be provided as the output 833 of the second network block 830.

[0089] Figure 8C 8 is a block diagram illustrating another embodiment of a hierarchical processing network structure 400 for base calling using a network block 840. Figure 8C , network block 840 receives input data 842 including a set of nucleic acid sequences. The sequence may be an unknown sequence from which base calls are made to perform base predictions. The sequence may have high quality sequences and / or low quality sequences. Network block 840 includes a base network 844 and a sequence filter 847, which may be the same or similar to base network 450 and sequence filter 460, respectively, such as Figure 4845. In some embodiments, the base network 844 includes one or more neural network-based models that can generate confidence scores 846 and data set quality indexes 845. The base network 844 provides the score 846 to the sequence filter 847. The base network 844 can also provide the data set quality index 845 to the sequence filter 847 so that the sequence filter 847 can perform an adaptive BCQF process. In the adaptive BCQF process, the high-quality base threshold number used to classify whether the sequence is a high-quality sequence varies based on the data set quality index 845. For example, if the data set quality index has a high value, the threshold number may be reduced, and vice versa. The sequence filter 847 can perform a BCQF process based on the confidence score 846 and the adaptive threshold number. The BCQF method is as described above. The sequence filter 847 generates a BCQF pass index 849 representing the quality of one or more sequences in the input data 842 based on the BCQF result. Therefore, the sequence filter 847 can filter the sequence in the input data 842 based on the BCQF pass index 849. The filtering can be performed using the methods described above (e.g., method 500). After filtering, a high-quality sequence is obtained, and the base network 844 can use the high-quality sequence to make base call predictions. The base call predictions and / or BCQF pass index 849 can be provided as output 843 of the network block 840.

[0090] Fig.8D 8 is a block diagram illustrating another embodiment of a hierarchical processing network structure 400 for base calling using a network block 860. Fig.8D , network block 860 receives input data 862 including a set of nucleic acid sequences. The sequence may be an unknown sequence from which base calls are made to perform base predictions. The sequence may have high quality sequences and / or low quality sequences. Network block 860 includes a base network 864 and a sequence filter 867, which may be the same or similar to base network 450 and sequence filter 460, respectively, such as Figure 4 861. In some embodiments, the base network 864 includes one or more neural network-based models that can generate confidence scores 866, data set quality indices 865, and SQFN pass indices 861. The base network 864 provides the scores 866 and the data set quality indices 865 to the sequence filter 867. The sequence filter 847 can perform an adaptive BCQF process similar to that described above. For example, the sequence filter 867 can perform an adaptive BCQF process based on the confidence scores 866 and the adaptive threshold number set by using the data set quality indices 865. The sequence filter 867 generates a BCQF pass index 869 representing the quality of one or more sequences in the input data 862 based on the BCQF results.

[0091] exist Fig.8D In the embodiment shown in , a SQFN pass index 861 is also generated. The SQFN pass index 861 represents the quality of one or more sequences in the input data 862 as determined by the base network 864. The sequence filter 867 can perform a logical operation (e.g., AND) by using the SQFN pass index 861 and the BCQF pass index 869. For example, in the following table, in the SQFN pass index 861 and the BCQF pass index 869, "true" and "false" or "1" and "0" represent "pass" or "fail", respectively. If the logical operation is an AND operation, the combined pass index generated by the logical operation is "passed" only when the SQFN pass index 861 and the BCQF pass index 869 are both "true" or "1". SQFN Pass Index BCQF Pass Index Combination Pass Index A B Y=AB 0 (false or failed) 0 (false or failed) 0 (false or failed) 0 (false or failed) 1 (true or passed) 0 (false or failed) 1 (true or passed) 0 (false or failed) 0 (false or failed) 1 (true or passed) 1 (true or passed) 1 (true or passed)

[0092] Still reference Fig.8D , the sequence filter 867 can thus filter the sequences in the input data 862 based on the combined pass index. The filtering can be performed using the method described above (e.g., method 500). After filtering, high-quality sequences are retained, and the base network 864 can make base call predictions using only high-quality sequences. The base call predictions and / or the combined pass index can be provided as output 863 of the network block 860.

[0093] Fig. 8E 8 is a block diagram illustrating another embodiment of a hierarchical processing network structure 400 for base calling using two network blocks 880 and 890. Fig. 8E , the first network box 880 receives input data 882 including a nucleic acid sequence group. The sequence may be an unknown sequence from which base calls are made to perform base prediction. The sequence may have a high-quality sequence and / or a low-quality sequence. The first network box 880 includes a base network 884, but without a sequence filter. The second network box 890 includes a base network 894 and a sequence filter 897. In some embodiments, the base network 884 uses one or more neural network-based models of the base network 884 to determine a first pass index (e.g., SQFN pass index 886). The first pass index (e.g., SQFN pass index 886) is provided to the base network 894 of the second network box 890. As described above, the SQFN pass index indicates whether a particular sequence is a high-quality sequence. Based on the SQFN pass index 886, the base network 894 processes the high-quality sequence provided by the first network box 880.

[0094] like Fig. 8EAs shown in, in some embodiments, the base network 884 in the first network box 880 also generates a data set quality index 885 and provides it to the sequence filter 897 of the second network box 890, so that an adaptive BCQF process can be performed. The base network 894 of the second network box 890 determines the confidence score associated with the sequence received by the second network box 890. The sequence filter 897 of the second network box 890 performs a BCQF process 898 and generates a second pass index (e.g., BCQF pass index 899). Therefore, the sequence filter 897 can filter the sequence received by the second network box to obtain a high-quality sequence. The BCQF process described above can be used for filtering. By using the data set quality index 885 to set the base threshold number for classifying the high-quality sequence, the BCQF process can be an adaptive process. The base network 894 can use the high-quality sequence obtained as a result of the filtering to perform base call prediction. Base call prediction, high-quality sequence and / or BCQF pass index 899 can be provided as the output 893 of the second network box 890.

[0095] FIG. 8A to FIG. 8E Various examples of network structure 400 are shown, including various combinations of using one or two network boxes and sequencing quality indicators (e.g., one or more of a dataset quality index, an SQFN pass index, a confidence score). It should be understood that other variations, combinations of boxes, or embodiments of network structure 400 may also be implemented without departing from the principles shown. For example, three or more network boxes may be used, and base call prediction may be performed by the first network box or the last network box.

[0096] FIG. 9A to FIG. 9G is a block diagram illustrating various embodiments of an underlying network according to different embodiments of the present invention. Fig.9A The base network 900 is shown to include a single network model 902 that is used to generate confidence scores 904 and base call predictions 906 . Fig. 9B The base network 910 is shown to include two network models 912 and 914. The first network model 912 is used to generate confidence scores 916, and the second network model 914 is used to make base call predictions 918. As described above, the two network models 912 and 916 can be trained differently using different training data sets (e.g., uncleaned training data and cleaned training data sets) to provide accurate base call predictions and improved confidence scores.

[0097] Fig. 9C Another embodiment of a base network 920 comprising a backbone network model 922 and two decoders 924 and 926 is shown. Fig. 9BThe two separate network models shown in can share a backbone network between the two network models, thereby partially combining the two network models, which can improve the reasoning speed of the network. The backbone network model 922 can use, for example, a one-dimensional CNN-based model, a transformer-based model, an RNN-based model, and / or any other desired neural network structure. Some examples of backbone network models are described in more detail below. Fig. 9C In the basic network 920 shown in FIG. 1 , a first decoder 924 is used to generate a confidence score 928, and a second decoder 926 is used to make a base call prediction 930. The two decoders 924 and 926 are used to generate different outputs, thereby improving computational efficiency. Similar to the above description, the two decoders 924 and 926 can be trained differently using different training data sets (e.g., uncleaned training data and cleaned training data sets), thereby providing accurate base call predictions and improved confidence scores.

[0098] Fig.9D A base network 930 is shown including a single network model 932 that is used to generate a SQFN pass index 934 and a data set quality index 936 . Fig.9E It is shown that the base network 940 includes three network models 942, 944 and 946. The first network model 942 is used to generate the SQFN pass index 943 and the data set quality index 945. The second network model 944 is used to generate the confidence score 947. The third network model 946 is used to make the base call prediction 949. Similar to the above description, the three network models 942, 944 and 946 can be trained differently using different training data sets (e.g., uncleaned training data sets and cleaned training data sets) to provide accurate base call predictions and improved confidence scores. By using multiple network models in the base network, the performance of the network frame can be improved.

[0099] Fig.9F An example is shown in which a base network 950 includes a first network model 952 for generating an SQFN pass index 953 and a data set quality index 955. The base network 950 also includes a backbone network model 954 and two decoders 956 and 958. The first decoder 956 is used to generate a confidence score 957. The second decoder 958 is used to make a base call prediction 959. Similar to the above description, the first network model 952, the first decoder 956 and the second decoder 958 can be trained differently using different training data sets (e.g., uncleaned training data sets and cleaned training data sets), thereby providing accurate base call predictions and improved confidence scores. By using multiple network models in the base network, the performance of the network box can be improved.

[0100] Figure 9G It is shown that the basic network 960 includes a backbone network model 962 shared between three decoders 964, 966 and 968. The backbone network model 922 can use, for example, a one-dimensional CNN-based model, a transformer-based model, an RNN-based model and / or any other desired neural network structure. Some examples of the backbone network model are described in more detail below. The first decoder 964 is used to generate an SQFN pass index 963 and a data set quality index 965. The second decoder 966 is used to generate a confidence score 967. The third decoder 968 is used to make a base call prediction 969. Similar to the above description, one or more of the three decoders 964, 966 and 968 can use different training data sets (e.g., uncleaned training data sets and cleaned training data sets) for different training, thereby providing accurate base call predictions and improved confidence scores. By using multiple decoders that share a backbone network model in the basic network, processing speed and efficiency can be improved.

[0101] As described above, some base network embodiments use a backbone network model and one or more decoders. The backbone network model can be a deep learning network model without the last few layers (e.g., layers for making final classification predictions and generating confidences). These layers can include pooling layers, linear layers, Softmax layers, etc. The decoder can have equivalent functionality to the last few layers. Fig. 10A 1 is a block diagram showing a backbone network model 1000 using a layer of a one-dimensional CNN according to one embodiment of the present invention. Fig. 10A As shown in , model 1000 uses multiple layers, including an input layer 1002 and multiple 1D convolutional layers 1004-1014. The input layer 1002 receives input data including multiple sequences. These sequences can be unknown nucleic acid sequences and include a mixture of high-quality sequences and low-quality sequences. These sequences can be represented by vectors. The sequences can represent clusters of fluorescent signals from multiple cycles. The input layer 1002 can be a linear layer. The linear layer is a feed-forward layer that can learn the offset and correlation rate between the input and output of the linear layer. The linear layer can learn to scale automatically so that it can reduce or expand the dimension of the input vector. In Fig. 10A In one embodiment shown in Fig. 10A As shown in , the input layer 1002 may have 5 channels, the 1D convolution layer 1004 has 16 channels; the 1D convolution layer 1006 has 32 channels, etc. The 1D convolution layer performs convolution operations in one direction, rather than two directions in the 2D convolution layer. For example, the input of the 1D convolution layer 1004 is a 1D vector (e.g., a 1D feature vector representing the signal at the center of the fluorescent signal cluster).

[0102] In some embodiments, each of the one-dimensional convolutional layers 1004-1014 has a kernel for performing a convolution operation. The kernel can have a size of, for example, 4 and a stride of 1. The stride is the number of pixels shifted on the input matrix. Therefore, if the stride is 1, the kernel (or filter) moves 1 pixel at a time. In some embodiments, in order to keep the size of the feature constant, the padding can be configured to be 3, one at the head and two at the tail. Padding refers to the number of pixels added to the image when the image is processed by the kernel of the one-dimensional convolutional layer. As Fig. 10A As shown in , the backbone network model 1000 does not include the last few layers or decoders. The output of the backbone network model 1000 can be provided to multiple decoders or other layers to generate different desired outputs.

[0103] Fig. 10B 1 is a block diagram of a backbone network model 1040 showing the layers of an encoder using a transformer-based neural network according to one embodiment of the present invention. The transformer neural network has an encoder-decoder architecture using one or more attention layers. The transformer neural network can process multiple input sequences or vectors in parallel. Therefore, the processing efficiency and training speed of the network are greatly improved. In addition, the transformer neural network uses one or more multi-head attention layers to better interpret or emphasize important aspects of the input embedding vector. The vanishing gradient problem is also eliminated or significantly reduced by the transformer neural network. Fig. 10B In the embodiment of the present invention, the backbone network model 1040 receives input data 1042, which can be a sequence or a vector. In some embodiments, the input data 1042 can be provided to a position encoding layer (not shown) to take into account the order of the feature vector elements. The position encoding layer includes a position encoder, which is a vector that provides a context based on the position of the elements in the vector. The position encoding layer generates a position-encoded vector, which is provided to the encoder 1020.

[0104] The encoder 1020 may be a self-attention based encoder. The encoder 1020 includes a multi-head attention layer 1026. The multi-head attention layer 1026 determines multiple attention vectors for each element of the position-encoded vector and uses a weighted average to calculate a final attention vector for each element of the position-encoded vector. The final attention vector captures the contextual relationship between the elements of the position-encoded vector. In some embodiments, the encoder 1020 also includes one or more normalization layers 1022 and 1028. The normalization layer controls the gradient scale. In some embodiments, as Fig. 10B As shown in , the normalization layer 1022 is located after the multi-head attention layer 1026. In some embodiments, the normalization layer can be located before the multi-head attention layer. Similarly, it can also be located in the multi-layer perceptron layer (MLP) Before or after 1030. The normalization layer normalizes the input to the next layer, which has the effect of stabilizing the learning process of the network and reducing the number of training iterations required to train the deep learning network. Normalization layers 1022 and 1028 can perform batch normalization and / or layer normalization.

[0105] Fig. 10B The encoder 1020 is also shown to include a multi-layer perceptron (MLP) 1030. The MLP 1030 is a feed-forward neural network. The MLP has a node layer, which includes: an input layer, one or more hidden layers, and an output layer. In addition to the input node, each node in the MLP is a neuron using a non-linear activation function. The MLP 1030 is applied to each normalized attention vector. The MLP 1030 can transform the normalized attention vector into a form acceptable to the next encoder or decoder in the network model 1040. In Fig. 10B In the example shown in , one encoder is used. Therefore, in Fig. 10B , after being normalized by the normalization layer 1028, the output of the MLP 1030 is the encoder output vector 1032. The encoder output vector 1032 is then provided to a decoder in a base network (e.g., base network 920, 950, or 960). In other embodiments, a stacked encoder structure with two encoders may be used. Thus, the output vector from encoder 1020 may also be provided as an input vector to the next encoder.

[0106] In the backbone network model using the transformer network model layer, all attention vectors (or those after normalization) are independent of each other. Therefore, they can be provided to the MLP 1030 in parallel. Therefore, the encoder 1020 can generate encoder output vectors 1032 for all input embedding vectors in the input data 1042 in parallel, thereby significantly improving the processing speed. Some embodiments of the 1D CNN network model and the transformer network model are described in more detail in the international application PCT / CN2021 / 141269. Exemplary Computing Device Embodiments

[0107] Fig.11 is an example block diagram of a computing device 1100 that may incorporate embodiments of the present invention. Fig.11 The machine system that merely illustrates aspects of the technical process described herein is not intended to limit the scope of the claims. Those of ordinary skill in the art will recognize other variations, modifications, and substitutions. In one embodiment, the computing device 1100 typically includes a monitor or graphical user interface 1102, a data processing system 1120, a communication network interface 1112, an input device 1108, an output device 1106, and the like.

[0108] like Fig.11 1, data processing system 1120 may include one or more processors 1104 communicating with a number of peripheral devices via a bus subsystem 1118. These peripheral devices may include input devices 1108, output devices 1106, communication network interfaces 1112, and storage subsystems such as volatile memory 1110 and non-volatile memory 1117. Volatile memory 1110 and / or non-volatile memory 1117 may store computer executable instructions, thus forming logic 1122, which, when applied to and executed by processor 1104, implements embodiments of the processes disclosed herein.

[0109] Input devices 1108 include devices and mechanisms for inputting information into the data processing system 1120. These may include keyboards, keypads, touch screens incorporated into the monitor or graphical user interface 1102, audio input devices (such as voice recognition systems, microphones), and other types of input devices. In various embodiments, input device 1108 may be embodied as a computer mouse, trackball, trackpad, joystick, wireless remote control, drawing tablet, voice command system, eye tracking system, etc. Input device 1108 typically allows a user to select objects, icons, control areas, text, etc. that appear on the monitor or graphical user interface 1102 via commands (such as clicking buttons, etc.). The graphical user interface 1102 may be used in step 1618 of method 1600 to receive user input to make corrections to bases or sequences during the data tagging process.

[0110] Output devices 1106 include devices and mechanisms for outputting information from data processing system 1120. These may include a monitor or graphical user interface 1102, speakers, printers, infrared LEDs, etc., as is well known in the art.

[0111] The communication network interface 1112 provides an interface to a communication network (e.g., communication network 1116) and devices external to the data processing system 1120. The communication network interface 1112 can be used as an interface for receiving data from other systems and transmitting data to other systems. Embodiments of the communication network interface 1112 can include an Ethernet interface, a modem (telephone, satellite, cable, ISDN), a (asynchronous) digital subscriber line (DSL), FireWire, USB, a wireless communication interface (such as Bluetooth or WiFi), a near field communication wireless interface, a cellular interface, etc. The communication network interface 1112 can be coupled to the communication network 1116 via an antenna, a cable, etc. In some embodiments, the communication network interface 1112 can be physically integrated on a circuit board of the data processing system 1120, or in some cases, can be implemented in software or firmware (such as a "soft modem", etc.). The computing device 1100 may include logic that enables communication over a network using protocols (such as HTTP, TCP / IP, RTP / RTSP, IPX, UDP, etc.).

[0112] Volatile memory 1110 and non-volatile memory 1114 are examples of tangible media configured to store computer-readable data and instructions that form logic to implement aspects of the processes described herein. Other types of tangible media include removable memory (e.g., pluggable USB memory devices, mobile device SIM cards), optical storage media (such as CD-ROMs, DVDs), semiconductor memories (such as flash memory, non-temporary read-only memory (ROM), battery-backed volatile memory), networked storage devices, etc. Volatile memory 1110 and non-volatile memory 1114 can be configured to store basic programming and data structures that provide the functions of the disclosed processes and other embodiments thereof that fall within the scope of the present invention. Logic 1122 that implements an embodiment of the present invention can be formed by volatile memory 1110 and / or non-volatile memory 1114 storing computer-readable instructions. The instructions can be read from volatile memory 1110 and / or non-volatile memory 1114 and executed by processor 1104. The volatile memory 1110 and the non-volatile memory 1114 may also provide a repository for storing data used by the logic 1122. The volatile memory 1110 and the non-volatile memory 1114 may include several memories, including a main random access memory (RAM) for storing instructions and data during program execution and a read-only memory (ROM) in which read-only non-transitory instructions are stored. The volatile memory 1110 and the non-volatile memory 1114 may include a file storage subsystem that provides persistent (non-volatile) storage for program and data files. The volatile memory 1110 and the non-volatile memory 1114 may include a removable storage system (such as removable flash memory).

[0113] Bus subsystem 1118 provides a mechanism for enabling the various components and subsystems of data processing system 1120 to communicate with each other as intended. Although communication network interface 1112 is schematically depicted as a single bus, some embodiments of bus subsystem 1118 may utilize multiple different busses.

[0114] It will be apparent to one of ordinary skill in the art that computing device 1100 may be a device such as a smart phone, a desktop computer, a laptop computer, a rack-mounted computer system, a computer server, or a tablet computer device. As is known in the art, computing device 1100 may be implemented as a collection of multiple networked computing devices. In addition, computing device 1100 will typically include operating system logic (not shown), the type and nature of which are known in the art.

[0115] One embodiment of the present invention includes a system, a method, and one or more non-transitory computer-readable storage media that tangibly stores computer program logic that can be executed by a computer processor. The computer program logic can be used to implement embodiments of the processes and methods described herein, including the method 300 for base calling, the method 400 for image preprocessing, the method 800 for cluster detection, the method 1000 for feature extraction, and various deep learning algorithms and processes.

[0116] Those skilled in the art will appreciate that computer system 1100 illustrates only one example of a system in which a computer program product according to an embodiment of the present invention may be implemented. As just one example of an alternative embodiment, the execution of instructions contained in a computer program product according to an embodiment of the present invention may be distributed across multiple computers, such as across computers in a distributed computing network.

[0117] Although the present invention has been specifically described with respect to the illustrated embodiments, it will be understood that various changes, modifications and adjustments may be made based on the present disclosure and are intended to be within the scope of the present invention. Although the present invention has been described in conjunction with what are currently considered to be the most practical and preferred embodiments, it should be understood that the present invention is not limited to the disclosed embodiments, but rather, the present invention is intended to cover various modifications and equivalent arrangements included within the scope of the basic principles of the present invention as described by the various embodiments cited above and below.

Claims

1. A computer-implemented method for enhancing the quality of base calls obtained from a high-throughput nucleic acid molecule sequencing process, characterized in that: The method comprises: obtaining input data including a first nucleic acid sequence group; determining one or more sequencing quality indicators based on the input data and one or more neural network-based models; filtering the first set of sequences using a sequence filter based on one or more of the plurality of sequencing quality indicators; Obtaining a second sequence group based on the filtering result, wherein the second sequence group has higher data quality than the first sequence group; and The second set of sequences is used to provide base call predictions by at least one of the one or more neural network-based models.

2. The method according to claim 1, characterized in that The one or more sequencing quality indicators include at least one of: Sequence Quality Filter Network (SQFN) pass index; Dataset quality index; and Confidence score.

3. The method according to any one of claims 1 and 2, characterized in that The plurality of sequencing quality indicators include confidence scores associated with bases in each sequence included in the first set of sequences, and wherein filtering the first set of sequences based on one or more of the plurality of sequencing quality indicators includes performing base call quality filtering (BCQF), the BCQF including, for each sequence in the first set of sequences: Based on a confidence threshold, evaluating whether the bases included in the sequence are high-quality bases; Counting the number of high-quality bases in the sequence; and Based on the number of high-quality bases in the sequence, determine whether the sequence is a high-quality sequence.

4. The method according to any one of claims 1 to 3, characterized in that The plurality of sequencing quality indicators comprises a sequence quality filter network (SQFN) pass index, and wherein, for each sequence in the first group of sequences, filtering the first group of sequences based on one or more of the plurality of sequencing quality indicators comprises: Based on the corresponding SQFN pass index, it is determined whether the sequence is a high-quality sequence.

5. The method according to claim 4, characterized in that Also included is determining a data set quality index based on the SQFN pass index of the sequences in the first sequence group.

6. The method according to any one of claims 1 to 5, characterized in that The input data comprises a training data set having labeled base call data, and wherein, for each sequence in the first set of sequences, filtering the first set of sequences based on one or more of the plurality of sequencing quality indicators comprises: determining, based on the labeled base call data and the sequence determined by the one or more neural networks, whether a base included in the sequence is correctly classified in the labeled base call data; Counting the number of correctly classified bases in the sequence; and Based on the number of correctly classified bases in the sequence, it is determined whether the sequence passes filtering.

7. The method according to claim 6, characterized in that Also includes: Based on determining that one or more sequences in the first group do not pass the filtering: Obtaining a retraining dataset that includes only sequences that pass the filtering; retraining the one or more neural network models using the retraining data, re-determining confidence scores and base call predictions based on the one or more retrained neural network-based models; as well as The retraining data is filtered based on one or more of the re-determined confidence scores and base call predictions.

8. The method according to any one of claims 1 to 7, characterized in that Filtering the first set of sequences based on one or more of the plurality of sequencing quality indicators comprises: performing base call quality filtering (BCQF) based on confidence scores included in the plurality of sequencing quality indicators; Based on the BCQF result, generating a BCQF pass index representing the quality of one or more sequences in the first sequence group; and Based on the BCQF pass index, the first sequence group is filtered to obtain the second sequence group.

9. The method according to any one of claims 1 to 8, characterized in that The one or more neural network-based models are included in one or more network boxes, the one or more network boxes including a first network box and a second network box; Wherein, determining the one or more sequencing quality indicators comprises: determining a first pass index using one or more neural network models of the first network block, providing the first pass index to one or more neural network based models of the second network block, and determining, based on the first pass index and the one or more neural network-based models of the second network block, confidence scores associated with bases included in the sequence received by the second network block; number.

10. The method according to claim 9, characterized in that Filtering the first sequence group includes: performing base call quality filtering (BCQF) based on the confidence scores associated with bases included in the sequence received by the second network block; generating a second pass index representing the quality of one or more sequences in the first group of sequences based on the BCQF result; and The sequences received by the second network block are filtered based on the second pass index to obtain the second sequence group.

11. The method according to any one of claims 1 to 10, characterized in that The one or more sequencing quality indicators include one or more dataset quality indexes and one or more confidence scores, wherein filtering the first set of sequences based on one or more of the plurality of sequencing quality indicators includes: performing base call quality filtering (BCQF) based on the one or more dataset quality indices and the one or more confidence scores; Based on the BCQF result, generating a BCQF pass index representing the quality of one or more sequences in the first sequence group; and The first sequence group is filtered by an index based on the BCQF to obtain the second sequence group.

12. The method according to any one of claims 1 to 11, characterized in that The one or more sequencing quality indicators include one or more dataset quality indexes, one or more first pass indexes, and one or more confidence scores, and wherein filtering the first set of sequences based on one or more of the plurality of sequencing quality indicators comprises: performing base call quality filtering (BCQF) based on the one or more dataset quality indices and the one or more confidence scores; Based on the BCQF results, a second pass index is generated; combining the first pass index and the second pass index to generate a combined pass index; and The second sequence group is filtered based on the combined pass index.

13. The method according to any one of claims 1 to 12, characterized in that The one or more neural network-based models are included in one or more network boxes, the one or more network boxes including a first network box and a second network box, Wherein, determining the one or more sequencing quality indicators comprises: determining one or more first pass indices and data set quality indices using one or more neural network models of said first network block, providing the first pass index to one or more neural network based models of the second network block, providing the data set quality index to the sequence filter of the second network block; and Based on the one or more neural network-based models of the second network block, confidence scores associated with bases included in the sequence received by the second network block are determined.

14. The method according to claim 13, characterized in that Filtering the first sequence group includes: performing base call quality filtering (BCQF) based on the one or more dataset quality indices and the confidence scores associated with bases included in the sequence received by the second network block; generating, based on the BCQF result, a second pass index representing the quality of the sequence received by the second network block; and The sequences received by the second network block are filtered based on the second pass index to obtain the second sequence group.

15. The method according to any one of claims 1 to 14, characterized in that The one or more neural network based models are included in one or more network boxes.

16. A system for enhancing the quality of base calling results obtained by a high-throughput nucleic acid molecule sequencing process, characterized in that: The system comprises: one or more processors of at least one computing device; and A memory storing one or more instructions, which, when executed by the one or more processors, cause the one or more processors to perform steps, the steps comprising: obtaining input data for base calling; determining one or more sequencing quality indicators and a first set of nucleic acid sequences based on the input data and one or more neural network-based models trained for base calling; filtering the first set of sequences using a sequence filter based on one or more of the plurality of sequencing quality indicators; Obtaining a second sequence group based on the filtering result, wherein the second sequence group has higher data quality than the first sequence group; and The second set of sequences is used to provide base call predictions by at least one of the one or more neural network-based models.

17. The system according to claim 16, characterized in that The one or more sequencing quality indicators include at least one of: Sequence Quality Filter Network (SQFN) pass index; Dataset quality index; and Confidence score.

18. The system according to claim 17, characterized in that The one or more neural network based models are part of a single network block configured to determine the confidence scores and the first sequence group.

19. The system according to claim 17, characterized in that: The one or more neural network-based models are distributed to a first network box and a second network box; The first network block is configured to determine the confidence score; as well as The second network block is configured to determine the first sequence group.

20. The system according to claim 17, characterized in that: At least one of the one or more neural network based models is included in a backbone network frame, The one or more neural network-based models include a first decoder and a second decoder, the first decoder and the second decoder are configured to receive output from the backbone network model, The first decoder is configured to determine one or more confidence scores, and The second decoder is configured to determine the first sequence group.

21. The system according to claim 16, characterized in that The one or more neural network based models are part of a single network block configured to determine one or more SQFN pass indices and one or more data set quality indices.

22. The system according to claim 16, characterized in that: The one or more neural network-based models are distributed in the first network box, the second network box, and the third network box; The first network block is configured to determine a SQFN pass index and a data set quality index; The second network block is configured to determine a confidence score; as well as The third network block is configured to determine the first sequence group.

23. The system according to claim 17, characterized in that: The one or more neural network-based models are distributed in the first network frame and the backbone network frame; The first network block is configured to determine a SQFN pass index and a data set quality index; The one or more neural network based models include a first decoder and a second decoder, the first decoder and the second decoder are configured to receive output from the backbone network block, The first decoder is configured to determine one or more confidence scores, and The second decoder is configured to determine the first sequence group.

24. The system according to claim 17, characterized in that: At least one of the one or more neural network based models is included in a backbone network frame, The one or more neural network-based models include a first decoder, a second decoder, and a third decoder, wherein the first decoder, the second decoder, and the third decoder are configured to receive outputs from the backbone network block, The first decoder is configured to determine the SQFN pass index and the data set quality index, The second decoder is configured to determine one or more confidence scores, and The third decoder is configured to determine the first sequence group.

25. A non-transitory computer-readable medium comprising a memory, characterized in that: The memory stores one or more instructions, which, when executed by one or more processors of at least one computing device, cause the at least one computing device to perform the method according to any one of claims 1-15.