Base calling method and system

By constructing a base recognizer based on a neural network, the problem of base recognition accuracy in multicolor high-throughput sequencing was solved. The spatial and contextual effects of fluorescent clusters were corrected using a cascade model, which improved the accuracy of sequencing results and computational efficiency.

CN119479831BActive Publication Date: 2026-01-09GENEMIND BIOSCIENCES CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411482436.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2026-01-09
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

In existing multicolor or multi-fluorescent channel high-throughput sequencing technologies, the accuracy of base identification is affected by factors such as crosstalk of fluorescence channel signals, phase imbalance, phase imbalance of fluorescent clusters, and the influence of the context sequence of the base to be measured, making it difficult to guarantee the accuracy and reliability of sequencing results, especially in high-density array chips.

Method used

A neural network-based base recognizer is constructed. By fully considering the spatial location, detection time, and contextual information of fluorescent clusters through training data, a cascaded model is used for base recognition to correct spatial crosstalk and contextual effects between physically adjacent amplicon, thereby improving the accuracy of base recognition.

Benefits of technology

It significantly improves the accuracy of sequencing results, making it particularly suitable for applications that require high accuracy, such as pathogen detection and liquid biopsy, and it also has advantages in terms of computational power consumption and computing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479831B_ABST
    Figure CN119479831B_ABST
Patent Text Reader

Abstract

The application provides a base calling method and system. The base calling method includes: processing input data of a plurality of amplicons by a neural network-based base caller and generating a surrogate representation of the input data, wherein the input data includes position information and intensity information of the amplicons on images collected respectively in the n-m1th to n+m2th rounds of sequencing, the file size of the input data is smaller than that of the corresponding images, n is a natural number greater than or equal to 1 and n is greater than m1, m1 and m2 are each independently an integer greater than or equal to 0; processing the surrogate representation by an output layer to generate an output, wherein the output identifies the probabilities of bases A, C, T and G introduced into each of the amplicons in the nth round of sequencing; and determining the base type introduced into each of the amplicons in the nth round of sequencing based on the output. The method can efficiently process the input data and obtain accurate base calling results.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data processing based on artificial intelligence, in particular, to the application of neural networks to processing biological information such as information collected by nucleic acid sequencing, more particularly, to a base calling method based on neural networks, a system, a computer program product, a computer device and a computer readable storage medium. BACKGROUND

[0002] The subject matter discussed in this section should not be taken as an admission that the background art is prior art. Similarly, discussion of a problem in this section in relation to the subject matter provided as background should not be taken as an admission that the problem was previously recognized in the art. The subject matter in this section merely represents different approaches, which in themselves can also correspond to embodiments of the technical solutions contained in the claims.

[0003] In the related art, machine learning, a subset of artificial intelligence, enables systems to improve automatically through experience, and is widely used in image recognition, speech recognition, natural language processing, recommendation systems, etc. Machine learning includes various algorithms, such as linear regression, decision trees, support vector machines, etc. Traditional machine learning methods usually require manual feature extraction, and data preprocessing and feature extraction screening are usually very important for effective training of the model. Deep learning, another subset of machine learning, can automatically extract features from raw data, and is often focused on neural networks, especially multi-layer neural networks (deep neural networks), such as convolutional neural networks or recurrent neural networks, etc., to automatically extract features. In terms of computing power, compared to relatively simple machine learning models that can run on ordinary computers, deep learning models are usually more complex and often require a large amount of computing resources and GPU support.

[0004] In the related art, sequencing generally refers to determining the primary structure or sequence of biological polymers, including nucleic acids such as DNA and RNA, etc., including determining the order of nucleotide bases (adenine A, guanine G, thymine T / uracil U, and cytosine C) of a given nucleic acid fragment. Such methods usually include identifying the base at one or more positions in the nucleic acid, i.e., base calling, to determine at least one segment or part of the sequence of the nucleic acid molecule.

[0005] The change in the signal and / or signal intensity corresponding to the binding of a nucleotide / base to a specific position of the nucleic acid molecule (template) to be tested can indicate the base type at that position on the nucleic acid molecule, e.g., different bases can be identified using different fluorescent molecules. The binding of a nucleotide / base to a specific position of the nucleic acid molecule to be tested is also referred to as the incorporation of a nucleotide / base into the nucleic acid molecule to be tested or base extension, which can be achieved by polymerization, ligation, hybridization, etc.

[0006] Specifically, on the platform that multiple image acquisitions of the base extension signal are performed by using an optical imaging system, and nucleic acid sequencing is realized based on processing the images, due to the influence of optical effects, spatial effects and / or chemical reactions such as chromatic aberration, crosstalk and / or phasing on image acquisition, template positioning and / or target signal strength, it is often difficult to accurately identify the base based on image processing.

[0007] Specifically, for example, the current mainstream commercially available sequencing platform based on chip surface fluorescence imaging detection to realize high-throughput sequencing, a fluorescent molecule is used to label the nucleic acid molecule to be detected which has undergone a specified biochemical reaction, and the nucleic acid molecule to be detected is often a clone containing multiple copies after amplification, so as to amplify the luminescence signal after the specified biochemical reaction on the chip surface or the luminescence signal detected by imaging; the amplified nucleic acid molecule to be detected often presents as a cluster or a ball on the surface, and in the sequencing platform of ILLUMINA company or Huada intelligent manufacturing, the nucleic acid to be detected connected on the chip surface is also called fluorescent cluster or DNA nanoball (DNB).

[0008] According to the difference in the arrangement of the nucleic acid molecules to be detected on the chip surface, the chips matched with the sequencing platforms can be divided into two categories, one is a chip with fluorescent clusters or groups randomly distributed on the chip surface, which is referred to as a random chip here; the other is a chip with fluorescent clusters or groups regularly arranged on the chip surface, which is referred to as an array chip here. The position or relative distance of the fluorescent clusters or groups on the surface of the array chip is known, so for high-density chips, the image of the array chip is more conducive to the analysis of downstream algorithms than that of the random chip. Regularly arranging fluorescent clusters or groups at specified surface positions such as regularly arranged holes is a common array chip. If the array chip needs to improve the density of the fluorescent clusters or groups, the usual way is to reduce the distance between the fluorescent clusters or groups or to reduce the distance between the positions such as holes that carry or accommodate the fluorescent clusters or groups, so as to improve the density of the fluorescent clusters or groups on the chip surface.

[0009] Although high-throughput sequencing technology has made significant progress in efficiency, cost and throughput, it still faces challenges in the accuracy of base calling. In the actual sequencing process, due to the crosstalk of different fluorescence channels, the crosstalk of adjacent fluorescent clusters or groups, the resolution of the optical imaging system is limited by the diffraction limit, the lag and lead phenomena in the sequencing reaction, and the influence of the context of the base to be tested on the biochemical reaction efficiency, etc., all of which can cause base calling errors, and thus affect the accuracy of the obtained sequencing sequences and the reliability of the analysis and detection results based on these sequencing sequences. These problems are particularly prominent on high-density array chips, because the distance between fluorescent clusters or groups is reduced, and the limitation of spatial resolution makes the signal crosstalk problem more serious; moreover, the fluorescent clusters or groups are physically too close to each other to be distinguished in space or two or more clusters physically overlap and it is difficult to identify their respective nucleic acid sequences, etc. Therefore, there is often a contradiction or difficulty in achieving both quantity and / or quality (accuracy) of the read nucleic acid sequences, and often needs to be traded off or compromised.

[0010] Machine learning models, including deep learning, have great application prospects in bioinformatics research, such as training machine learning models to predict template positions, base calling, nucleotide / base sequence calling, quality scoring or evaluation of called bases or sequences, etc. Related documents such as US20230005253A1, CN117976042A, CN117912550A, US11676685B2 disclose using machine learning models to process sequencing images or sequencing data to predict the type of base incorporated into the nucleic acid molecule to be tested or the quality of the detected nucleotide / base sequence, etc. SUMMARY

[0011] The embodiments of the present application aim to at least solve one of the above technical problems or at least provide a practical commercial choice.

[0012] The inventors made this technical solution based on the discovery of the following technical problems and the results of multiple test studies:

[0013] In the base recognition methods used in current multi-color or multi-fluorescence channel high-throughput sequencing such as two-color sequencing, three-color sequencing or four-color sequencing, the signals of each fluorescence cluster on a certain region at a certain time point are often independently calculated and processed to recognize the base type incorporated by each fluorescence cluster at the time point and region; and the interference of different fluorescence channel signal crosstalk and phase imbalance (incomplete or asynchronous sequencing reactions of multiple nucleic acid molecules in the same fluorescence cluster, including advance or lag) on current base recognition is often considered, so that crosstalk correction, phasing and prephasing correction are generally performed, such as the base recognition and signal correction methods disclosed in EP3077943B1 and CN113012757B, which include calculating a fixed correction parameter or coefficient, such as calculating the brightness distribution of all fluorescence clusters and calculating a uniform phasing and prephasing correction parameter according to the brightness distribution, and using the correction parameter to uniformly correct the signal intensity / brightness of different fluorescence channels in the round or correct the signal intensity / brightness of all fluorescence channels in the round.

[0014] For this purpose, the inventors found that factors that may interfere with accurate base recognition and detection include not only the previously mentioned fluorescence channel signal crosstalk and fluorescence cluster phase imbalance, but also energy transfer and dissipation of fluorescence signals of adjacent bases before and after the to-be-detected base, energy transfer and dissipation of fluorescence signals of physically adjacent fluorescence clusters, and the influence of the context sequence of the to-be-detected base on the biochemical reaction efficiency of the current base position, which also affect the accuracy of base recognition or sequencing. Moreover, it is difficult to accurately correct the reaction signals of each base position of each fluorescence cluster at different time points (rounds) in a region (field of view) using a fixed or uniform correction coefficient, and therefore the accuracy of base recognition or sequencing is still difficult to achieve the desired accuracy.

[0015] Therefore, the inventors constructed training data, trained and tested different neural network models or cascade models containing neural networks multiple times, and trained a neural network-based base recognizer. The neural network-based base recognizer fully considers the influence of the spatial position of the fluorescence cluster, the detection time (front and rear sequencing actions) and the context (front and rear sequencing results) on the current base recognition. The base recognizer can quickly and accurately obtain the base recognition result.

[0016] In particular, the base calling method provided by the embodiments of the present application comprises: processing input data of a plurality of amplicons by a neural network-based base caller to generate a surrogate representation of the input data, wherein the input data comprises position information and intensity information of the amplicons on images respectively captured in n-m1th to n+m2th rounds of sequencing, the sequencing adopts a surface-imaging-based multi-fluorescence channel sequencing technology, the amplicons are to-be-tested nucleic acid molecules comprising a plurality of identical polynucleotide molecules, the amplicons are connected on the surface, n is a natural number greater than or equal to 2 and n is greater than m1, m1 and m2 are respectively independent natural numbers greater than or equal to 1; processing the surrogate representation by an output layer to generate an output, wherein the output identifies probabilities of bases A, C, T and G introduced into each of the amplicons in the n th round of sequencing; and determining a base type introduced into each of the amplicons in the n th round of sequencing based on the output.

[0017] The embodiments of the present application also provide a base calling system for implementing all or part of the steps of the base calling method of any of the above embodiments, which comprises: a processing unit for processing input data of a plurality of amplicons by a neural network-based base caller to generate a surrogate representation of the input data, wherein the input data comprises position information and intensity information of the amplicons on images respectively captured in n-m1th to n+m2th rounds of sequencing, the sequencing adopts a surface-imaging-based multi-fluorescence channel sequencing technology, the amplicons are to-be-tested nucleic acid molecules comprising a plurality of identical polynucleotide molecules, the amplicons are connected on the surface, n is a natural number greater than or equal to 1 and n is greater than m1, m1 and m2 are respectively independent integers greater than or equal to 0; an output unit for processing the surrogate representation by an output layer to generate an output, wherein the output identifies probabilities of bases A, C, T and G introduced into each of the amplicins in the n th round of sequencing; and a base type determination unit for determining a base type introduced into each of the amplicins in the n th round of sequencing based on the output.

[0018] The embodiments of the present application also provide a computer program product comprising instructions, when a computer executes part or all of the instructions, the base calling method of any of the above embodiments is executed.

[0019] The embodiments of the present application provide a computing device comprising: a memory for storing data comprising computer executable programs; and one or more processors for executing the computer executable programs to implement the base calling method in any of the above embodiments or examples. In some embodiments, the computing device is a sequencer.

[0020] The embodiment of the present application also provides a computer readable storage medium for storing a program for execution by a computer, and execution of the program includes completing the base recognition method of any of the above embodiments or examples.

[0021] In some embodiments, the sequencing system includes one or more of the base recognition system, the computer program product, the computing device, or the computer readable storage medium in any of the above embodiments.

[0022] With the trained neural network or cascade model (neural network-based base recognizer) which fully considers the influence of the spatial position of the to-be-tested nucleic acid molecule and the context sequence information on the current base recognition result of the to-be-tested nucleic acid molecule, and relatively small input data, the base recognition method or system or product in any of the above embodiments can efficiently process the input data and obtain accurate base recognition results.

[0023] Specifically, the training data constructed includes signal intensity information and position information of multiple amplicons at the same region and the same time point, and intensity and position information of multiple amplicons of the region in multiple rounds before and after, so that the trained model fully considers the influence of the spatial position relationship of the amplicons, the context, and the reaction information of the amplicons at different time points before and after on the base recognition of the current reaction position of the specified amplicon. The neural network-based base recognizer can correct the spatial crosstalk between physically adjacent multiple amplicons, and can also perform independent context correction for each amplicon, not only can correct the influence of phasing and prephasing, but also can eliminate part of the pattern errors caused by the context base arrangement characteristics, etc., significantly improving the accuracy of the sequencing result, and is particularly suitable for application detection with high requirements for sequencing result accuracy, such as pathogen detection, liquid biopsy, etc. Moreover, the file size of the input data is smaller than that of the corresponding image, compared with the base recognition model using image as input data, the base recognition using the model has significant advantages in computing power consumption and calculation speed.

[0024] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0025] The above and / or additional aspects and advantages of the embodiments of the present application will become apparent and be readily appreciated from the following description, taken in conjunction with the following drawings in which:

[0026] Figure 1 is a flowchart of the base recognition method of the embodiment of the present application;

[0027] Figure 2 is a flowchart of the base recognition method of the embodiment of the present application;

[0028] Figure 3 is a flowchart of a base recognition method of an embodiment of the present application;

[0029] Figure 4 is a structural diagram of a U-Net of an embodiment of the present application to which a CBAM module is added to a down-sampling layer;

[0030] Figure 5 is a flowchart of processing input data by a neural network-based base recognizer of an embodiment of the present application to predict a base recognition result;

[0031] Figure 6 is a diagram of converting an entire image into a group of blocks and writing the group of blocks as a group of matrices as input data in an embodiment of the present application;

[0032] Figure 7 is a structural diagram of a base recognition system of an embodiment of the present application;

[0033] Figure 8 is a flowchart of constructing a neural network-based base recognizer of an embodiment of the present application;

[0034] Figure 9 is a flowchart of processing input data by a U-Net to obtain a base recognition result in an embodiment of the present application;

[0035] Figure 10 is a structural diagram of a U-Net to which a spatial attention and a channel attention mechanism are added in an embodiment of the present application;

[0036] Figure 11 is a structural diagram of an input convolution layer (in_conv) in an embodiment of the present application;

[0037] Figure 12 is a structural diagram of a CBAM in an embodiment of the present application;

[0038] Figure 13 is a structural diagram of a down-sampling layer Down in an embodiment of the present application;

[0039] Figure 14 is a structural diagram of an up-sampling layer Up in an embodiment of the present application. DETAILED DESCRIPTION

[0040] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings. The embodiments described below by reference to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as limiting the present application.

[0041] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. "A set" or "a plurality" means two or more.

[0042] As used herein, the terms "first", "second", and so on do not indicate or imply relative importance or a quantity or order of the technical features indicated. In the description of the present application, the meaning of "a plurality" is two or more, unless explicitly specified otherwise.

[0043] Unless otherwise defined, "connected", "coupled" and similar terms are used broadly and encompass both direct and indirect connections, couplings and the like, and are not limited to mechanical or physical connections, couplings or the like. Connections can be fixed or detachable, integral or not, mechanical, electrical or communicative, direct or indirect, internal or external, and can be through physical attraction or chemical bonding such as incorporation through polymerization, etc. The skilled person will understand the specific meaning of the term in the corresponding example based on the description of the specific embodiment, including the context and common knowledge.

[0044] As used herein, "sequencing" refers to nucleic acid sequencing, also known as "nucleic acid sequencing" or "genetic sequencing", which refers to determining the primary structure of a nucleic acid molecule, i.e. the order of bases in the nucleic acid molecule. Sequencing can be achieved using sequencing by synthesis (SBS), sequencing by ligation (SBL) or sequencing by hybridization (SBH), etc. Unless otherwise specified, as used herein, sequencing by synthesis includes the commonly understood SBS (typical ILLUMINA / Solexa technology) that uses a polymerase to catalyze the incorporation of nucleotides into a nucleic acid molecule to be tested (polymerization reaction) and detects the corresponding reaction signal to identify the type of nucleotide incorporated, and also includes sequencing similar to SBS that uses a polymerase or non-polymerase to controllably introduce or ligate nucleotides to a nucleic acid molecule to be tested, and directly or indirectly detects the corresponding signal to determine the type of nucleotide ligated, such as SBL, SBH and the like based on surface fluorescence imaging detection.

[0045] The sequencing can be performed by a sequencing platform. According to embodiments of the present application, the sequencing platform that can be optionally used includes, but is not limited to, the Hiseq, Miseq, Nextseq and Novaseq series of sequencing platforms of Illumina, the Ion Torrent platform of Thermo Fisher / Life Technologies, the BGISEQ and MGISEQ / DNBSEQ platform of Huada, the AVITI of Element Bioscience, and the single molecule sequencing platform. The sequencing mode can be selected as single-end sequencing, double-end sequencing, or the sequencing mode supported by the selected automatic sequencing platform.

[0046] In some examples, the sequencing sequence or read is obtained by multiple rounds of sequencing by sequencing-by-synthesis. For example, the nucleic acid molecule to be tested and the polymerase and modified nucleotides are contacted and placed in conditions suitable for polymerization reaction, the modified nucleotides are controllably incorporated into the nucleic acid molecule to be tested or controllably single-base extension is achieved, and the corresponding reaction signal is detected, based on which the type of nucleotide incorporated into the nucleic acid molecule to be tested in this reaction is determined, and multiple controllable single-base extensions and corresponding signal detections are performed, so as to detect the type of nucleotide or base incorporated into the nucleic acid molecule to be tested in multiple or multiple rounds of reactions based on the reaction signal information, so as to read out a part of the sequence of the nucleic acid molecule to be tested.

[0047] The nucleic acid molecule to be tested, also referred to as nucleic acid template or template, can be a single molecule without amplification, or a molecular cluster or group or long chain containing multiple identical polynucleotide molecules after amplification, such as the fluorescent cluster or DNA nanoball (DNB) formed by bridge amplification or rolling circle amplification adopted by the current mainstream sequencing platform on the market. The nucleic acid molecule to be tested can be in the form of single strand, double strand and / or hybridization with probe or primer complex.

[0048] The corresponding reaction signal can be a fluorescence signal, or can be converted into image data (such as color image or grayscale image) formed by collecting the fluorescence signals. Thus, the image data is processed and analyzed to detect the nucleotide incorporated into the nucleic acid molecule to be tested in each or each round of reaction, so as to determine a part of the base sequence of the nucleic acid molecule to be tested.

[0049] Specifically, in some examples, sequencing is implemented based on surface fluorescence imaging detection, the nucleic acid molecules to be detected are attached to a solid surface, for example, the nucleotides can be modified to carry or be capable of binding a fluorescent label, and to carry a removable inhibiting group that prevents other nucleotides from being polymerized and attached to the next position of the nucleic acid molecules to be detected (such modified nucleotides are also referred to as reversible terminators), and after each polymerization reaction or single-base extension reaction is completed, excitation is performed to cause the fluorescent label to emit light, and the light emission signals are collected to obtain an image of the nucleic acid molecules to be detected at the specified surface position that undergoes single-base extension reaction; then, the inhibiting group and the fluorescent label are removed for the next or next round of polymerization reaction and signal collection (photographing), and the polymerization-photographing-removal is repeated multiple times or multiple rounds to obtain a set of image information associated with the nucleotides attached to the nucleic acid molecules to be detected in each single-base extension reaction.

[0050] It can be understood that the nucleic acid molecules to be detected at the specified position on the surface emit fluorescence after polymerization reaction, which generally appears as bright spots or bright spots with higher background signal intensity in the corresponding position of the image collected in this round of reaction. Therefore, according to the information corresponding to the specific chemical characteristics (nucleic acid molecules to be detected) contained in the set of images, it can be determined whether the nucleic acid molecules to be detected at the specified position undergo polymerization reaction, and the type of nucleotide attached to the nucleic acid molecules to be detected by biochemical reaction can be detected by combining the corresponding relationship between the preset fluorescence emission signal and the type of nucleotide, thereby determining at least part of the sequence of the nucleic acid molecules to be detected, obtaining the so-called read.

[0051] The set of positions corresponding to the chemical characteristics (nucleic acid molecules to be detected) on the surface is determined based on the image, that is, the template is obtained, which can be determined by identifying the features (positions) corresponding to the chemical characteristics in the image, or can be determined by identifying other features in the image that have a specific association relationship with the spatial position relationship of the chemical characteristics, for example, for a regular array surface containing a marker region (also commonly referred to as a tracer region), the spatial position relationship between the marker region and the reaction region (the region where the nucleic acid molecules to be detected are located) is generally preset and known, and the signal from the marker region is obviously distinguishable from the signal from the reaction region or the signal from the marker region appears as a detectable feature on the image. By identifying the signal from the marker region to determine the position of the marker region on the image, the positions of the amplicons (nucleic acid molecules to be detected) in the reaction region can be determined.

[0052] It should be noted that the nucleotides referred to herein include ribonucleic acid or deoxyribonucleic acid, including natural nucleotides or derivatives or modifications thereof (also referred to as modified nucleotides or the like). In this article, sometimes the base contained in the nucleotide is used to refer to the nucleotide, and those skilled in the art can clearly understand it according to the conventional knowledge and / or context.

[0053] The so-called one round of sequencing can include one repeat of base extension reaction. For example, four different nucleotides (dATP, dTTP, dGTP and dCTP) can be put into the same polymerization reaction system with a plurality of nucleic acid molecules to be tested, so that each nucleotide can be excited to emit a signal distinguishable from other types of nucleotides, thereby determining the type of nucleotide incorporated or introduced at a position of any nucleic acid molecule to be tested through one repeat reaction. For example, the four nucleotides are respectively labeled with fluorescent markers of four different light emission wavelengths, and four-color or four-channel single molecule sequencing or high-throughput sequencing is performed; for another example, the four nucleotides are respectively labeled with fluorescent markers of three different light emission wavelengths and no fluorescent marker (cold nucleotide), and two-color or three-color high-throughput sequencing is performed, one round of sequencing includes one repeat, and the base type at a position on any nucleic acid template can be detected.

[0054] The so-called one round of sequencing can also include multiple repeats. For example, four nucleotides can be sequentially contacted with a plurality of nucleic acid molecules to be tested on the surface, and base extension and collection of corresponding reaction signals such as photographing are performed respectively, and one round of sequencing includes four times of base extension; for another example, four nucleotides are contacted with a plurality of nucleic acid molecules to be tested on the surface in any combination, for example, two-by-two combination or one-three combination, two combinations are respectively subjected to base extension and collection of corresponding reaction signals such as photographing, and one round of sequencing includes two times of base extension. In some embodiments, one repeat is also sometimes referred to as one round, and those skilled in the art can understand it according to the description of specific embodiments including the associated context.

[0055] In this paper, the so-called amplicon is an amplified nucleic acid molecule to be tested, and an amplicon is a cluster or chain or group containing a plurality of identical polynucleotide sequences; the so-called fluorescent cluster, group or ball is an amplicon that can be excited to emit light such as fluorescence after a specified biochemical reaction. The amplicon can emit light and be imaged and detected during sequencing, and the base incorporated or linked to the nucleic acid molecule to be tested after the specified biochemical reaction can be detected by processing the signal and / or corresponding image information collected by imaging, thereby determining at least part of the sequence of the nucleic acid molecule to be tested.

[0056] The "chromatic aberration" (CA) referred to in the embodiments herein means the phenomenon that an optical lens cannot focus light of various wavelengths at the same point [Max Born; Emil Wolf. Principles of Optics: Electromagnetic Theory of Propagation, Interference and Diffraction of Light (7th Edition). Cambridge University Press. October 13, 1999: 334. ISBN 0521642221.]; in imaging, chromatic aberration means that each color in the spectrum cannot be focused at the same point on the optical axis, and for a sequencing platform that involves imaging the same object (e.g., one or more nucleic acid molecules) using light of multiple wavelengths, at least, chromatic aberration will cause the object to have different positions / coordinates in multiple images of the object taken at different wavelengths, or in other words, the object does not actually move, but it appears to move in multiple images of the object taken at different wavelengths due to chromatic aberration.

[0057] The so-called crosstalk (or laser-crosstalk or spectra-crosstalk), also known as "spectral crosstalk" or "spectral cross-over", generally refers to the phenomenon that the signal of one base spreads into the signal of another base; for a sequencing platform that uses different fluorescent molecules to identify different bases, if the emission spectra of two or more fluorescent molecules selected have overlap, it is possible to detect the spread of the signal of one fluorescent molecule into another fluorescent channel in one round of sequencing.

[0058] The so-called phase misphasing, phase imbalance, misphasing or phase difference generally refers to the phenomenon that the reactions between nucleic acid molecules in a group, such as a nucleic acid molecule cluster, are not synchronized in a chemical reaction, including phasing or sequence lag and prephasing or sequence lead; in a sequencing platform that uses different fluorescent molecules to identify different bases, it is manifested as the phenomenon that the signal of a fluorescent molecule corresponding to a base at a specific position is not zero in more than one round of sequencing. Generally, sequencing is performed using nucleotides labeled with fluorescent molecules and blocking groups, and the blocking group on the nucleotide can prevent other nucleotides from binding to the next position of the template, and the blocking group is, for example, an azide connected to the 3' position of the sugar group of the nucleotide; the shedding of the blocking group or the failure to remove the blocking group before the next base is extended will both cause phase misphasing.

[0059] In the embodiments of the present application, the terms amplicon, fluorescent cluster or ball, nucleic acid molecule cluster, colony, nucleic acid molecule to be detected, template, nucleic acid template, template dot, bright spot corresponding to a chemical feature or fluorescent cluster on the surface, etc. can be used interchangeably unless otherwise specified. The terms signal intensity, intensity, brightness, and fluorescence brightness, etc. can also be used interchangeably. The terms surface, chip, array, etc. can also be used interchangeably in the embodiments of the present application.

[0060] Please refer to Figure 1 In some embodiments, the base calling method comprises: S10 processing input data of a plurality of amplicons by a neural network-based base caller and generating a surrogate representation of the input data, wherein the input data comprises position information and intensity information of the amplicons on images respectively captured in n-m1th to n+m2th rounds of sequencing, the sequencing adopts a surface-imaging-based multi-fluorescence channel sequencing technology, the amplicons are nucleic acid molecules to be detected comprising a plurality of identical polynucleotide molecules, the amplicons are connected on the surface, the file size of the input data is smaller than that of the corresponding images, n is a natural number greater than or equal to 1 and n is greater than m1, m1 and m2 are respectively integers greater than or equal to 0; S20 processing the surrogate representation by an output layer to generate an output, wherein the output identifies the probabilities of bases A, C, T and G introduced into each of the amplicons in the n th round of sequencing; and S30 determining the base type introduced into each of the amplicons in the n th round of sequencing based on the output.

[0061] The neural network-based base caller is a single neural network model or a composite model or a cascade model comprising a neural network trained to predict the base type based on input data. The neural network referred to here is a multi-layer neural network comprising at least one hidden layer.

[0062] The input data, sometimes also referred to as input features or inputs, comprises information representing or reflecting target signals, i.e. amplicons that have undergone a specified biochemical reaction, in images, and the file size of the input data is smaller than that of the corresponding images.

[0063] The file size, sometimes also referred to as size or data volume, reflects the storage space required to store the data or the storage space occupied.

[0064] The amplicon is an amplification product comprising a plurality of identical polynucleotide molecules, which are connected on a solid surface and often present as clusters (identical nucleic acid molecules arranged in close proximity to form clusters) or balls (identical nucleic acid molecules connected head to tail / series, or formed into balls under external or internal forces), thus, it is also sometimes referred to as cluster, molecule cluster / ball, fluorescent cluster, nanoball, colony, etc. in this document.

[0065] The embodiments of the present application do not limit the way of obtaining surface amplificates, which can be obtained by liquid phase amplification or solid phase amplification using various amplification methods, such as chain displacement amplification on solid surface, such as template walking or bridge PCR to obtain amplificates connected to the surface, or rolling circle amplification in liquid phase to obtain amplificates, and then loading or adhering the amplificates to the designated position of the solid surface to obtain amplificates connected to the surface.

[0066] In the multi-fluorescence channel high-throughput sequencing technology based on surface imaging, the amplificates or amplificates connectable are randomly distributed on the surface of the chip (random chip) or regularly distributed (pattern chip or array chip). In each round of sequencing process, the luminescence signal (fluorescence signal) of the amplificates from a certain imaging or scanning area (field of view) is collected to obtain a plurality of images corresponding to the number of fluorescence channels, such as four-color high-throughput sequencing, a field of view generally collects four images in one round of sequencing, such as three-color high-throughput sequencing, a field of view generally collects two images in one round of sequencing, such as two-color high-throughput sequencing, a field of view generally collects two images in one round of sequencing. The images collected by sequencing include images directly collected by sequencing, images collected by sequencing after preprocessing, or images generated based on the reconstruction or recombination of images collected by sequencing, which are sometimes referred to as original images in this paper.

[0067] The present application is not limited as to the method of determining the location of amplicons on the image. It is understood that the image taken by sequencing contains features corresponding to the luminescent amplicons in the field of view, and the signals of the features are often stronger than the background, and often appear as isolated or optically distinguishable bright spots or bright patches on the image. The amplicons on the image can be located based on the identification or detection of the features, such as the image from the random chip surface, and the more reliable bright patches corresponding to the amplicons can be identified based on the signal intensity and / or morphology of the features on the image, and the template (the set of locations of the amplicons or the set of locations of the bright patches corresponding to the amplicons) can be determined. The location of the amplicon on the image or the location of the connectable amplicon on the surface can also be located based on the relative positional relationship between the location of the connectable amplicon on the surface and the specific marker, such as the image from the array chip surface, by identifying the features on the image corresponding to the specific marker, such as the marker line or marker pattern (tracer), and based on the predetermined spatial positional relationship between the specific marker and the region of the amplicon or the surface region of the connectable amplicon (e.g., the amplicon is connected in the well), and the predetermined arrangement rule of the surface position region of the amplicon or the connectable amplicon, the positional information of the amplicon or the surface region of the connectable amplicon on the image (the positional information of the well connected with the amplicon) can be determined. Further, the intensity information of the location can be determined as the intensity information of the amplicon based on the determined positional information of the amplicon on the image.

[0068] The position information can be represented by a coordinate system, which can be a unified coordinate system based on registration or alignment of multiple images of the same field of view. The position information can be represented by absolute or relative coordinate values, or can be assigned other values, such as numbers, to represent the relative position relationship of the amplicon on the image to reveal the position information of the amplicon, including the field of view (surface area) from which the amplicon comes and its position relationship with other amplicons in the field of view, etc. For example, the feature of the amplicon on the image occupies one or more pixels or multiple sub-pixels, and the position information of the amplicon on the image can be represented by the position of any pixel or sub-pixel where the feature is located, or the position of the sub-pixel center of the feature determined by methods such as quadratic function interpolation. The intensity of the amplicon on the image, also known as brightness, can be represented by the pixel value, sub-pixel value or sub-pixel value at the position of the amplicon on the image. For example, the feature of the amplicon on the image occupies one or more pixels or multiple sub-pixels, and the intensity information of the amplicon on the image can be represented by the intensity value of any pixel or sub-pixel where the feature is located, or the intensity value of the pixel or sub-pixel where the center or centroid of the feature / amplicon is located. The intensity information of the amplicon on the image can be the original signal intensity value on the image, or the processed signal intensity value, such as the intensity value after background removal and / or at least one correction of color difference, crosstalk and phase, or the intensity value after standardization or normalization processing, etc. The crosstalk correction, phase correction, etc. can be performed by the methods disclosed in the published documents, such as EP3077943B1, CN113012757B, etc. The entire contents of the cited documents are incorporated herein.

[0069] The so-called alternative representation can refer to the output after inputting one or more convolutional layers or hidden layers of the trained base recognizer, or the new representation after the input features are converted by the encoder in the neural network. This new representation is also commonly referred to as encoding or feature map.

[0070] Using a trained neural network single model or a cascade model containing a neural network (neural network-based base recognizer) that takes into account the influence of the spatial position of the nucleic acid molecule to be tested and the context sequence information on the current base recognition result of the nucleic acid molecule to be tested, and relatively small size input data, the base recognition method can efficiently process the input data and obtain accurate base recognition results.

[0071] Specifically, the training data constructed by the inventors includes signal intensity information and position information of multiple amplicons at the same time point (the same round of sequencing) in multiple regions, intensity and position information of the amplicons in the same region in multiple rounds, so that the trained model can fully consider the spatial position relationship of the amplicons, the context, and the influence of the reaction information of the amplicons at different time points on the identification of the current reaction position base of the specified amplicon. The base identifier based on the neural network can correct the spatial crosstalk between physically adjacent multiple amplicons, and can also perform independent context correction for each amplicon. Not only can it correct the influence of phasing and prephasing, but also can eliminate pattern errors caused by context base arrangement characteristics to a certain extent. Moreover, the file size of the input data is smaller than that of the corresponding image, and compared with the base identification model whose input data is image, the base identification using the model has significant advantages in terms of computing power consumption and calculation speed.

[0072] In certain preferred embodiments, n is greater than or equal to 2, and each of m1 and m2 is independently greater than or equal to 1. In this way, the input data containing the amplicons of the target round of sequencing and the multiple rounds of sequencing before and after it are input to the neural network-based base identifier, so that the spatial position relationship of the amplicons, the context, and the influence of the different time points (the rounds of sequencing before and after the target round) on the base extension and identification of the specified amplicon in the current round can be fully considered and corrected. Therefore, the base type introduced by the amplicon in the target round of sequencing can be accurately predicted.

[0073] In certain embodiments, the neural network-based base identifier is a single model. Specifically, the neural network-based base identifier 1000 is a multi-layer neural network model (semantic segmentation model) trained to achieve a semantic segmentation task. Specifically, the semantic segmentation model is selected from at least one of a network with an encoder-decoder structure or a transformer-based network. In the embodiments of the present application, when the neural network-based base identifier is a single model, it is also directly referred to as a semantic segmentation model.

[0074] In some embodiments, the neural network-based base caller is a cascade model comprising two or more models connected in series. The cascade model comprises a first sub-model and a second sub-model connected in series, each of the first sub-model and the second sub-model independently selected from one of a semantic segmentation model and a machine learning model, the semantic segmentation model selected from at least one of a network with an encoder-decoder structure or a transformer-based network, and the machine learning model selected from at least one of a tree model. In some embodiments, at least one of the models in the cascade model is a semantic segmentation model. In this way, the cascade model can efficiently consider the influence of space, context, etc. on the accuracy of the prediction results, while reducing the requirement for computing power.

[0075] Further, please refer to Figure 2 In some embodiments, the base calling method comprises: S50 processing, by the first sub-model, the input data of each of the nth-m1th round to the nth+m2th round of sequencing of each amplicon to generate a first prediction result of each of the specified rounds, respectively, the first prediction result of each of the specified rounds comprising the probability of each of the specified rounds of sequencing introducing a base of A, C, T and G into each amplicon; and S52 processing, by the second sub-model, the first prediction result of each of the specified rounds and the round information in which each of the specified rounds is located to generate the output. In some embodiments, the first sub-model is a semantic segmentation model. In this way, the cascade model can efficiently consider the influence of space, context, etc. on the accuracy of the prediction results, while being fast and efficient.

[0076] Please refer to Figure 3 In some embodiments, the base calling method comprises: S60 processing, by the first sub-model, the input data of the nth-m1th round to the nth+m2th round of each of the specified amplicons to generate a second prediction result of each of the specified rounds of each of the specified amplicons, the second prediction result of each of the specified rounds of each of the specified amplicons comprising the probability of each of the specified rounds of sequencing introducing a base of A, C, T and G into each of the specified amplicons; and S62 processing, by the second sub-model, the second prediction result of each of the specified rounds of each of the specified amplicons to generate the output. In some embodiments, the second sub-model is a semantic segmentation model. In this way, the cascade model can efficiently consider the influence of space, context, etc. on the accuracy of the prediction results, while being fast and efficient.

[0077] In some embodiments, in the case that the neural network-based base caller is a cascade model, the cascade model comprises a semantic segmentation model and a machine learning model, and the machine learning model is a gradient boosting decision tree (GBDT) using at least one of LightGBM and XGBoost algorithms. In some tests, it is found that the combination of the semantic segmentation model and the machine learning model can effectively consider the influence of space and context on the accuracy of the prediction results, and is fast and efficient.

[0078] The semantic segmentation model in the embodiments of the present application is, for example, a convolutional neural network (CNN), and specifically, can be selected from at least one of U-Net, DeepLab, SegFormer, PSPNet, FCN, GCN, and their respective variants.

[0079] More specifically, in some embodiments, the semantic segmentation model is U-Net or SegFormer. Preferably, it is U-Net or SegFormer with an attention mechanism, such as U-Net or SegFormer with a CBAM module (Convolutional Block Attention Module) or an SE module (Squeeze-and-Excitation) added to the encoding part, or U-Net with a transformer-based multi-head attention mechanism added to the decoding part and the skip connection part. The inventors have trained and tested these network structures and their variants using the same training data, and the training and testing results show that the use of these models or their variants can obtain a prediction model with high classification result accuracy.

[0080] U-Net (arXiv: 1505.04597 [cs.CV]) is a convolutional neural network (CNN) architecture for semantic segmentation, first proposed by Ronneberger et al. [Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation [C] / / Medical image computing and computer-assisted intervention - MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. Springer International Publishing, 2015: 234-241.]. U-Net consists of a symmetric encoder-decoder structure, where the encoder part is responsible for extracting features of the input data, obtaining high-level feature representations through stepwise down-sampling operations such as convolution and max-pooling, and the decoder part is used to recover the spatial resolution of the input data by stepwise up-sampling, making it consistent with the input. In order to enhance the resolution and retain the details, U-Net uses skip connections to pass the features of the encoder to the decoder at each layer.

[0081] In some specific examples, to improve the performance of U-Net in semantic segmentation tasks, attention mechanisms can be added to the encoder or decoder, for example, adding SE modules [Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-Excitation Networks. arXiv preprint arXiv: 1709.01507.], which can strengthen significant features and suppress redundant information by calculating attention weights for each channel. Adding SE modules to the encoder part of U-Net can make the network pay more attention to important channels and improve the encoding efficiency. For another example, adding CBAM modules [Woo S, Park J, Lee JY, et al. CBAM: Convolutional block attention module [C] / / Proceedings of the European conference on computer vision (ECCV). 2018: 3-19.], which combine channel attention and spatial attention, and through information weighting in channels and space, the network pays more attention to important features. If CBAM modules are added to the encoding part, the network will more effectively extract input data features and optimize the features in space and channels. In addition, U-Net variants can also introduce transformer-based multi-head attention mechanisms in the decoder and the skip connection part. Multi-head attention can model detailed information in a wider context, so that the decoder can more effectively restore details and achieve higher accuracy in segmentation tasks.

[0082] SegFormer (arXiv: 2105.15203[cs.CV]) is a lightweight, Transformer-based semantic segmentation architecture designed for efficient and accurate semantic segmentation. SegFormer includes two key components, an encoder and a decoder. The encoder consists of Transformer encoders (e.g., Vision Transformer) that can capture long-range dependencies while generating multi-scale feature representations. The decoder of SegFormer directly fuses the multi-scale features generated by the encoder and produces the final segmentation results. SegFormer already has a built-in attention mechanism based on Transformer, and further adding attention mechanisms to SegFormer can further enhance its performance. For example, adding CBAM or SE modules to the encoder part of SegFormer can further highlight meaningful features while suppressing irrelevant features; for example, adding multi-head attention mechanisms to the decoder part can enhance the interaction between feature maps, and multi-head attention can model the interaction between different scale features, improving the accuracy and resolution of the segmentation results.

[0083] In some specific examples, such as in network structures like U-Net, the so-called neural network-based base caller 100 or the semantic segmentation model 200 contained therein includes: an encoder 120 containing multiple down-sampling layers 102, each down-sampling layer 102 containing at least one pooling operation and at least one convolution operation, and at least part of the down-sampling layers 102 adding CBAM modules 1020 or SE modules 1022; a decoder 140 containing multiple up-sampling layers 104, and the number of up-sampling layers 104 being the same as the number of down-sampling layers 102, each up-sampling layer 104 including at least one deconvolution operation or bilinear interpolation operation; and a skip connection 160 connecting the down-sampling layer 102 and the corresponding up-sampling layer 104.

[0084] Further, in some specific examples, the encoder 120 further includes an input convolutional layer 101, and the input data is processed by the input convolutional layer 101 to generate a substitute representation as the input of the first down-sampling layer 102 of the encoder 120, and the input convolutional layer 101 contains multiple convolution operations; and / or, the decoder 140 further includes an output convolutional layer 105, and the output convolutional layer 105 processes the substitute representation generated by the last up-sampling layer 104 of the decoder 140 to generate the output.

[0085] Figure 4 A structural diagram of a U-Net in which a CBAM module 1020 is added to a down-sampling layer in some embodiments is shown. In the figure, "Down" represents a down-sampling layer 102, "up" represents an up-sampling layer 104, and the arrow pointing from the down-sampling to the up-sampling represents the skip connection 160.

[0086] In some more specific examples, the neural network-based base caller 100 or the segmentation model 200 further has one or more of the following technical features (a)-(e): (a) The encoder 120 contains four down-sampling layers 102. (b) Each down-sampling layer 102 contains a CBAM module 1020 or a SE module 1022 and a Down structure 1024 connected to the CBAM module 1020 or the SE module 1022; in a specific example, the Down structure 1024 contains one pooling and two convolution operations. (c) The CBAM module 1020 contains a channel attention module 10202 and a spatial attention module 10204 connected in series, the channel attention module 10202 includes a first pooling unit 10202a, an MLP unit 10202b and a first output layer 10202c, and the spatial attention module 10204 includes a second pooling unit 10204a, a convolution unit 10204b and a second output layer 10204c. In the channel attention module 10202, the input is respectively processed by global max pooling and global average pooling of the first pooling unit 10202a to generate two channel descriptors, the two channel descriptors are processed by a shared MLP in the MLP unit 10202b to obtain channel weights, and then the result of weighting the input by the channel weights is output by the first output layer 10202c; in the spatial attention module 10204, the output from the first output layer 10202c is respectively processed by global max pooling and global average pooling of the second pooling unit 10204a to generate two spatial descriptors, the two spatial descriptors are processed by a convolution layer of the convolution unit 10204b to obtain spatial weights, and then the result of weighting the output of the first output layer 10202c by the spatial weights is output by the second output layer 10204c. (d) The SE module 1022 contains a channel attention module 10222, the channel attention module 10222 includes a third pooling unit 10222a, a fully connected unit 10222b and a third output layer 10222c, one channel of the input is processed by global average pooling of the third pooling unit 10222a to generate a corresponding channel descriptor, the channel descriptors are processed by two fully connected layers and an activation function of the fully connected unit 10222b to obtain channel weights, and then the result of weighting the input by the channel weights is output by the third output layer 10222c. (e) Each up-sampling layer 104 contains one bilinear interpolation or one deconvolution operation (Transposed Convolution, also known as inverted convolution), and a concatenation operation that concatenates the output of the previous layer and the result of the corresponding down-sampling layer 102 transmitted to the up-sampling layer 104 by the skip connection 160 in the channel dimension.

[0087] In particular, in (b), the Down structure of the down-sampling layers refers to the encoding path of the convolutional neural network, which is used to extract high-level features of the input data, for example in the network structure of U-Net. These layers obtain feature representations at different scales by progressively reducing the spatial resolution of the feature maps, i.e. down-sampling. Common components of the Down structure include convolution, pooling layers, batch normalization, and Dropout layers (optional), each down-sampling layer or block usually contains two convolution operations (e.g. 3x3 convolution kernel) and usually adopts ReLU activation function. Each convolution extracts features of the image, thereby gradually constructing a more abstract representation. The stride of the convolution operation is usually 1 to ensure that the resolution of the feature map does not change at this step. After the convolution operation, a Max Pooling layer is usually used, which can set a 2x2 pooling kernel and use a stride of 2. This step reduces the resolution of the feature map by selecting the maximum value in the pooling region, thereby down-sampling the feature map. The role of the Max Pooling layer is to reduce the spatial dimension of the feature map while retaining the most important feature information, thereby improving the computational efficiency of the model. Sometimes, a Batch Normalization layer is added after the convolution operation. It normalizes the mean and variance of the convolution output, stabilizes the training process and speeds up the convergence. Batch normalization helps to reduce the internal covariate shift and improve the expressiveness of the network. In addition, in some applications, a Dropout layer is added in the Down structure to reduce overfitting. Dropout helps to improve the generalization ability of the model by randomly discarding some neurons.

[0088] In (e), bilinear interpolation is used for preliminary up-sampling, which can stretch the feature map to a higher resolution than the original resolution. Bilinear interpolation is a non-learned, simple and fast up-sampling method. Deconvolution is a learned up-sampling method that generates high-resolution feature maps from low-resolution feature maps by applying convolution layers in reverse. For example, in U-Net, the up-sampled feature maps generated by deconvolution are concatenated with the feature maps of the corresponding layers in the encoding path (down-sampling stage) in the channel dimension, thereby retaining low-level detail information and enhancing the network's ability to recognize edge and other detailed features. In a comparative test, it was found that bilinear interpolation is fast, but if it is faced with complex data, the effect is not as good as deconvolution.

[0089] In some embodiments, for the cascaded model, the first sub-model and the second sub-model are independently trained using non-overlapping training data. In this way, overfitting can be prevented, so that the trained model loses or is insufficient in generalization ability for new data. It is also convenient to train the sub-models in parallel and quickly obtain the trained cascaded model.

[0090] In some embodiments, training the neural network-based base caller or semantic segmentation model based on the training data comprises computing gradients of a loss function using backpropagation to guide random gradient descent to update the model parameters to minimize the loss function, wherein the loss function is a hybrid loss function comprising at least one of a cross-entropy loss, a dice loss, a mean squared error, and one or more custom loss functions, which can be parameters commonly set for a sequencing scenario to evaluate sequencing results or intermediate results, such as sequencing error rate, pattern error feature parameters, etc. In one example, the custom loss function of the hybrid loss function comprises a mismatch loss, which reflects the proportion or content of bases that are inconsistent between the predicted results and the true results.

[0091] In related embodiments, the loss function is a function that measures the error between the predicted value and the true value. Given a training data set, the error between the predicted value and the true value is only related to the model parameters, and thus it is denoted as a function of the model parameters. In machine learning or deep learning, the function that measures the error is called the loss function. Optimization algorithms such as stochastic gradient descent (SGD, such as mini-batch gradient descent algorithm) are used to find numerical solutions by iteratively updating the model parameters to optimize the loss function. Backpropagation (BP) is a common method used in conjunction with optimization algorithms such as SGD to train the network, and the basic idea can be summarized in the following steps: forward propagation (calculating the network output result), calculating the error (calculating the error between the network output result and the true result (the error is usually measured by the loss function)), backpropagating the error (calculating the gradient of the loss function with respect to each parameter), and updating the parameters (updating the network parameters according to the calculated gradient).

[0092] In one example, the optimization algorithm or optimizer used is SGD; during the model training process, SGD calculates the gradient by randomly selecting a portion of the samples, and then updates the parameters of the network according to the gradient. In one example of training a U-Net network, the U-Net defines the structure and function of the network, and the SGD optimizes the parameters of the U-Net, uses multiple loss functions (hybrid loss functions) during backpropagation, and includes custom loss functions to guide the training direction of the network; the custom loss function is, for example, a mismatch loss function, which can be defined as the proportion of the number of bases that are different between the training results and the true labels to the total number of base calls in the true labels.

[0093] In some specific examples, the hybrid loss function comprises a cross-entropy loss, a dice loss, and a mismatch loss, and tests find that it is beneficial to train a model with a higher accuracy of base calling results to make the relative weight of the dice loss the second highest or the highest. In some more specific examples, the weights of the cross-entropy loss, the dice loss, and the mismatch loss in the hybrid loss function are 1:3:4, 1:4:4, 1:5:4, 1:6:4, 1:7:4, 1:7:5, etc., respectively, and all can train a prediction model with better performance. It should be noted that the embodiments of the present application are limited by specific numerical values, unless otherwise specified, and are allowed or include a fluctuation of ±10% of the original numerical value.

[0094] In some examples, the so-called neural network-based base caller or semantic segmentation model is trained in a supervised learning or semi-supervised learning manner.

[0095] The embodiments of the present application do not limit the way of obtaining the training data used for training the model. The base of each round of template nucleic acid incorporation can be recognized by a conventional base calling method or software or tool to obtain the base calling result of each round or the sequence of at least a part of the template nucleic acid. In some examples, in the training data, the base on the reference sequence corresponding to the base calling result of the amplicon is taken as the target value (the target value is also referred to as the true result or the correct answer or the true label in the embodiments of the present application).

[0096] In addition, in the process of constructing the training data, the sequences with poor sequencing quality or unreliable or relatively low reliability alignment positions are removed, which can improve the reliability of the training data, and based on the training data with high reliability, it is beneficial to train a prediction model with accurate and reliable base calling results. For example, in a certain example, for sequences that fail to successfully align to the reference genome or have unreliable or low reliability alignment results, the bases of these sequences or positions are assigned as “N” or recorded as “N”, and the bases or sequences recorded as N are not involved in subsequent training, which can improve the reliability of the training data.

[0097] In some specific examples, m1 is equal to or greater than m2. In multiple tests, it is found that, especially in the cascade model, it is often beneficial to reduce the pattern error when m1 is greater than or equal to m2, for example, m1 is 9, 7, 6, 5, or 3, and the corresponding m2 is 5, 6, 4, 5, or 3, and the prediction results of the model are all good, especially the high-frequency pattern error is significantly reduced.

[0098] In gene sequencing, the types of errors can be classified into high-frequency pattern errors and non-high-frequency pattern errors according to their frequency and patterns. Both types of errors generally need to be processed in sequencing data analysis to ensure the accuracy of sequencing results.

[0099] High-frequency pattern errors generally refer to systematic errors that often occur in sequencing data. Such errors are often caused by limitations of sequencing technology, characteristics of reagents, or defects of data processing algorithms. High-frequency pattern errors often have the following characteristics: systematicity, usually occurring multiple times under the same experimental conditions; technology dependence, different sequencing technologies will produce different types of high-frequency pattern errors; predictability, due to their systematicity, these errors are often predictable and can usually be corrected by specific methods. Common error types include, for example, base repeat errors: prone to occur in long homopolymer regions, such as measuring “AAAA” as “AAA” or “AAAAA”; sequence preference errors: inaccurate recognition of certain bases or sequence fragments by sequencing technology; and insertion / deletion errors (Indels): insertion or deletion of bases in the sequence.

[0100] Non-high-frequency pattern errors generally refer to errors that are not common or randomly occur in sequencing data. These errors are often caused by accidental technical problems, variations in sample processing, or random noise. Non-high-frequency pattern errors often have the following characteristics: randomness, no obvious systematicity or repeatability, randomly distributed in the data; unpredictability, due to their randomness, these errors are difficult to predict and systematically correct; and accidentality, often caused by one-time technical problems or accidental processing errors. Common error types include, for example, random base errors: misreading of individual bases, such as incorrectly identifying “A” as “G”; accidental insertion / deletion errors: insertion or deletion of bases without a specific pattern; and noise-induced errors: random errors caused by experimental noise or environmental factors.

[0101] The embodiments of the present application do not limit the acquisition method of input data, i.e., the acquisition method of the position information and intensity information of surface amplification.

[0102] In some embodiments, the surface is a patterned surface, e.g., presenting a regular array of wells, the surface comprises marker regions and reaction regions, the spatial positional relationship of the marker regions and the reaction regions is known and the signals from the marker regions and from the reaction regions have a clear difference on the image or the signals from the marker regions appear as specific identifiable features on the image, a plurality of amplicons are respectively linked in a plurality of separated wells of the reaction regions, the template is constructed by processing at least one fluorescence channel image in the images collected in the first round of sequencing, and the input data is obtained based on the template, including: identifying the features on the fluorescence channel image corresponding to the marker regions to determine the positions of the marker regions on the image, and determining the position information of the amplicons based on the positions of the marker regions to construct the template; and mapping the position information of the amplicons on the template to the images collected in each designated round of sequencing respectively to determine the position information and intensity information of the amplicons on the images collected in each designated round of sequencing respectively.

[0103] Specifically, the trackline on the surface can be identified according to the signals from the marker regions, e.g., tracklines, and then the positions of all regularly arranged amplicon clusters can be calculated directly based on the positions of the tracklines according to the positional relationship between the tracklines and the wells connecting the amplicon clusters of the reaction regions preset when the surface design layout is designed.

[0104] In other embodiments, the template is constructed by processing the images collected in the first round or the first p1 rounds of sequencing, and the input data is obtained based on the template, p1 being a natural number greater than 0 and less than or equal to 4, including: identifying the bright spots on each image corresponding to the amplicons that have undergone a specified biochemical reaction, and merging the bright spots to obtain the template constructed; and mapping the position information of the amplicons on the template to the images collected in each designated round of sequencing respectively to determine the position information and intensity information of the amplicons on the images collected in each designated round of sequencing respectively. This method is suitable for obtaining the input data of sequencing based on the imaging detection of random array surfaces or regular array surfaces. Certain embodiments utilize the technical solutions disclosed in CN112289381B to construct the template and determine the positions of the amplicons that have undergone reactions on the images of designated rounds and the intensity values of the amplicons, etc.

[0105] In some embodiments, the template is constructed by processing images collected in p2 rounds of sequencing before the input data is obtained based on the template, p2 is a natural number greater than or equal to 6, and the method comprises: detecting a base introduced to each unit position on each image in each round according to the intensity of the unit position to determine the detected base sequence of each unit position, the image contains a plurality of unit positions, and the size of a unit position is less than or equal to the size of a pixel of the image, and a size corresponding to a feature of an amplicon on which a specified biochemical reaction occurs on the image occupies one or more unit positions; clustering the unit positions based on the similarity of the detected base sequences of the unit positions, including grouping a plurality of unit positions with a plurality of base detection sequences with a similarity satisfying a preset level as corresponding to the same amplicon to construct the obtained template, and the template contains the position information of the amplicons on the surface; and mapping the position information of the amplicons on the template to the images collected in each specified round of sequencing, respectively, to determine the position information and intensity information of the amplicons on the images collected in each specified round of sequencing, respectively. The method is suitable for obtaining input data based on imaging detection of a random array surface or a regular array surface to realize sequencing.

[0106] In some examples, the intensity information is corrected intensity information. The correction can be selected from at least one of crosstalk correction and / or phasing and prephasing.

[0107] The input data contains the position information and intensity information of the amplicons on the image, and the size of the input data is smaller than the size of the image corresponding thereto. Therefore, in some aspects, the input data can be said to be compressed image information or data extracted or screened from the image information, and can be represented in a simplified image form or a non-image form, preferably in a form suitable for computer or processor calculation and processing. In preferred examples, the input data is represented in a non-image form, such as a matrix, a multidimensional matrix, an array, a multidimensional array, etc., only recording or presenting the relative position information and intensity information of the amplicons on the image. In this way, compared with the input original image, the file size of the input data can be greatly reduced, the requirements for computing power and / or storage can be reduced, and the computer can quickly perform various operations or processing on the input data to quickly obtain a prediction result. After compression and / or extraction of specific image information and conversion into other data forms, in some examples, the file size of the input data is less than or equal to half of the file size of the corresponding image. In some high-density random or array surfaces, the file size of the input data is less than or equal to one-third of the file size of the corresponding image, or even less than or equal to one-fifth of the file size of the original image, etc.

[0108] In some embodiments, the file size of the input data is less than or equal to half of the file size of the corresponding image. In some examples, particularly for regular surfaces, the file size of the input data is less than or equal to one third of the file size of the corresponding image. In some tests, particularly for high-density regular surfaces, the file size of the input data is even less than or equal to one fourth or one fifth of the file size of the corresponding image. In this way, the storage or space occupied is significantly reduced, and the running speed is significantly improved, relative to directly inputting image data.

[0109] In particular, in some examples, the input data is represented as a plurality of groups of matrices, a group of matrices comprising a plurality of matrices, a group of matrices reflecting the position information and intensity information of the amplicons on an image collected in a round of sequencing, a matrix comprising a plurality of rows and a plurality of columns, a matrix reflecting the information of the amplicons on all or a block of an image of a fluorescent channel in a round of sequencing, an element of the matrix corresponding to an amplicon, the row and column positions of the element of the matrix reflecting the position information of the corresponding amplicon on the corresponding image, and the value of the element of the matrix reflecting the intensity information of the corresponding amplicon on the image of the fluorescent channel in the round of sequencing. In some embodiments, the matrix is also referred to as a fluorescence intensity matrix. In this way, the amount of information or size of the input data is greatly reduced compared to the input image, which is beneficial to reducing the requirements for storage or computing power, and is also beneficial to improving the running speed of the calculation.

[0110] In more specific examples, a piece of input data can correspond to a plurality of amplicons in a region or field of view (FOV), and a piece of input data, for example, one or one group of "fluorescence intensity matrices", can correspond to the position and intensity information of the amplicons that have undergone a specified reaction in the image collected in a round of sequencing of a FOV. In this way, the size of the input data is significantly reduced compared to the input original image, which can reduce the requirements for computing power and / or storage.

[0111] With regard to the number of matrices contained in a group and the dimensions of the matrices, it can be understood that a group of matrices corresponding to a group of images of a FOV in a round of sequencing of multiple fluorescent channels, for example, four fluorescence intensity matrices of amplicons from four fluorescent channels in a round of sequencing of a FOV of four-color sequencing, for example, three fluorescence intensity matrices of amplicons from three fluorescent channels in a round of sequencing of a FOV of three-color sequencing, for example, two fluorescence intensity matrices of amplicons from two fluorescent channels in a round of sequencing of a FOV of double-color high-throughput sequencing, can all be referred to as a three-dimensional matrix. In combination with the conventional knowledge, a person skilled in the art can understand the specific meaning of the description related to the matrix, the number of matrices or the dimensions according to the description of the specific embodiments including the context.

[0112] Figure 5The flowchart shows the process of the base caller based on neural network processing input data to predict base calling results. Among them, scheme 1 shows the case of a single model of the base caller based on neural network, while scheme 2.1 and scheme 2.2 show the case of a cascade model of two base callers based on neural network. Moreover, scheme 2.1 shows that the input data is first processed by the spatial crosstalk correction sub-model (model 1 in scheme 2.1), and then the output is processed by the context crosstalk correction sub-model (model 2 in scheme 2.1), that is, the cascade order or calling order of the cascade model used by scheme 2.1 is to call the spatial crosstalk sub-model first, and then call the context crosstalk correction sub-model; while scheme 2.2 shows that the input data is first processed by the context crosstalk correction sub-model (model 1 in scheme 2.2), and then the output is processed by the spatial crosstalk correction sub-model (model 2 in scheme 2.2), that is, the cascade order or calling order of the cascade model used by scheme 2.2 is that the context crosstalk correction sub-model is called first, and then the spatial crosstalk correction sub-model is called.

[0113] Further, in some examples, the advantages of the input matrix compared with the input original image are not only in the difference in image and matrix size. For example, in four images from a field of view of a regular surface collected in a round of sequencing of four-color fluorescent imaging sequencing, assuming that the size of each image is 2000*2000 pixels, and the size of the template point matrix (fluorescence intensity matrix) reflecting the position and intensity information of the amplification on the image in matrix form is 800*800, then the size is reduced to 0.4^2 times of the original. Moreover, the intensity information of the 800*800 template points of the input matrix can be the intensity information corresponding to the sub-pixel level positioning position, such as the intensity obtained by bilinear interpolation or other interpolation methods, which is not the intensity value corresponding to a pixel in the image. Therefore, in terms of information quantity, the 800*800*4 template intensity matrix information reflects the information of four 2000*2000*4 matrices plus 800*800 position matrices. It can be said that the information quantity of the input matrix reflects the position information and intensity information of the amplification determined directly or indirectly from the image information, including the processed image information. In this way, the input data contains target information, the size is significantly reduced, and it is also in the form of matrix that is easy for machine operation and calculation, which further facilitates fast prediction to obtain accurate and reliable base calling results.

[0114] In some other more specific examples, one matrix reflects the information of the amplicons in one block of one fluorescent channel image of one round of sequencing, and the input data further includes a matrix reflecting the position information of each block in the image from which it comes. In this way, the matrix corresponding to one fluorescent channel image is converted into a set of relatively small matrices, or in other words, one fluorescent channel image is converted into a set of blocks, the information of the amplicons in the simplified blocks is extracted into matrix form, and the processing speed of the input data is further improved in the same hardware and software operating environment, and the prediction result is obtained more quickly.

[0115] Please refer to Figure 6 In some specific examples, the image is converted into a set of block matrices, and all the block matrices of different sizes are zero-padded to the same size as input, so the entire image is converted into a total of 3*3, i.e., 9 blocks. The length and width of the block can be different, so after being split into 9 blocks, the edges of all the blocks with a size less than b*y can be padded with 0 values to a size of b*y. Finally, an additional position coding layer is added according to the position number of the block in the full image, for example, the block at the lower right corner is numbered 9, and a layer of Pos(P) position coding is filled, with all values being 9, to obtain the final input matrix.

[0116] Please refer to Figure 7 The base recognition system 1000 of the embodiments of the present application can implement all or part of the steps of the base recognition method in any of the above embodiments or examples. The base recognition system 1000 includes: a processing unit 1010 configured to process input data of a plurality of amplicons by a neural network-based base recognition device 100 and generate a substitute representation of the input data, wherein the input data includes position information and intensity information of the amplicons on images respectively collected in the n-m1th to n+m2th rounds of sequencing, the sequencing is surface imaging-based multi-fluorescent channel sequencing technology, the amplicon is a nucleic acid molecule to be tested containing a plurality of identical polynucleotide molecules, and the amplicon is connected to a surface, n is a natural number greater than or equal to 1 and n is greater than m1, m1 and m2 are each independently an integer greater than or equal to 0; an output unit 1030 configured to process the substitute representation by an output layer to generate an output, wherein the output identifies the probabilities of bases A, C, T, and G introduced into each of the amplicons in the n th round of sequencing; and a base type determination unit 1050 configured to determine the type of the base introduced into each amplicon in the n th round of sequencing based on the output.

[0117] It should be noted that the explanations or descriptions of the technical features and advantages of the base calling method in any of the above embodiments, examples or aspects also apply to the base calling system 1000, and thus will not be repeated in detail. For example, in some embodiments, n is greater than or equal to 2, and each of m1 and m2 is independently greater than or equal to 1. In this way, the influence of the front and rear wheel information on the base type of the specified amplicon incorporated by the current wheel can be further focused on.

[0118] In some embodiments, the neural network-based base caller is a trained single model, which is a neural network capable of performing a semantic segmentation task (semantic segmentation model). The semantic segmentation model can be selected from at least one of a network with an encoder-decoder structure or a transformer-based network.

[0119] In other embodiments, the neural network-based base caller is a trained cascade model, which includes a first sub-model and a second sub-model connected in series. Each of the first sub-model and the second sub-model is independently selected from one of a semantic segmentation model and a machine learning model. The semantic segmentation model is selected from at least one of a network with an encoder-decoder structure or a transformer-based network. The machine learning model refers to a traditional machine learning model, such as at least one selected from a tree model.

[0120] In some embodiments, the system 1000 includes a first sub-model module 1012 configured to process the input data of each of the n-m1th to n+m2th round of sequencing of the amplicon by the first sub-model to generate a first prediction result of the specified round, respectively. The first prediction result of the specified round includes the probability of each of A, C, T and G being incorporated into each of the amplicons by the sequencing of the round. The system 1000 also includes a second sub-model module 1014 configured to process the first prediction result of each specified round and the round information of each specified round by the second sub-model to generate the output. In some specific examples, the first sub-model is a semantic segmentation model, and the semantic segmentation model is selected from at least one of a network with an encoder-decoder structure or a transformer-based network.

[0121] In some embodiments, the system 1000 comprises: a first sub-model module 1016 for processing, by a first sub-model, the input data of the n-m1th round to the n+m2th round of each specified amplicon to generate the second predicted results of each specified round of the amplicon, the second predicted results of each specified round of the amplicon comprising the probabilities of bases A, C, T and G introduced into the amplicon by sequencing in each specified round; and a second sub-model module 1018 for processing, by a second sub-model, the second predicted results of each specified round of each specified amplicon to generate the output. In some specific examples, the second sub-model is a semantic segmentation model selected from at least one of a network with an encoder-decoder structure or a transformer-based network.

[0122] In particular, the semantic segmentation model is selected from at least one of U-Net, DeepLab, SegFormer, PSPNet, FCN, GCN and their respective variants.

[0123] In some embodiments involving a cascaded model comprising a machine learning model, the machine learning model generally refers to a traditional machine learning model, such as the machine learning model being a gradient boosting decision tree (GBDT) using at least one of LightGBM and XGBoost algorithms.

[0124] In some embodiments, the neural network-based base caller 100 in the system 1000 is a single semantic segmentation model or a cascaded model comprising a semantic segmentation model, and the semantic segmentation model is a U-Net or a SegFormer. Preferably, the U-Net or the SegFormer is added with an attention mechanism; or the U-Net is added with a CBAM module or an SE module in the encoding part. In another specific example, the semantic segmentation model is a U-Net added with a transformer-based multi-head attention mechanism in the decoding part and the skip connection part.

[0125] In some embodiments, the semantic segmentation model in the neural network-based base caller 100 in the system 1000 is a single model or a cascaded model comprising: an encoder 120 comprising a plurality of down-sampling layers 102, each down-sampling layer 102 comprising at least one pooling operation and at least one convolution operation, and at least a portion of the down-sampling layers 102 being added with a CBAM module 1020 or an SE module 1022; a decoder 140 comprising a plurality of up-sampling layers 104 and the number of up-sampling layers 104 being the same as the number of down-sampling layers 102, each up-sampling layer 104 comprising at least one deconvolution operation or bilinear interpolation operation; and a skip connection 160 connecting the down-sampling layers 102 and the corresponding up-sampling layers 104.

[0126] Further, in some specific examples, the encoder 120 further comprises an input convolutional layer 101, the input data is processed by the input convolutional layer 101 to generate the substitute representation as the input of the first down-sampling layer 102 of the encoder 120, the input convolutional layer 101 comprises multiple convolution operations; and / or, the decoder 140 further comprises an output convolutional layer 105, the output convolutional layer 105 processes the substitute representation generated by the last up-sampling layer 104 of the decoder 140 to generate the output.

[0127] In some more specific examples, the neural network-based base caller 100 in the system 1000 or the semantic segmentation model 200 therein further has one or more of the following technical features (a)-(e): (a) the encoder 120 contains four down-sampling layers 102. (b) Each down-sampling layer 102 contains a CBAM module 1020 or an SE module 1022 and a Down structure 1024 connected to the CBAM module 1020 or the SE module 1022; in a specific example, the Down structure 1024 contains one pooling and two convolution operations. (c) The CBAM module 1020 contains a channel attention module 10202 and a spatial attention module 10204 connected in series, the channel attention module 10202 includes a first pooling unit 10202a, an MLP unit 10202b and a first output layer 10202c, and the spatial attention module 10204 includes a second pooling unit 10204a, a convolution unit 10204b and a second output layer 10204c. In the channel attention module 10202, inputs are respectively processed by global max pooling and global average pooling of the first pooling unit 10202a to generate two channel descriptors, the two channel descriptors are processed by a shared MLP in the MLP unit 10202b to obtain channel weights, and then the result of weighting the inputs by the channel weights is output by the first output layer 10202c; in the spatial attention module 10204, the output from the first output layer 10202c is respectively processed by global max pooling and global average pooling of the second pooling unit 10204a to generate two spatial descriptors, the two spatial descriptors are processed by a convolution layer of the convolution unit 10204b to obtain spatial weights, and then the result of weighting the output of the first output layer 10202c by the spatial weights is output by the second output layer 10204c. (d) The SE module 1022 contains a channel attention module 10222, the channel attention module 10222 includes a third pooling unit 10222a, a fully connected unit 10222b and a third output layer 10222c, one channel of the input is processed by global average pooling of the third pooling unit 10222a to generate a corresponding channel descriptor, the channel descriptors are processed by two fully connected layers and an activation function of the fully connected unit 10222b to obtain channel weights, and then the result of weighting the input by the channel weights is output by the third output layer 10222c. (e) Each up-sampling layer 104 contains one bilinear interpolation or one deconvolution operation, and a concatenation operation that concatenates the output of the previous layer and the result of the corresponding down-sampling layer 102 transmitted to the up-sampling layer 104 by the skip connection 160 in the channel dimension.

[0128] In some embodiments, the neural network-based base caller 100 in the system 1000 is a cascaded model comprising a first sub-model and a second sub-model, and the first sub-model and the second sub-model are trained independently using non-overlapping training data, respectively.

[0129] In some embodiments, training the neural network-based base caller 100 or the semantic segmentation model 200 therein based on the training data comprises calculating gradients of a loss function using backpropagation to guide random gradient descent to update model parameters to minimize the loss function, wherein the loss function is a hybrid loss function comprising at least one of a cross-entropy loss, a dice loss, a mean square error, and one or more custom loss functions, and the custom loss functions comprise a mismatch loss reflecting a proportion or content of bases that are inconsistent between a prediction result and a ground truth.

[0130] In some embodiments, the hybrid loss function comprises the cross-entropy loss, the dice loss, and the mismatch loss, and a weight ratio of the three is 1:3:4, 1:4:4, 1:5:4, 1:6:4, 1:7:4, or 1:7:5.

[0131] In certain embodiments, the neural network-based base caller 100 or the semantic segmentation model 200 therein is trained by supervised learning or semi-supervised learning, and the training is performed with bases on a reference sequence corresponding to a base calling result of a subamplicon as ground truth.

[0132] In some specific examples, it is found that m1 being equal to or greater than m2 is conducive to further reducing mode errors, especially when the neural network-based base caller is a cascaded model.

[0133] In some embodiments, the system 1000 further comprises an input data acquisition unit 1008 configured to acquire input data of a subamplicon.

[0134] In some embodiments, the surface presents a regular array of wells, the surface comprises marker regions and reaction regions, spatial positions of the marker regions and the reaction regions are known, and signals from the marker regions present specific features on the images, and a plurality of amplicons are respectively attached in a plurality of separated wells of the reaction regions. In the input data obtaining unit 1008, a template is constructed by processing at least one of the images acquired in the first round of sequencing, and the input data is obtained based on the template, including: identifying the features on the images corresponding to the marker regions to determine the positions of the marker regions, and determining the position information of the amplicons based on the positions of the marker regions to obtain the template; and mapping the position information of the amplicons on the template to the images acquired in each of the specified rounds of sequencing respectively to determine the position information and intensity information of the amplicons on the images acquired in each of the specified rounds of sequencing respectively.

[0135] In some other embodiments, in the input data obtaining unit 1008, a template is constructed by processing the images acquired in the first round or the first p1 rounds of sequencing, and the input data is obtained based on the template, p1 being a natural number greater than 0 and less than or equal to 4, including: identifying bright spots on each of the images corresponding to the amplicons on which the specified biochemical reactions occur, and merging the bright spots to obtain the template; and mapping the position information of the amplicons on the template to the images acquired in each of the specified rounds of sequencing respectively to determine the position information and intensity information of the amplicons on the images acquired in each of the specified rounds of sequencing respectively.

[0136] In some other embodiments, in the input data obtaining unit 1008, a template is constructed by processing the images acquired in the first p2 rounds of sequencing, and the input data is obtained based on the template, p2 being a natural number greater than or equal to 6, including: detecting the bases introduced to each unit position on each of the images based on the intensity of each unit position to determine the detected base sequence of each of the unit positions, each unit position having a size less than or equal to the size of one pixel of the images, and each feature on the images corresponding to the amplicons on which the specified biochemical reactions occur occupying one or more unit positions; clustering the unit positions based on the similarity of the detected base sequences on the unit positions, including grouping a plurality of unit positions with a plurality of base detection sequences with a similarity satisfying a preset level as corresponding to the same amplicon to obtain the template, the template comprising the position information of the amplicons on the surface; and mapping the position information of the amplicons on the template to the images acquired in each of the specified rounds of sequencing respectively to determine the position information and intensity information of the amplicons on the images acquired in each of the specified rounds of sequencing respectively.

[0137] In some embodiments, the intensity information is corrected intensity information. For example, the intensity value after at least one of crosstalk correction and / or phase correction.

[0138] In some embodiments, the file size of the input data is less than or equal to half of the file size of the corresponding image. For example, the file size of the input data is less than or equal to one third, one fourth, or even one fifth of the file size of the corresponding image, particularly for patterned surfaces (regular surfaces).

[0139] In some specific examples, the input data is represented as a plurality of matrices, each matrix including a plurality of rows and columns, each matrix reflecting position and intensity information of a field of view amplicon on an image acquired in a round of sequencing, each element of the matrix corresponding to an amplicon, the row and column positions of the element of the matrix reflecting position information of the corresponding amplicon on the corresponding image, and the value of the element of the matrix reflecting intensity information of the corresponding amplicon on the image of the round of sequencing.

[0140] In some preferred examples, each matrix reflects information of a field of view amplicon on a block of an image of a round of sequencing, and the input data further includes matrices reflecting position information of each block on the image from which the block is derived.

[0141] A computer program product of an embodiment of the present application includes instructions, when executed by a computer, to perform the method of any of the above embodiments.

[0142] A computing device of an embodiment of the present application includes a memory for storing data, including a computer executable program, and one or more processors for executing the computer executable program to implement the method of any of the above embodiments. In some examples, the processor includes a GPU.

[0143] A computer readable storage medium of an embodiment of the present application stores a program for execution by a computer, execution of the program including completion of the method of any of the above embodiments. The computer readable storage medium includes read only memory, random access memory, magnetic or optical disks, and the like.

[0144] The embodiments of the present application also provide a method for training a base caller based on deep learning, which comprises training a neural network single model or a cascade model containing a neural network by using training data with the same content and data form as the input data in any of the above embodiments, so as to obtain a trained neural network. The base calling result of the training data is known and can be determined by a conventional base calling method. The features of the input data, the model training method or mode in any of the above base calling method embodiments are applicable to the training method of the base caller based on deep learning, unless otherwise specified. The method for training the base caller based on deep learning can obtain the neural network-based base caller in any of the above embodiments.

[0145] Figure 8 The flowchart for constructing a neural network-based base caller is shown. Scheme 1 shows the construction flow of a single model, and schemes 2.1 and 2.2 show the construction flow of two cascade models.

[0146] In some embodiments, an early stopping mechanism is introduced, for example, if the training does not achieve better results within 30 or 50 epochs (training periods), the training is stopped, the optimal training result parameters are saved, the robustness of the network is improved, and overfitting is prevented.

[0147] The following further example describes the specific operation mode or implementation process that can be included in some of the above embodiments. However, this should not be understood as limiting the scope of the subject matter of the present application to only the following further example description or embodiments.

[0148] Model training

[0149] The following example introduces the modeling steps of the model. It is assumed here that the model needs to consider the information of m adjacent contexts of each base, where m is freely selected according to the characteristics of the platform. Here, the array chip sequencing scene of four-color fluorescence imaging is taken as an example to introduce two different modeling schemes.

[0150] Data extraction and preprocessing can include the following steps:

[0151] 1) Select a suitable library and reference genome. The preferred genome is as large as possible, as comprehensive as possible to cover the application samples of the platform and as much as possible to exclude polymorphic sites, and sequencing is completed on the array chip. The total number of sequencing rounds generally needs to be consistent with the longest sequencing round of the corresponding instrument in actual application. In order to obtain more comprehensive training data, sequencing is preferably repeated multiple times on different instruments of the same type by different operators using different batches of reagents and chips to accumulate the training data set.

[0152] Take four-color fluorescent sequencing as an example, each field of view (FOV) will obtain 4 pictures in each round of sequencing, which are A, G, C, T four fluorescence channel under the fluorescent image.

[0153] 2) Use existing or conventional known base recognition tools to analyze and recognize the base of the sequencing data, and obtain the base recognition result, that is, the fastq file recording all the fluorescent cluster tags (ID) and their corresponding base sequences.

[0154] 3) Align the above fastq file with the reference genome to obtain the alignment result, that is, the bam file or sam file.

[0155] 4) For the nth round of sequencing, the target value is the base recognition "correct answer" of each fluorescent cluster corresponding to the nth round of sequencing, that is, the corresponding base on the reference genome after alignment. The base recognition correct answer corresponding to each fluorescent cluster is arranged into a matrix according to the actual spatial position of the fluorescent cluster, that is, the target value matrix. For example, assuming that the fluorescent cluster in the 2nd row and the 8th column of the chip should be A in the nth round of sequencing, then A is written in the [1, 7] coordinate position (counting from 0) of the target value matrix. It should be noted that in order to ensure the reliability of the training data, sequences with poor sequencing quality or unreliable alignment positions can be excluded. For sequences that fail to successfully align to the reference genome or have unreliable alignment results, the corresponding fluorescent cluster position in the target value matrix can be recorded as "N", and the N position is not involved in the subsequent training.

[0156] 5) According to the known array chip design layout, the center positions of all holes (i.e. the regions where the fluorescent clusters are located) on each fluorescent picture in each round are located, and the fluorescent intensity of the center position of each fluorescent cluster is obtained from the original fluorescent image according to the above positioning. In the four-color fluorescent sequencing scenario, each fluorescent cluster corresponds to four fluorescent intensity values of A, G, C, and T in each round of sequencing. Here, the fluorescent intensity of the original fluorescent picture corresponding to the center position of each fluorescent cluster is called "original fluorescent brightness".

[0157] 6) Optional operation: according to the characteristics of the original fluorescent brightness distribution on each picture, the original fluorescent brightness is subjected to background removal operation to remove the influence of the background of the image. Here, the fluorescent brightness after background removal is called "background-removed fluorescent brightness".

[0158] 7) Optional Operations: Based on the background-free fluorescence brightness distribution characteristics of rounds n-1 and n, perform phasing correction on the background-free fluorescence brightness of round N; based on the background-free fluorescence brightness distribution characteristics of rounds n and n+1, further perform prephasing correction; then, based on the crosstalk parameters of the fluorescence signals of the four fluorescence channels (determined by the spectra of the fluorescent molecules and the parameters of the optical instruments), correct the signal crosstalk of the fluorescence channels. The final background-free fluorescence brightness after phasing, prephasing, and fluorescence crosstalk correction is called the "corrected fluorescence brightness".

[0159] 8) The brightness matrix input to the deep learning model as a feature parameter can be any one of the following: original fluorescence brightness, background-removed fluorescence brightness, or corrected fluorescence brightness. Here, we refer to it collectively as "fluorescence brightness." Each fluorescence cluster will obtain a fluorescence brightness value in each fluorescence image. For each fluorescence image, the fluorescence brightness corresponding to each fluorescence cluster is arranged into a matrix according to the actual spatial position of the cluster. For example, the fluorescence cluster in the 2nd row and 8th column on the chip has its corresponding fluorescence brightness written in the [1,7] coordinates of the matrix (counting from 0). Therefore, for each fluorescence image, a matrix with the same number of rows and columns as the number of rows and columns of holes on the chip surface will be obtained; this is called the "fluorescence brightness matrix."

[0160] Single model

[0161] 1. The training dataset can be established by referring to the above steps. Taking the sequencing data of the nth round of a certain FOV as an example, the target value of this data is the target value matrix of the nth round of the FOV. The feature values ​​can be all fluorescence brightness matrices from the n-m1th to the n+m2th round, a total of (m1+m2+1)*4 (including the m1+m2+1th round, with 4 fluorescence channels in each round). m1 and m2 can be equal or not equal. The following example uses m1=m2=m.

[0162] 2. Use the training dataset described above to build a deep learning model. The final model's performance is evaluated by assessing the error rate of the output sequence.

[0163] 3. The model building process involves numerous data selection and model parameter tuning methods. Among these, the parts most closely related to sequencing and significantly impacting the accuracy of model predictions after testing include:

[0164] When constructing the training dataset, it is best to choose data with high confidence in the alignment results for training. For example, when constructing the target value matrix in 4) above, if the corresponding sequence in the sam or bam file has insertions or missing values, it may be suspected that the alignment position is not reliable enough. In this case, the target value matrix values ​​of the entire corresponding sequence can be written as 'N'.

[0165] In the subsequent training process, the positions with target value N can not participate in the training. For example, when calculating the loss function, the positions corresponding to N can not be considered.

[0166] The loss function needs to be optimized according to the specific model. Tests have found that considering parameters for evaluating sequencing quality in the sequencing scenario helps to obtain a better model. For example, when calculating the loss function, in addition to considering the traditional loss function calculation method, parameters specific to the sequencing scenario can also be added, such as error rate of sequencing, mode error characteristic parameters, etc.

[0167] For sequencing with a large number of rounds, the sequencing quality and data characteristics may differ greatly in different stages (e.g., when the number of rounds is small in the early stage and when the number of rounds is large in the later stage). Different models can be established for different round segmentation intervals. For example, model A can be used for 0-50 rounds, model B can be used for 51-100 rounds, and so on.

[0168] The single model training of the embodiments of the present application is also sometimes referred to as scheme 1, and the cascade model training is referred to as scheme 2.

[0169] Cascade model

[0170] 1. The main idea of scheme 2 is to connect two or more sub-models in series. The following takes the case of a cascade model with two sub-models in series as an example. Specifically, according to the parts that the two sub-models mainly focus on or the main influence on the prediction results, the two sub-models are sometimes also referred to as a spatial crosstalk correction model and a context correction model. The order of the two sub-models is not limited, and the spatial crosstalk can be corrected first, or the context influence can be corrected first. Generally, the spatial crosstalk correction model is a neural network, and the context influence correction model can be a traditional machine learning such as a tree model or a neural network. Here, the training of a cascade model with a spatial crosstalk correction model (model 1) first and a context influence correction model (model 2) second is taken as an example for description, which is scheme 2.1 of Figure 8 .

[0171] 2. Construct the training data set of model 1. One piece of training data corresponds to the sequencing data of one FOV and one round (assuming the nth round). The target value of this piece of data is the target value matrix of the nth round of the FOV, and the feature value is all the fluorescence intensity matrices of the nth round, a total of 4.

[0172] 3. Based on the training data set of model 1, a deep learning model 1 is established. Model 1 outputs the base recognition result matrix of each round of sequencing, and also outputs the probability of base classification calculated by the model at each position of each round, i.e., the probability of each position being A, G, C, or T base. The above base classification probability is also output in the form of a matrix, and the matrix coordinates correspond one-to-one to the spatial positions of the fluorescence clusters.

[0173] 4、Build the training dataset of model 2. One piece of data corresponds to the relevant data information of a certain fluorescent cluster position in a certain round (assuming the nth round) of sequencing. The target value is the reference base of the fluorescent cluster in the nth round (i.e. the "standard answer" after alignment), and the feature value is the base classification probability value corresponding to the nth-m to nth+m round, a total of (2m+1)*4 values. In addition, the feature value can also include round information (i.e. n value), 4 base classification probability size distribution relationship related parameters, etc. It can be understood that when training model 2, the input of model 2 is the output of the trained model 1.

[0174] 5、Based on the training dataset of model 2, train model 2. Model 2 can be a deep learning model or a machine learning model. Model 2 outputs the final base recognition result of each fluorescent cluster in each round.

[0175] 6、The order of model 1 and model 2 can also be exchanged for modeling, that is, first build and train model 2, and then train model 1; only the fluorescence intensity and base classification probability in the above steps need to be exchanged and re-modeled, that is, the feature value input to model 1 when training model 1 is the output of the trained model 2. Please refer to Figure 8 Solution 2.2.

[0176] It should be noted that training model 1 and model 2 separately using completely independent and disjoint training data can prevent potential overfitting and help obtain a more practical and universal cascade model.

[0177] Model prediction (base recognition)

[0178] The model calling method corresponding to the single model and cascade model in the above example please refer to Figure 5 In some sequencing scenarios, the input data is the sequencing round and the fluorescence intensity matrix. Therefore, in the case where the data before and after multiple rounds need to be considered, only the fluorescence intensity matrix before and after multiple rounds needs to be retained, and the original picture does not need to be retained. This can greatly save the storage requirements of the online analysis (sequencing while analyzing) system and improve the calculation speed of online analysis.

[0179] Figure 9 The process of using the trained U-Net to process input data to obtain base recognition results is shown.

[0180] In addition, for the patterned surface / chip with marked areas, the marked areas divide the reaction area into multiple blocks. Under a certain data processing calculation level, it is found in some tests that if the matrix corresponding to the full image is input at one time, the input matrix is too large, and the existence of cross points between blocks affects the training speed of the network and the accuracy of the training results. Therefore, the full image matrix can be selected according to the actual block size, cut into a matrix of corresponding size, and additional position coding can be given according to the position of the block in the full image to reflect the position information.

[0181] For example, please refer to Figure 6 The matrices of different sizes are all zero-padded to the same size as input, and the entire image is converted into a total of 3*3, i.e., 9 blocks, with different lengths and widths of the blocks; after splitting into 9 blocks, the edges of all blocks with a size less than b*y are padded with 0 values to a size of b*y; finally, an additional position coding layer is added according to the position number of the block in the full image, for example, the block at the lower right corner is numbered 9, and a layer of Pos(P) position coding is filled, with all values being 9, to obtain the final input matrix. The obtained output result is cropped to the original size according to the matrix information and spliced to obtain the base call result of the full image.

[0182] Similar processing can also be performed when constructing training data for model training, for example, position coding is added during data preprocessing to reflect the position information of the current block in the entire image. By introducing such position information, the differences in light intensity of different block positions caused by the optical system can be further eliminated or reduced, and the network can overcome these differences during training.

[0183] Embodiment 1

[0184] This embodiment verifies the potential of the deep learning model to correct spatial fluorescence crosstalk and does not consider the influence of context on the current base call. This embodiment uses a fixed area in a high-density array chip sequencing image, uses a four-color fluorescence second-generation sequencing process, P150 (150+150 double-end sequencing), the fluorescence intensity matrix uses the "corrected fluorescence intensity" matrix, and the training data set sample is a human genome sample.

[0185] In this embodiment, each training data corresponds to the sequencing information corresponding to a certain FOV and a certain round of sequencing. The training data is 1000 FOV array chip sequencing data, P150 (150+150 double-end sequencing), and is derived from different experiments, using different reagents and chip batches.

[0186] Network structure

[0187] The network structure based on U-Net is added with spatial attention and channel attention mechanisms, as shown in Figure 10 After the data X of 32*4*400*320(B: batch size, C: channel, H: height, and W: width) is input, the following processing procedures are experienced:

[0188] 1) First, an in_conv layer and a CBAM layer are experienced, and the structures thereof are shown in Figure 11 、 Figure 12 The in_conv layer includes a 3*3 convolution, which changes the size to 32*32*400*320, and then normalization and ReLU activation functions are experienced, and finally a 3*3 convolution, normalization and ReLU activation functions are experienced, and finally X0 of 32*32*400*320 is obtained.

[0189] Then, the CBAM layer is input, and in the channel attention of the CBAM layer, data X0 is respectively experienced a global maximum pooling layer and a global average pooling layer, and then a shared parameter neural network MLP. In the MLP, the data is first experienced a 1*1 convolution, the channel number is changed to 1 / 16 of the original, and then an activation function ReLU is experienced, and finally a 1*1 convolution is experienced to restore the channel number. Finally, two data are added and a Sigmoid activation function is experienced to obtain X 0C , which is element-wise multiplied with X0 to obtain the final output X 0MID .

[0190] For ease of understanding, an example is given. For example, the input is (N, C, H, W), and the output of the global pooling is (N, C), and the specific value is the mean and maximum value of each channel. After the input features are pooled, the result is two one-dimensional vectors, and the two vectors are added after MLP operation to obtain a one-dimensional channel attention. Finally, the attention is multiplied with the input features to obtain the final output. It can be understood that the weights of each channel are calculated and assigned to the input features.

[0191] In the spatial attention of the CBAM, data X 0C is based on the channel to obtain a global pooling and a global maximum, and then the two results obtained are spliced based on the channel, compressed to 1 channel after a 7*7 convolution, and finally a Sigmoid activation function is experienced to obtain X 0S , which is element-wise multiplied with X 0MID to obtain the final CBAM output X 0FIN .

[0192] 2) secondly, 4 layers of down-sampling layers, the first three layers each include a CBAM layer and a Down layer, and the last layer only has a Down layer, and the Down layer structure is as shown in Figure 13 . Each time a Down layer is passed, the number of channels C of the data is doubled (except for the last Down), and the height H and width W are halved.

[0193] Take the first layer as an example (input X 0FIN , output X 1FIN ): X 0FIN is input into the Down layer, and after a 2*2 max pooling layer with a stride of 2, the data becomes 32*32*200*160, and then a 3*3 convolution is performed, and the size becomes 32*64*200*160, and then normalization and ReLU activation functions are performed, and finally a 3*3 convolution, normalization and ReLU activation function are performed, and finally 32*64*200*160 X1 is obtained.

[0194] X1 is input into the CBAM, and after the steps described in 1), the final output X 1FIN is obtained. 1FIN :32*64*200*160, X 2FIN :32*128*100*80, X 3FIN :32*256*50*40, X4:32*256*25*20.

[0195] 3) then, 4 layers of up-sampling layers (Up), as shown in Figure 14 . Each layer of up-sampling has two inputs, one is the output result of the previous layer (the first down-sampling layer is X4, and the second layer is the output Y1 of the first layer), and the other is the corresponding down-sampling result (the first layer corresponds to X 3FIN , and the last layer corresponds to X 0FIN ). Each time a up-sampling layer is passed, the number of channels C of the data is halved (except for the last up-sampling), and the height H and width W are doubled. Take the first layer of up-sampling as an example (input X4, X 3FIN , output Y1):

[0196] The input X4 is bilinearly interpolated to become 32*256*50*40, and then X4 is concatenated with X 3FIN according to the channel dimension C, and the result is 32*512*50*40. Then a 3*3 convolution is performed, and the size becomes 32*256*50*40, and then normalization and ReLU activation functions are performed, and finally a 3*3 convolution, normalization and ReLU activation function are performed, and finally 32*128*50*40 Y1 is obtained.

[0197] Y1 and X 2FIN The next up-sampling layer is input to obtain the output result Y2, and so on until the output result Y4 is obtained.

[0198] 4) Finally, Y4 is input into a 1*1 convolution out_conv, and the output channel number is the classification number, to obtain the final prediction result Out: 32*5*400*320.

[0199] Input description: ACGT images collected by the camera under the same fov every Cycle, the hole coordinate information and light intensity information of the center block of each image are extracted, a two-dimensional matrix with the format (High: 381, Width: 304) is generated according to the hole position, each element in the matrix represents the light intensity value of the corresponding hole, and four matrices are combined into a three-dimensional matrix (4, High: 381, Width: 304) as training data. When the training data is input, the light intensity value will be standardized to the range of 0-255 first, and then changed to 4*400*320 after padding to input the network model for training (the padding method is edge padding, which expands the right and lower sides of the original matrix to the specified size). Through the calculation of a large amount of data, the mean (0.260, 0.275, 0.273, 0.267) and variance (0.219, 0.210, 0.211, 0.213) of the input matrix are used to standardize the input data, and the mean of the standardized data is 0 and the standard deviation is 1.

[0200] The label of the data comes from the base information compared with the reference genome. ACGT is numbered according to 1234, input into a 381*304 matrix according to the hole position, and then expanded to 400*320 after the same edge padding processing.

[0201] Output description: The output result is a two-dimensional matrix (400*320), and the value (1, 2, 3, 4) of each element represents the base (A, C, G, T) recognition result of the corresponding hole.

[0202] When measuring the final training effect, the right lower corner area filled during training is cropped from the output result, and the size of the cropped area is 381*304. The points with a value of 0 in the label made by the reference genome are excluded, and the points in the output result are excluded according to the same position information, and then the difference value ratio between the two is calculated to obtain the final base recognition error rate.

[0203] Training method description: The stochastic gradient descent algorithm (SGD) is used as the optimizer during training, with an initial learning rate of 0.05, a momentum of 0.9, a weight decay of 1e-4, a batch size of 32, a total of 300 epochs of training, and the optimal training parameters are saved. During the training process, the learning rate is updated every step. The loss function used during backpropagation is a hybrid function of cross-entropy loss, dice loss, and a custom miss_match loss, with a weight ratio of, for example, 2:2:1 or 1:5:3. When calculating the loss, the weight of the background (pixels with a value of 0) is set to 0 and does not participate in the loss calculation, and only the loss function size of the ACGT four types is calculated. An early stopping strategy is set during training to prevent network overfitting. When the minimum loss of the validation result has not been updated for 30 epochs, training is stopped and the model parameters of the epoch with the smallest loss are saved.

[0204] The test data is an array of 100 FOV chip sequencing data, PE150 (double-end sequencing 150+150), from different experiments, using different reagents and chip batches, and the samples are all human genome samples. After using the above model, the overall sequencing error rate is reduced by 28%, and the alignment rate is increased by 1.2%.

[0205] Example 2

[0206] This example is used to verify the potential of deep learning or machine learning models to correct the influence of context on the identification of the current base, without considering the influence of spatial fluorescence crosstalk on base identification, i.e. the spatial arrangement information of each fluorescence cluster is not input into the model for training. This example uses a fixed area in a high-density array chip sequencing image, uses a four-color fluorescence second-generation sequencing process, PE150 (double-end sequencing 150+150), and uses a "corrected fluorescence intensity" matrix for the fluorescence intensity matrix. The training data set samples are human genome samples.

[0207] Model establishment: The training data is in units of individual fluorescence clusters, and each training data represents the sequencing information corresponding to a certain sequencing round of a certain fluorescence cluster, which can not include spatial position information of the fluorescence cluster. The training data is 80 FOV array chip sequencing data, P150 (150+150 double-end sequencing), from different experiments, using different reagents and chip batches.

[0208] Feature value: The corrected fluorescence intensity values corresponding to the current sequencing round and the previous and next three sequencing rounds of the fluorescence cluster, as well as the number of sequencing rounds corresponding to the current round. Target value: The aligned reference sequence base corresponding to the current sequencing round of the fluorescence cluster, i.e. the "standard answer" of the current sequencing round. Here, a lightGBM machine learning model is established.

[0209] Model effect: The test data is 20 FOV array chip sequencing data, PE150 (double-end sequencing 150+150), from different experiments, using different reagents and chip batches, and the samples are all human genome samples. After using the above model, the overall error rate of sequencing is reduced by 37%, and the alignment rate is increased by 2.3%.

[0210] Example 3

[0211] This example is to verify the effect of the complete algorithm / model. A fixed area in a high-density array chip sequencing picture is used, using a four-color fluorescent second-generation sequencing process, PE150 (double-end sequencing 150+150), and the fluorescence intensity matrix uses the "corrected fluorescence intensity" matrix. The training data set samples are human genome samples.

[0212] Model building: The modeling method and process can refer to the description in the previous model training part, and the specific can refer to the description in the solution 2.1 part. Figure 8

[0213] The modeling method of model 1 can refer to the related description of example 1. Each training data corresponds to the sequencing information corresponding to a certain round of sequencing of a certain FOV. The training data is 1000 FOV array chip sequencing data, P150 (150+150 double-end sequencing), from different experiments, using different reagents and chip batches.

[0214] The modeling method of model 2 can refer to the related description of example 2. The training data is in single fluorescent cluster units, and each training data represents the sequencing information corresponding to a certain round of sequencing of a certain fluorescent cluster, which does not contain the spatial position information of the fluorescent cluster. The training data is 80 FOV array chip sequencing data, P150 (150+150 double-end sequencing), from different experiments, using different reagents and chip batches. The 80 FOV training data used by model 2 does not overlap with the 1000 FOV training data used by model 1. In addition, the training data (sequencing data) of model 2 needs to be processed by the trained model 1 to obtain the base classification probability value, and then the base classification probability value corresponding to each fluorescent cluster is input into model 2 as a feature parameter for training. The construction of model 2 uses a lightGBM model.

[0215] Model effect: The test data is 20 FOV array chip sequencing data, PE150 (double-end sequencing 150+150), from different experiments, using different reagents and chip batches, and the samples are all human genome samples. After using the above model, the overall error rate of sequencing is reduced by 37%, and the alignment rate is increased by 2.3%. ​

[0216] In the description of the specification, the description "one embodiment", "some embodiments", "exemplary embodiment", "example", "some examples", "specific example" or "embodiment" etc. means that the particular feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the application. Descriptive expressions of the above terms in the specification do not necessarily refer to the same embodiment or example. Moreover, the described particular feature, structure, material or characteristic can be combined in any appropriate manner in one or more embodiments or examples.

[0217] In addition, each functional unit in each embodiment of the specification can be integrated in one processing module, or each unit can exist physically separately, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0218] Although the embodiments of the application have been shown and described above, it should be understood that the above embodiments are exemplary and are not to be construed as limiting the application, and those of ordinary skill in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the application.

Claims

1. A base recognition method, characterized in that, include: The input data of multiple amplicons is processed by a neural network-based base recognizer to generate an alternative representation of the input data. The input data includes the position and intensity information of each amplicon on the images acquired in each sequencing round from n-m1 to n+m2. The sequencing adopts a multi-fluorescent channel sequencing technology based on surface imaging. The amplicon is a nucleic acid molecule to be tested containing multiple identical polynucleotide molecules. The amplicon is attached to the surface. The input data is presented in a non-image form. The file size of the input data is smaller than the file size of its corresponding image. n is a natural number greater than or equal to 1 and n is greater than m1. m1 and m2 are each independent integers greater than or equal to 0. The substitution representation is processed by an output layer to produce an output, wherein the output identifies the probability that the bases introduced into each of the amplicons in the nth round of sequencing are A, C, T, and G; and Based on the output, the base types introduced into each of the amplicon in the nth round of sequencing are determined.

2. The method according to claim 1, characterized in that, n is greater than or equal to 2, and m1 and m2 are each independently greater than or equal to 1.

3. The method according to claim 2, characterized in that, The neural network-based base recognizer is a trained semantic segmentation model, which is selected from at least one of a network with an encoder-decoder structure or a transformer-based network.

4. The method according to claim 2, characterized in that, The neural network-based base recognizer is a trained cascaded model, which includes a first sub-model and a second sub-model connected in series. The first sub-model and the second sub-model are each independently selected from a semantic segmentation model and a machine learning model. The semantic segmentation model is selected from at least one of a network with an encoder-decoder structure or a transformer-based network. The machine learning model is selected from at least one of a tree model.

5. The method according to claim 4, characterized in that, The method includes: The first sub-model processes the input data of each sequencing round from the n-m1th round to the n+m2th round of the amplicon to generate a first prediction result for a specified round, wherein the first prediction result for a specified round includes the probability that the bases introduced into each amplicon in that round of sequencing are A, C, T and G. The second sub-model processes the first prediction results of each specified round and the round information of each specified round to generate the output.

6. The method according to claim 5, characterized in that, The first sub-model is a semantic segmentation model.

7. The method according to claim 4, characterized in that, The method includes: The input data from round n-m1 to round n+m2 of a specified amplicon are processed by the first sub-model to generate a second prediction result for each specified round of the amplicon. The second prediction result for each specified round of the amplicon includes the probability that the bases introduced into the amplicon in each specified round of sequencing are A, C, T and G. The output is generated by processing the second prediction results of each specified round of each specified amplicon through the second sub-model.

8. The method according to claim 7, characterized in that, The second sub-model is a semantic segmentation model.

9. The method according to any one of claims 4, 6, and 8, characterized in that, The semantic segmentation model is selected from at least one of U-Net, DeepLab, SegFormer, PSPNet, FCN, GCN and their respective variants.

10. The method according to any one of claims 4, 5, and 7, characterized in that, The machine learning model is a gradient boosting decision tree (GBDT), which employs at least one of the LightGBM and XGBoost algorithms.

11. The method according to claim 9, characterized in that, The semantic segmentation model is selected from one of the following (i)-(iv): (i) is U-Net or SegFormer; (ii) For U-Net or SegFormer with added attention mechanisms; (iii) For U-Net with CBAM or SE modules added to the encoding section; (iv) U-Net incorporates a transformer-based multi-head attention mechanism in the decoding and skip connection parts.

12. The method according to any one of claims 4, 6, 8 and 11, characterized in that, The neural network-based base recognizer or the semantic segmentation model includes: An encoder containing multiple downsampling layers, each downsampling layer containing at least one pooling operation and at least one convolution operation, and at least some downsampling layers have added CBAM or SE modules; A decoder comprising multiple upsampling layers, wherein the number of upsampling layers is the same as the number of downsampling layers, and each upsampling layer includes at least one deconvolution operation or bilinear interpolation operation; and A skip connection that connects the downsampling layer and the corresponding upsampling layer.

13. The method according to claim 9, characterized in that, The neural network-based base recognizer or the semantic segmentation model includes: An encoder containing multiple downsampling layers, each downsampling layer containing at least one pooling operation and at least one convolution operation, and at least some downsampling layers have added CBAM or SE modules; A decoder comprising multiple upsampling layers, wherein the number of upsampling layers is the same as the number of downsampling layers, and each upsampling layer includes at least one deconvolution operation or bilinear interpolation operation; and A skip connection that connects the downsampling layer and the corresponding upsampling layer.

14. The method according to claim 12, characterized in that, The encoder further includes an input convolutional layer, the input data of which is processed by the input convolutional layer to produce an alternative representation as the input of the encoder's first downsampling layer, the input convolutional layer containing multiple convolution operations; and / or, the decoder further includes an output convolutional layer, the output convolutional layer processing the alternative representation produced by the decoder's last upsampling layer to produce the output.

15. The method according to claim 13, characterized in that, The encoder further includes an input convolutional layer, the input data of which is processed by the input convolutional layer to produce an alternative representation as the input of the encoder's first downsampling layer, the input convolutional layer containing multiple convolution operations; and / or, the decoder further includes an output convolutional layer, the output convolutional layer processing the alternative representation produced by the decoder's last upsampling layer to produce the output.

16. The method according to claim 12, characterized in that, The neural network-based base recognizer or the segmentation model also has one or more of the following technical features (a)-(e): (a) The encoder comprises four downsampling layers; (b) Each of the downsampling layers includes a CBAM module or an SE module and a Down structure connected to the CBAM module or SE module, the Down structure including one pooling operation and two convolution operations; (c) The CBAM module includes a connected channel attention module and a spatial attention module. The channel attention module includes a first pooling unit, an MLP unit and a first output layer. The spatial attention module includes a second pooling unit, a convolution unit and a second output layer. In the channel attention module, the input is processed by the global max pooling and global average pooling of the first pooling unit to generate two channel descriptors. The two channel descriptors are processed by the shared MLP in the MLP unit to obtain channel weights. Then, the first output layer outputs the result of weighting the input using the channel weights. In the spatial attention module, the output from the first output layer is processed by the global max pooling and global average pooling of the second pooling unit to generate two spatial descriptors. The two spatial descriptors are processed by the convolutional layer of the convolutional unit to obtain spatial weights. Then, the output of the second output layer is output using the spatial weights to weight the output of the first output layer. (d) The SE module includes a channel attention module, which includes a third pooling unit, a fully connected unit, and a third output layer. An input channel is generated as a corresponding channel descriptor after global average pooling by the third pooling unit. The channel descriptors are processed by two fully connected layers and an activation function of the fully connected unit to obtain channel weights. Then, the third output layer outputs the result of weighting the input using the channel weights. (e) Each of the upsampling layers includes a bilinear interpolation or a deconvolution operation, and a concatenation operation that concatenates the output of the previous layer with the result of the corresponding downsampling layer passed to the upsampling layer via a skip connection in the channel dimension.

17. The method according to any one of claims 13-15, characterized in that, The neural network-based base recognizer or the segmentation model also has one or more of the following technical features (a)-(e): (a) The encoder comprises four downsampling layers; (b) Each of the downsampling layers includes a CBAM module or an SE module and a Down structure connected to the CBAM module or SE module, the Down structure including one pooling operation and two convolution operations; (c) The CBAM module includes a connected channel attention module and a spatial attention module. The channel attention module includes a first pooling unit, an MLP unit and a first output layer. The spatial attention module includes a second pooling unit, a convolution unit and a second output layer. In the channel attention module, the input is processed by the global max pooling and global average pooling of the first pooling unit to generate two channel descriptors. The two channel descriptors are processed by the shared MLP in the MLP unit to obtain channel weights. Then, the first output layer outputs the result of weighting the input using the channel weights. In the spatial attention module, the output from the first output layer is processed by the global max pooling and global average pooling of the second pooling unit to generate two spatial descriptors. The two spatial descriptors are processed by the convolutional layer of the convolutional unit to obtain spatial weights. Then, the output of the second output layer is output using the spatial weights to weight the output of the first output layer. (d) The SE module includes a channel attention module, which includes a third pooling unit, a fully connected unit, and a third output layer. An input channel is generated as a corresponding channel descriptor after global average pooling by the third pooling unit. The channel descriptors are processed by two fully connected layers and an activation function of the fully connected unit to obtain channel weights. Then, the third output layer outputs the result of weighting the input using the channel weights. (e) Each of the upsampling layers includes a bilinear interpolation or a deconvolution operation, and a concatenation operation that concatenates the output of the previous layer with the result of the corresponding downsampling layer passed to the upsampling layer via a skip connection in the channel dimension.

18. The method according to any one of claims 4-8, 11, and 13-16, characterized in that, The first sub-model and the second sub-model are trained independently using non-overlapping training data.

19. The method according to claim 9, characterized in that, The first sub-model and the second sub-model are trained independently using non-overlapping training data.

20. The method according to claim 12, characterized in that, The first sub-model and the second sub-model are trained independently using non-overlapping training data.

21. The method according to claim 17, characterized in that, The first sub-model and the second sub-model are trained independently using non-overlapping training data.

22. The method according to claim 12, characterized in that, Training the neural network-based base recognizer or the semantic segmentation model based on training data includes calculating the gradient of the loss function using backpropagation to guide stochastic gradient descent to update model parameters in order to minimize the loss function, wherein... The loss function is a hybrid loss function that includes at least one of cross-entropy loss, Dice loss, mean squared error, and one or more custom loss functions. The custom loss function includes mismatch loss, which reflects the proportion or content of bases that are inconsistent between the predicted result and the actual result.

23. The method according to any one of claims 13-16 and 19-21, characterized in that, Training the neural network-based base recognizer or the semantic segmentation model based on training data includes calculating the gradient of the loss function using backpropagation to guide stochastic gradient descent to update model parameters in order to minimize the loss function, wherein... The loss function is a hybrid loss function that includes at least one of cross-entropy loss, Dice loss, mean squared error, and one or more custom loss functions. The custom loss function includes mismatch loss, which reflects the proportion or content of bases that are inconsistent between the predicted result and the actual result.

24. The method according to claim 22, characterized in that, The hybrid loss function includes cross-entropy loss, dice loss, and mismatch loss, and the weight ratio of the three is 1:3:4, 1:4:4, 1:5:4, 1:6:4, 1:7:4, or 1:7:

5.

25. The method according to claim 23, characterized in that, The hybrid loss function includes cross-entropy loss, dice loss, and mismatch loss, and the weight ratio of the three is 1:3:4, 1:4:4, 1:5:4, 1:6:4, 1:7:4, or 1:7:

5.

26. The method according to claim 22, characterized in that, The neural network-based base recognizer or the semantic segmentation model is trained using supervised or semi-supervised learning methods, with the bases on the reference sequence corresponding to the base recognition result of the amplicon used as the real results for this training.

27. The method according to claim 23, characterized in that, The neural network-based base recognizer or the semantic segmentation model is trained using supervised or semi-supervised learning methods, with the bases on the reference sequence corresponding to the base recognition result of the amplicon used as the real results for this training.

28. The method according to any one of claims 1-8, 11, 13-16, 19-22, and 24-27, characterized in that, m1 is equal to or greater than m2.

29. The method according to claim 9, characterized in that, m1 is equal to or greater than m2.

30. The method according to claim 12, characterized in that, m1 is equal to or greater than m2.

31. The method according to claim 17, characterized in that, m1 is equal to or greater than m2.

32. The method according to claim 23, characterized in that, m1 is equal to or greater than m2.

33. The method according to any one of claims 1-8, 11, 13-16, 19-22, 24-27 and 29-32, characterized in that, The surface is a regular array of pores, comprising labeled regions and reaction regions. The spatial relationship between the labeled regions and the reaction regions is known, and signals from the labeled regions exhibit specific features in the image. Multiple amplicon groups are respectively connected to multiple separate pores in the reaction regions. A template is constructed by processing at least one image acquired in the first round of sequencing, and the input data is obtained based on the template. include, The template is obtained by identifying features on the image corresponding to the marked region to determine the location of the marked region, and determining the location information of the amplicon based on the location of the marked region; as well as The position information of the amplicon on the template is mapped onto the images acquired in each specified round of sequencing, so as to determine the position and intensity information of the amplicon on the images acquired in each specified round of sequencing.

34. The method according to claim 12, characterized in that, The surface is a regular array of pores, comprising labeled regions and reaction regions. The spatial relationship between the labeled regions and the reaction regions is known, and signals from the labeled regions exhibit specific features in the image. Multiple amplicon groups are respectively connected to multiple separate pores in the reaction regions. A template is constructed by processing at least one image acquired in the first round of sequencing, and the input data is obtained based on the template. include, The template is obtained by identifying features on the image corresponding to the marked region to determine the location of the marked region, and determining the location information of the amplicon based on the location of the marked region; as well as The position information of the amplicon on the template is mapped onto the images acquired in each specified round of sequencing, so as to determine the position and intensity information of the amplicon on the images acquired in each specified round of sequencing.

35. The method according to claim 17, characterized in that, The surface is a regular array of pores, comprising labeled regions and reaction regions. The spatial relationship between the labeled regions and the reaction regions is known, and signals from the labeled regions exhibit specific features in the image. Multiple amplicon groups are respectively connected to multiple separate pores in the reaction regions. A template is constructed by processing at least one image acquired in the first round of sequencing, and the input data is obtained based on the template. include, The template is obtained by identifying features on the image corresponding to the marked region to determine the location of the marked region, and determining the location information of the amplicon based on the location of the marked region; as well as The position information of the amplicon on the template is mapped onto the images acquired in each specified round of sequencing, so as to determine the position and intensity information of the amplicon on the images acquired in each specified round of sequencing.

36. The method according to claim 23, characterized in that, The surface is a regular array of pores, comprising labeled regions and reaction regions. The spatial relationship between the labeled regions and the reaction regions is known, and signals from the labeled regions exhibit specific features in the image. Multiple amplicon groups are respectively connected to multiple separate pores in the reaction regions. A template is constructed by processing at least one image acquired in the first round of sequencing, and the input data is obtained based on the template. include, The template is obtained by identifying features on the image corresponding to the marked region to determine the location of the marked region, and determining the location information of the amplicon based on the location of the marked region; as well as The position information of the amplicon on the template is mapped onto the images acquired in each specified round of sequencing, so as to determine the position and intensity information of the amplicon on the images acquired in each specified round of sequencing.

37. The method according to any one of claims 1-8, 11, 13-16, 19-22, 24-27 and 29-32, characterized in that, A template is constructed by processing images acquired in the first or pre-p1 sequencing rounds, and the input data is obtained based on the template, where p1 is a natural number greater than 0 and less than or equal to 4, including... Identify bright spots on each image corresponding to amplicon that has undergone a specified biochemical reaction, and merge these bright spots to obtain the template; as well as The position information of the amplicon on the template is mapped onto the images acquired in each specified round of sequencing, so as to determine the position and intensity information of the amplicon on the images acquired in each specified round of sequencing.

38. The method according to any one of claims 1-8, 11, 13-16, 19-22, 24-27 and 29-32, characterized in that, A template is constructed by processing images acquired in the first p2 rounds of sequencing, and the input data is obtained based on the template, where p2 is a natural number greater than or equal to 6, including... The detected base sequence at each unit position is determined by detecting the base introduced at that unit position based on the intensity at each unit position on each image in each round. The image contains multiple unit positions, and the size of a unit position is less than or equal to the size of one pixel of the image. The size of a feature on the image corresponding to an amplicon that has undergone a specified biochemical reaction occupies one or more unit positions. Clustering is performed on these unit positions based on the similarity of the detected base sequences, including grouping multiple unit positions of multiple base detection sequences with similarity levels meeting a preset level into those corresponding to the same amplicon, to obtain the template, the template containing the position information of these amplicones on the surface; and The position information of the amplicon on the template is mapped onto the images acquired in each specified round of sequencing, so as to determine the position and intensity information of the amplicon on the images acquired in each specified round of sequencing.

39. The method according to any one of claims 1-8, 11, 13-16, 19-22, 24-27, 29-32 and 34-36, characterized in that, The intensity information is the corrected intensity information.

40. The method according to claim 39, characterized in that, The correction is selected from at least one of crosstalk correction and / or phase correction.

41. The method according to any one of claims 1-8, 11, 13-16, 19-22, 24-27, 29-32, 34-36 and 40, characterized in that, The file size of the input data is less than or equal to half the file size of its corresponding image.

42. The method according to claim 41, characterized in that, The file size of the input data is less than or equal to one-third of the file size of its corresponding image.

43. The method according to any one of claims 1-8, 11, 13-16, 19-22, 24-27, 29-32, 34-36, 40, and 42, characterized in that, The input data is represented as multiple sets of matrices. Each set of matrices includes multiple matrices. Each set of matrices reflects the position and intensity information of the amplicon on the image acquired in one round of sequencing. Each matrix contains multiple rows and columns. Each matrix reflects the information of the amplicon on all or a part of a fluorescence channel image in one round of sequencing. Each element of the matrix corresponds to one amplicon. The row and column positions of the elements of the matrix reflect the position information of the corresponding amplicon on the corresponding image. The values ​​of the elements of the matrix reflect the intensity information of the corresponding amplicon presented on the fluorescence channel image in that round of sequencing.

44. The method according to claim 41, characterized in that, The input data is represented as multiple sets of matrices. Each set of matrices includes multiple matrices. Each set of matrices reflects the position and intensity information of the amplicon on the image acquired in one round of sequencing. Each matrix contains multiple rows and columns. Each matrix reflects the information of the amplicon on all or a part of a fluorescence channel image in one round of sequencing. Each element of the matrix corresponds to one amplicon. The row and column positions of the elements of the matrix reflect the position information of the corresponding amplicon on the corresponding image. The values ​​of the elements of the matrix reflect the intensity information of the corresponding amplicon presented on the fluorescence channel image in that round of sequencing.

45. The method according to claim 43, characterized in that, A matrix reflects information about the amplicon on a block of a fluorescence channel image in a round of sequencing, and the input data also includes a matrix reflecting the location information of each block in the image from which it comes.

46. ​​The method according to claim 44, characterized in that, A matrix reflects information about the amplicon on a block of a fluorescence channel image in a round of sequencing, and the input data also includes a matrix reflecting the location information of each block in the image from which it comes.

47. The method according to claim 3, characterized in that, The semantic segmentation model is selected from at least one of U-Net, DeepLab, SegFormer, PSPNet, FCN, GCN and their respective variants.

48. The method according to claim 47, characterized in that, The semantic segmentation model is selected from one of the following (i)-(iv): (i) is U-Net or SegFormer; (ii) For U-Net or SegFormer with added attention mechanisms; (iii) For U-Net with CBAM or SE modules added to the encoding section; (iv) U-Net incorporates a transformer-based multi-head attention mechanism in the decoding and skip connection parts.

49. The method according to claim 47, characterized in that, The semantic segmentation model includes: An encoder containing multiple downsampling layers, each downsampling layer containing at least one pooling operation and at least one convolution operation, and at least some downsampling layers have added CBAM or SE modules; A decoder comprising multiple upsampling layers, wherein the number of upsampling layers is the same as the number of downsampling layers, and each upsampling layer includes at least one deconvolution operation or bilinear interpolation operation; and A skip connection that connects the downsampling layer and the corresponding upsampling layer.

50. The method according to claim 49, characterized in that, The encoder further includes an input convolutional layer, the input data of which is processed by the input convolutional layer to produce an alternative representation as the input of the encoder's first downsampling layer, the input convolutional layer containing multiple convolution operations; and / or, the decoder further includes an output convolutional layer, the output convolutional layer processing the alternative representation produced by the decoder's last upsampling layer to produce the output.

51. The method according to claim 49 or 50, characterized in that, The neural network-based base recognizer or the segmentation model also has one or more of the following technical features (a)-(e): (a) The encoder comprises four downsampling layers; (b) Each of the downsampling layers includes a CBAM module or an SE module and a Down structure connected to the CBAM module or SE module, the Down structure including one pooling operation and two convolution operations; (c) The CBAM module includes a connected channel attention module and a spatial attention module. The channel attention module includes a first pooling unit, an MLP unit and a first output layer. The spatial attention module includes a second pooling unit, a convolution unit and a second output layer. In the channel attention module, the input is processed by the global max pooling and global average pooling of the first pooling unit to generate two channel descriptors. The two channel descriptors are processed by the shared MLP in the MLP unit to obtain channel weights. Then, the first output layer outputs the result of weighting the input using the channel weights. In the spatial attention module, the output from the first output layer is processed by the global max pooling and global average pooling of the second pooling unit to generate two spatial descriptors. The two spatial descriptors are processed by the convolutional layer of the convolutional unit to obtain spatial weights. Then, the output of the second output layer is output using the spatial weights to weight the output of the first output layer. (d) The SE module includes a channel attention module, which includes a third pooling unit, a fully connected unit, and a third output layer. An input channel is generated as a corresponding channel descriptor after global average pooling by the third pooling unit. The channel descriptors are processed by two fully connected layers and an activation function of the fully connected unit to obtain channel weights. Then, the third output layer outputs the result of weighting the input using the channel weights. (e) Each of the upsampling layers includes a bilinear interpolation or a deconvolution operation, and a concatenation operation that concatenates the output of the previous layer with the result of the corresponding downsampling layer passed to the upsampling layer via a skip connection in the channel dimension.

52. The method according to claim 49 or 50, characterized in that, Training the neural network-based base recognizer or the semantic segmentation model based on training data includes calculating the gradient of the loss function using backpropagation to guide stochastic gradient descent to update model parameters in order to minimize the loss function, wherein... The loss function is a hybrid loss function that includes at least one of cross-entropy loss, Dice loss, mean squared error, and one or more custom loss functions. The custom loss function includes mismatch loss, which reflects the proportion or content of bases that are inconsistent between the predicted result and the actual result.

53. The method according to claim 51, characterized in that, Training the neural network-based base recognizer or the semantic segmentation model based on training data includes calculating the gradient of the loss function using backpropagation to guide stochastic gradient descent to update model parameters in order to minimize the loss function, wherein... The loss function is a hybrid loss function that includes at least one of cross-entropy loss, Dice loss, mean squared error, and one or more custom loss functions. The custom loss function includes mismatch loss, which reflects the proportion or content of bases that are inconsistent between the predicted result and the actual result.

54. The method according to claim 52, characterized in that, The hybrid loss function includes cross-entropy loss, dice loss, and mismatch loss, and the weight ratio of the three is 1:3:4, 1:4:4, 1:5:4, 1:6:4, 1:7:4, or 1:7:

5.

55. The method according to claim 52, characterized in that, The neural network-based base recognizer or the semantic segmentation model is trained using supervised or semi-supervised learning methods, with the bases on the reference sequence corresponding to the base recognition result of the amplicon used as the real results for this training.

56. A base recognition system, characterized in that, include: A processing unit is configured to process input data from multiple amplicons using a neural network-based base recognizer and generate an alternative representation of the input data. The input data includes position and intensity information of each amplicon on images acquired during sequencing rounds n-m1 to n+m2. The sequencing employs surface imaging-based multi-fluorescence channel sequencing technology. The amplicon is a nucleic acid molecule to be tested containing multiple identical polynucleotide molecules, and the amplicon is attached to the surface. The input data is presented in a non-image format, and the file size of the input data is smaller than the file size of its corresponding image. n is a natural number greater than or equal to 1 and greater than m1, where m1 and m2 are each independently an integer greater than or equal to 0. An output unit is configured to process the substitution representation through an output layer to produce an output, wherein the output identifies the probabilities that the bases introduced into each of the amplicons in the nth round of sequencing are A, C, T, and G; and The base type determination unit determines the base type introduced into each of the amplicon in the nth round of sequencing based on the output.

57. A computer program product, characterized in that, The method of any one of claims 1-55 is performed when the computer executes some or all of the instructions.

58. A computing device, characterized in that, include: Memory is used to store data, including computer-executable programs; as well as One or more processors are configured to execute the computer-executable program to implement the method of any one of claims 1-55.

59. The computing device according to claim 58, characterized in that, The processor includes a GPU.

60. A computer-readable storage medium for storing a program executable by a computer, wherein executing the program includes performing the method of any one of claims 1-55.

Citation Information

Patent Citations

  • Methods, apparatus, and computer products for constructing sequencing templates based on images

    CN112289381B

  • Methods and systems for identifying bases in nucleic acids

    CN113012757B

  • Method for determining read quality and sequencing method

    CN117912550A

  • Method for determining mass fraction of read segment, sequencing method and device

    CN117976042A

  • Methods and systems for analyzing image data

    EP3077943B1