System and method for automatically generating base recognition model
Through a system that automatically generates a base recognition model, the model development process is simplified by using automated machine learning processes, solving the problems of intensive and low efficiency in the existing technology, and achieving efficient and accurate base recognition model development.
Patent Information
- Application Number
- CN202510244861.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, a lot of manual operations are required in the development of machine learning models in the field of sequencing or sequencing, resulting in low efficiency, high error rate and difficult to manage, especially with high requirements for model configuration and optimization in different scenarios.
A system is provided that automatically generates a base recognition model, the system includes an input layer, a machine learning layer, and an output layer. The machine learning layer processes the sequencing data input by the user through the training data generation module, generates multiple pairs of associated features and labels, and the training module then trains the specified model based on these data, including multiple iteration training and performance evaluation to obtain the base recognition model.
The automated machine learning process is realized, the modeling process is simplified, and the dependence on professionals is reduced, so that non-machine learning experts can also develop base recognition models on their own, improving efficiency and accuracy.
Smart Images

Figure CN120220825A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing based on artificial intelligence. Specifically, it relates to a system and method for automatically generating a base recognition model, and more specifically, to a system and method for generating a base recognition model based on an automatic training tool. Background Art
[0002] The subject matter discussed in this section should not be considered prior art merely because it is mentioned in this section. Similarly, technical problems mentioned in this section or associated with the subject matter provided as background art should not be considered to have been previously recognized in the prior art. The subject matter in this section only represents different methods, which may themselves also correspond to specific embodiments of the technical solutions included in the claims.
[0003] In related technologies, sequencing generally includes detecting relevant signals to identify the types of nucleotides or bases linked to one or more nucleotide positions in a nucleic acid molecule, that is, determining at least a partial nucleotide sequence of the nucleic acid molecule through base calling. Changes in the signals and / or signal intensities corresponding to specific positions of the nucleic acid molecule to be tested can indicate the base types at those positions on the nucleic acid molecule. For example, different nucleotides can be labeled with different fluorescent molecules or optically distinguishable luminescent markers. In one round of sequencing, multiple nucleotides are allowed to contact the nucleic acid molecule to be tested, and then the luminescent signals are detected to distinguish the types of nucleotides or bases linked to the template in this round of sequencing. Typical examples include the sequencing platforms of Illumina or MGI. It is also possible to separate the optical signal detection and nucleotide incorporation or ligation into independent steps to achieve one round of sequencing, such as the sequencing platform of Element Bioscience; there is also, for example, the Ion Torrent sequencing platform of ThermoFisher that does not rely on fluorescence signal or optical signal detection, but uses an electrochemical sensor to detect the pH change caused by the release of H+ ions when nucleotides are incorporated into the nucleic acid molecule to be tested to identify the types of incorporated nucleotides or bases; in addition, there are also single-molecule sequencing platforms based on various nanopores such as PacBio and Oxford Nanopore. Compared with the sequencing rounds mentioned above, the sequencing of such platforms is continuous. For example, the fluorescence signals emitted during the process of observing the incorporation of modified nucleotides with fluorescently labeled phosphate groups into a single-molecule template (the nucleic acid molecule to be tested) using a zero-mode waveguide (ZMW) nanopore are used to identify the types of incorporated bases in real time, or the base sequence of the nucleic acid molecule is directly determined by physically detecting the electrical signals generated when different bases pass through the nanopore without involving biochemical enzyme-catalyzed processes such as polymerization reactions. Currently, each sequencing platform (also known as a sequencing system or sequencer) sold by various companies on the market includes base calling software adapted to this platform, model, or detection principle to determine and distinguish the base types in the sequence of the nucleic acid molecule to be tested based on the detection and analysis of relevant signals.
[0004] In related technologies, machine learning includes various algorithms, such as traditional machine learning algorithms like linear regression, decision trees, support vector machines, and deep learning of neural networks. Applying machine learning to the field of sequencing or gene detection, for example, using machine learning to develop base recognition algorithms, base recognition result classification scores, or quality score algorithms, etc., there have been some relevant disclosures. For example, Wei-Chun Kao, et al.: "naiveBayesCall: An Efficient Model-Based Base-Calling Algorithm for High-Throughput Sequencing", 2010-04-25, RESEARCH IN COMPUTATIONAL MOLECULAR BIOLOGY, SPRINGER BERLIN HEIDELBERG, BERLIN, HEIDELBERG, pages 233–247, XP019141683, ISBN: 978-3-642-12682-6; CN118942549A, CN112789680, CN118429967A, CN118429965A, CN118116469A, CN117976042A, and US10068053B, etc., are all incorporated herein by reference in their entireties.
[0005] Common machine learning models for application directions in the field of sequencing or sequencing-related fields, including the model training, development, or usage processes involved in the above-listed disclosed solutions, often involve a relatively large number of manual operations or interventions. For example, in the early stage, it is often necessary to manually collect and preprocess data, select models, adjust hyperparameters, train models, and evaluate models. Manual intervention inevitably leads to the need to configure and optimize machine learning models for different real-world scenarios, resulting in relatively intensive manual operations. And relatively intensive manual operations generally tend to be prone to errors, low efficiency, or difficult management. Moreover, this process usually also requires the workers or operators to have rich relevant professional knowledge and experience, such as having algorithm development experience or a professional background, especially when the machine learning model is relatively complex and involves configuring and optimizing different types of algorithms. Summary of the Invention
[0006] The embodiments of the present application aim to at least partly solve at least one of the above technical problems or at least provide a practical commercial option.
[0007] Embodiments of the present application provide a system for automatically generating a base recognition model. The system includes: an input layer for receiving user input from a user, the user input including sequencing data and optionally one or more tasks, wherein the sequencing data is obtained by detecting nucleic acid molecules attached to multiple positions on a surface through surface imaging sequencing technology; a machine learning layer including a training data generation module and a training module. The training data generation module is configured to process the user input from the input layer to generate training data, the training data including multiple pairs of associated features and labels, the features describing the intensities corresponding to nucleic acid molecules at one or more positions, and the labels including one selected from bases A, T, C, and G. The training module is configured to train a specified model based on at least a portion of the training data, including performing multiple iterative trainings on the specified model using a training set and evaluating the performance of the specified model after each iterative training using a validation set, so as to obtain a base recognition model. The training set and the validation set are each independently a non-overlapping portion of the training data; and a result output layer for outputting the base recognition model. In related embodiments of the present application, the system for automatically generating base recognition is also referred to as an automatic training tool.
[0008] Embodiments of the present application also provide a method for automatically generating a base recognition model or a method for developing a base recognition model. The method includes accessing user input using the system for automatically generating a base recognition model in any of the embodiments, and running the system using at least one hardware processor to present the best base recognition model to the user. The so-called user input includes sequencing data and optionally one or more tasks, and the sequencing data is obtained by detecting nucleic acid molecules attached to multiple positions on a surface through surface imaging sequencing technology.
[0009] Embodiments of the present application also provide a base recognition model generated by the system for automatically generating a base recognition model in any of the embodiments.
[0010] Embodiments of the present application also provide a computer-readable storage medium storing multiple instructions for controlling a processor to execute the method for developing a base recognition model in any of the embodiments.
[0011] Embodiments of the present application also provide a system for automatically generating a base recognition model. The system includes a memory storing a program and one or more processors coupled to the memory, and the processors run the program to implement the method for developing a base recognition model in any of the embodiments.
[0012] Embodiments of the present application also provide a sequencing system or a sequencing platform, which includes the system for automatically generating a base recognition model in any of the above embodiments or examples.
[0013] The system, method, or related system for automatically generating a base recognition model according to any of the above embodiments realizes an automated machine learning process, simplifies the modeling process, and reduces the dependence on professionals, such as algorithm developers, enabling non-machine learning experts, such as sequencing platform users without relevant algorithm development backgrounds or experiences, to develop base recognition models by themselves.
[0014] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and / or additional aspects and advantages of the embodiments of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0016] Figure 1 is a schematic structural diagram of the system for automatically generating a base recognition model according to an embodiment of the present application;
[0017] Figure 2 is a schematic diagram of sequencing data in which the entire image is converted into a set of blocks and the set of blocks is written as a set of matrices as input according to an embodiment of the present application;
[0018] Figure 3 is a schematic structural diagram of the system for automatically generating a base recognition model according to an embodiment of the present application;
[0019] Figure 4 is a schematic structural diagram of the system for automatically generating a base recognition model according to an embodiment of the present application;
[0020] Figure 5 is a schematic structural diagram of the system for automatically generating a base recognition model according to an embodiment of the present application;
[0021] Figure 6 is a schematic structural diagram of the system for automatically generating a base recognition model according to an embodiment of the present application;
[0022] Figure 7 is a schematic structural diagram of the system for automatically generating a base recognition model according to an embodiment of the present application;
[0023] Figure 8 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0025] In this text, unless otherwise specified, the singular forms "a", "an", etc. used in a limiting manner include their plural referents (one or more than one). "A group" or "a plurality" means two or more than two.
[0026] In this text, "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity or order of the indicated technical features. In the description of this application, "a plurality" means two or more than two, unless otherwise specifically stated.
[0027] Unless otherwise specified, the terms "connected" and "joined" in this text should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection, or a connection that allows mutual communication; it can be a direct connection or an indirect connection through an intermediate medium, and can be the communication inside two components or the interaction relationship between two components; it can be a connection through physical adsorption or other acting forces, or a chemical connection such as a chemical bond, such as incorporation through a polymerization reaction. Those skilled in the art can understand the specific meaning of this term in the corresponding examples according to the specific embodiments described, including the context and common knowledge.
[0028] In this text, "sequencing" refers to nucleic acid sequence determination, the same as "nucleic acid sequencing" or "gene sequencing", which means determining the base sequence of the primary structure of a nucleic acid molecule, and can be achieved by methods such as sequencing by synthesis (SBS), sequencing by ligation (SBL), or sequencing by hybridization (SBH). Unless otherwise stated, the so-called sequencing by synthesis in this text, in addition to including the commonly understood SBS that uses polymerase to catalyze nucleotide incorporation into the nucleic acid molecule to be tested (polymerization reaction) and detects the corresponding reaction signals to identify the types of incorporated nucleotides (the typical ILLUMINA / Solexa technology), also includes sequencing similar to SBS that uses polymerase or non-polymerase to controllably introduce or ligate nucleotides to the nucleic acid molecule to be tested, directly or indirectly, simultaneously or successively detecting the corresponding signals to determine the types of one or more ligated nucleotides, such as SBL, SBH, SBB (sequencing by binding), or SBE (sequencing by expansion) that achieve sequencing based on surface fluorescence imaging detection.
[0029] In the process of sequencing based on solid-phase surface signal detection, sequencing can also be said to measure the intensity values corresponding to the positions of one or more nucleic acid molecules. The intensity values here are sometimes simply referred to as intensities, and can be any signal such as the intensity of electrical or electromagnetic radiation such as visible light. A base can have one intensity value, or multiple intensity values, or multiple bases can have one or fewer intensity values than the number of bases. Additionally, the intensity values can be for specific positions on the solid-phase surface or for multiple positions of the nucleic acid molecule. The intensity values can be converted or mapped to pre-determined numerical values (such as integers in a binary or decimal number system) or continuous numerical values or a certain range.
[0030] Sequencing can be carried out through a sequencing platform. According to the embodiments of the present application, the selectable sequencing platforms include but are not limited to the Hiseq, Miseq, Nextseq, and Novaseq series sequencing platforms of Illumina, the Ion Torrent platform of Thermo Fisher / Life Technologies, the BGISEQ and MGISEQ / DNBSEQ platforms of BGI Genomics, the AVITI of Element Bioscience, and single-molecule sequencing platforms; the sequencing method can be single-end sequencing, paired-end sequencing, or the sequencing methods supported by the selected automated sequencing platform, etc.
[0031] The "sequencing process" or "sequencing run" in the embodiments of the present application refers to measuring the intensity values corresponding to the positions of one or more nucleic acids in batches. For example, in a scenario where sequencing involves multiple cycles (also called rounds or sequencing rounds) of imaging nucleic acid molecules on a substrate such as the surface of a solid-phase substrate after a specified biochemical reaction, a series of intensity values are obtained during the same sequencing run. Generally, the nucleic acid molecules in the first sequencing run do not participate in the second sequencing run independent of the first sequencing run, such as not being included in an image collected at a certain time point (a certain round) of the second sequencing run.
[0032] In some examples, sequencing by synthesis is used for multiple rounds of sequencing to obtain a sequencing sequence or read. For example, the nucleic acid molecule to be tested is contacted with a polymerase and modified nucleotides and placed under conditions suitable for polymerization reaction, and the modified nucleotides are controllably incorporated into the nucleic acid molecule to be tested or single-base extension is controllably achieved, and the corresponding reaction signal is detected. Based on this signal, the type of nucleotide incorporated into the nucleic acid molecule to be tested in this reaction is determined. In this way, multiple controllable single-base extensions and corresponding signal detections are carried out, so as to detect the types of nucleotides or bases incorporated into the nucleic acid molecule to be tested in multiple or multiple rounds of reactions according to the reaction signal information, in order to sequence a part of the nucleic acid molecule to be tested.
[0033] The so-called nucleic acid molecule to be tested, also known as nucleic acid template or template, can be an unamplified single molecule, or a molecular cluster or mass or long chain containing multiple identical polynucleotide molecules after amplification, such as the fluorescence clusters (clusters) or DNA nanoballs (DNBs) formed by bridge amplification or rolling circle amplification adopted by current commercially available mainstream sequencing platforms. The nucleic acid molecule to be tested can be presented in the form of single-stranded, double-stranded, and / or triple-stranded or more-stranded complexes hybridized with probes or primers.
[0034] The corresponding reaction signals can be, for example, fluorescence signals, and can be presented as image data (such as color images or grayscale images) formed by collecting these fluorescence signals. Thus, these image data are processed and analyzed to detect the nucleotides incorporated into the nucleic acid molecule to be tested in each reaction or each round of reaction, so as to determine a part of the base sequence of the nucleic acid molecule to be tested.
[0035] Specifically, in some examples, sequencing is achieved based on surface fluorescence imaging detection. The nucleic acid molecule to be tested is connected to a solid surface. For example, nucleotides can be modified to carry or be capable of binding fluorescence labels, and carry excisable inhibitory groups that can prevent other nucleotides from polymerizing and connecting to the next position of the nucleic acid molecule to be tested (such modified nucleotides are also called reversible terminators). And after each polymerization reaction or single-base extension reaction is completed, it is excited to make the fluorescence label emit light, and these emission signals are collected to obtain an image of the nucleic acid molecule to be tested where single-base extension reaction occurs at a specified surface position; then, the inhibitory group and fluorescence label are removed, etc. to perform the next or the next round of polymerization reaction and signal collection (taking pictures). Repeating such polymerization reaction - taking pictures - excision multiple times or multiple rounds to obtain image information related to the nucleotides connected to the nucleic acid molecule to be tested in each single-base extension reaction.
[0036] It can be understood that when the nucleic acid molecule to be tested at a specified surface position undergoes a polymerization reaction and emits fluorescence, it generally appears as bright spots or bright patches (spots) with a signal intensity higher than the background at the corresponding position in the image collected in this round of reaction. Therefore, according to the information corresponding to specific chemical characteristics (the nucleic acid molecule to be tested undergoing polymerization reaction) contained in these image sets, such as intensity and / or morphology, etc., it can be judged whether there is a nucleic acid molecule to be tested at the specified position and whether the nucleic acid molecule to be tested undergoes a polymerization reaction. Further, in combination with the preset corresponding relationship between fluorescence emission signals and nucleotide types, the type of nucleotide connected to the nucleic acid molecule to be tested in the biochemical reaction can be detected, and thus at least a part of the sequence of the nucleic acid molecule to be tested can be determined to obtain the so-called read segment.
[0037] Determining a set of positions corresponding to chemical features (nucleic acid molecules to be detected) on a surface based on an image, that is, obtaining a template, can be determined by identifying features corresponding to chemical features on the detection image such as bright spots (positions), or can also be determined by identifying other features on the image that have a specific association relationship with the spatial position relationship of the chemical features. For example, for a surface of a regular array containing a labeling region (also often referred to as a tracer region), generally the spatial position relationship between the labeling region and the reaction region (the region where the nucleic acid molecules to be detected are located) is preset and known, and the signal from the labeling region and the signal from the reaction region are clearly distinguishable or the signal from the labeling region presents as a detectable feature on the image. By identifying the position of the labeling region on the image, the positions of each amplicon (nucleic acid molecule to be detected) in the reaction region can be determined.
[0038] It should be noted that the nucleotides referred to in this article include ribonucleic acid or deoxyribonucleic acid, including natural nucleotides or their derivatives or modified forms (also referred to as modified nucleotides or modified nucleotides, etc.). In this article, sometimes the nucleotide is also referred to by the base it contains, and those skilled in the art can clearly understand according to conventional knowledge and / or context. In addition, the nucleic acid sequences, nucleotide sequences, oligonucleotide sequences or base sequences referred to in this article can be used interchangeably unless otherwise specified, and sometimes are directly abbreviated as sequences. In addition, sometimes one or more nucleotides or bases are also referred to as sequences, and those skilled in the art can clearly understand according to conventional knowledge and / or context.
[0039] In some embodiments, a round of sequencing can include a single base extension reaction (one repeat). For example, four different nucleotides (dATP, dTTP, dGTP, and dCTP) can be placed in the same polymerization reaction system with multiple nucleic acid molecules to be detected, so that each nucleotide can be excited to emit a signal distinguishable from other types of nucleotides, thereby determining the type of nucleotide incorporated or introduced at a position of any nucleic acid molecule to be detected based on the information obtained from one repeat reaction. For example, the four nucleotides are respectively labeled with fluorescent labels of four different emission bands for four-color or four-channel single-molecule sequencing or high-throughput sequencing; or for another example, the four nucleotides are respectively labeled with fluorescent labels of three different emission bands and unlabeled (cold nucleotides) for two-color or three-color high-throughput sequencing. One round of sequencing includes one repeat, and one round of sequencing can detect the type of base at a position on any nucleic acid template.
[0040] The so-called amplicon is the nucleic acid molecule to be tested after amplification. An amplicon is a cluster or strand or mass containing multiple identical polynucleotide sequences; the so-called fluorescent cluster, mass or sphere is an amplicon that can be excited to emit light such as fluorescence after a specified biochemical reaction. The amplicon can emit light and be imaged and detected during the sequencing process. By processing the signals collected by imaging and / or the corresponding image information, the bases incorporated or linked to the nucleic acid molecule to be tested after the specified biochemical reaction can be detected, thereby determining at least a part of the sequence of the nucleic acid molecule to be tested.
[0041] In the embodiments of the present application, unless otherwise specified, the so-called amplicon, fluorescent cluster or mass or sphere, nucleic acid molecule cluster, clone mass, nucleic acid molecule to be tested, template, nucleic acid template, template spot, bright spot corresponding to chemical features or fluorescent clusters on the surface, etc. can generally be used interchangeably. The so-called signal intensity, intensity, brightness, fluorescence brightness, etc. can generally also be used interchangeably. The so-called surface, chip, array, etc. can generally also be used interchangeably in the embodiments of the present application.
[0042] The machine learning model referred to in the embodiments of the present application, also simply referred to as the model, refers to a determination technique for predicting the type of output base or base combination based on known results (training data); in some embodiments, it also includes a technique for predicting the quality score of the type of output base or base combination. The so-called quality score characterizes the accuracy or error rate or reliability of the predicted base or base combination result. The known results can be sequences or bases considered or assumed to be accurate, that is, it is assumed that the sequence or base is correct. In some embodiments, the sequence or base assumed to be accurate is the result expected to be predicted by the model, and sometimes it is also called the corresponding ground truth base. The model can be supervised learning using the training data.
[0043] In the embodiments of the present application, the concepts of labels or target values involved in machine learning are usually synonymous and refer to the results that the model needs to predict. The bases or base sequences in the labels or target values in machine learning are assumed to be accurate sequences (bases or base combinations), which are recognized as accurate sequences and are sometimes also called ground truth bases. The measured or generated training data, including the labels therein, may be inaccurate, but it is assumed to be accurate during model training. The assumed sequences can be determined in various ways. As described in the embodiments of the present application, the reference sequence, for example, the bases or base combinations corresponding to the positions where the sequencing results are aligned to the reference sequence, is used as the assumed sequence. The sequences determined or verified by detection methods or means generally considered to have higher accuracy, such as Sanger sequencing and / or DNA synthesis, are used as the assumed sequence. The consistent results determined by multiple detection methods or means are used as the assumed sequence. Or the sequences or base combinations with a high probability of being accurate, such as bases with a quality score reaching or exceeding Q40, or sequences or base combinations with Q40 reaching or exceeding 80 or 85 or 90, are used as the assumed sequence, etc. Here, Q40 represents an error rate or error probability of one in ten thousand. If it is not the desired goal to train the model to output highly accurate prediction results, the measured results with low accuracy can also be used as the assumed sequence and as the training data. However, it can be understood that the accuracy of the prediction results of the model learned accordingly is generally not high either.
[0044] The base recognition referred to in the embodiments of the present application is sometimes also called base class recognition, base detection, base class detection, base reading, base determination, or base interpretation, etc., and refers to the determination of bases or base combinations at specified or unspecified positions in nucleic acid molecules. The base determination can be carried out independently or as part of a specified base, such as A or T, or a specified base combination, such as ACT or GTCA, etc. The base determination can be for the same genomic position, for example, the scores or probabilities of detecting multiple bases at this position are close to each other, or for different positions. The scores output from the machine learning model can be used to determine the base types at specified or unspecified positions. For example, scores can be provided or assigned to each base or base combination, and the determination of bases based on the scores can be part of the model prediction. In addition, some models can provide scores, and these scores can be used in subsequent processes. The scores or values can reflect probabilities or possibilities. The probability scores of each base or base combination often sum up to a fixed value, such as 1. The scores reflecting possibilities generally do not need to sum up to a fixed value, but of course, they can also sum up to a fixed value. For example, the possibility scores of each are limited to between 0 and 1, so that the possibility scores can also sum up to 1. In the embodiments of the present application, the values or scores that provide or assign the probabilities or possibilities of each base or base combination are sometimes collectively referred to as base recognition probabilities.
[0045] In addition, in some embodiments of the present application, during the process of the quality score algorithm or quality score model or related model learning further involved, the accuracy or reliability of the base determination result is scored based on the quality score model generated by training, so as to quantitatively evaluate or appraise the accuracy or credibility of such base determination results, which can be used to compare the sequencing quality of different sequencing methods or sequencing platforms, and to determine whether the accuracy of the bases or sequences identified in one or more rounds of sequencing or one or more sequencing runs meets the expectation and whether to continue the current sequencing run, etc.
[0046] Please refer Figure 1 , some embodiments provide a system 100 for automatically generating a base recognition model. The system 100 includes: an input layer 120 for receiving user input from a user, including sequencing data and optionally one or more tasks, wherein the sequencing data is obtained by detecting nucleic acid molecules connected to multiple positions on a surface through surface imaging sequencing technology; a machine learning layer 140 including a training data generation module 142 and a training module 144. The training data generation module 142 is used to process the user input from the input layer 120 to generate training data, which includes multiple pairs of associated features and labels. The so-called features describe the intensity corresponding to nucleic acid molecules at one or more positions, and the so-called labels include one selected from the bases A, T, C, and G. The training module 144 is used to train a specified model based on at least a part of the training data, including performing multiple iterative trainings on the specified model using a training set and evaluating the performance of the specified model after each iterative training using a validation set, so as to obtain a better base recognition model. The so-called training set and validation set are each independently a non-overlapping part of the training data; and an output layer 160 for outputting the better base recognition model. In related embodiments, the system 100 for automatically generating base recognition is also referred to as an automatic training tool.
[0047] The sequencing data included in the user input received by the input layer 120 can be data obtained from one or more nucleic acid samples through one or more library constructions and one or more sequencing runs on one or more sequencing platforms. For example, the electricity or electromagnetic radiation (such as visible light) detected during the sequencing process, which can be in the form of an image or image data or other forms or formats. The sequencing data can be obtained by self-sequencing, or provided by others or publicly available, or a mixture or combination of sequencing data from various sources. In some embodiments, the sequencing data provided by the user input can be in the form of an access link or path to reduce storage requirements and facilitate the system 100 to receive and connect to read the data. The model output by the output layer 160 can also be in the form of an access path or link, for example, stored or saved under a specified path to facilitate calling and running.
[0048] In some embodiments, the sequencing data is derived from a sequencing platform that implements sequencing based on surface imaging detection. The sequencing data includes images of nucleic acid molecules at multiple positions on one or more surfaces collected from one or more rounds of sequencing detection in one or more sequencing runs. The so-called images or sequencing images include images directly collected during the sequencing process, images collected after preprocessing of the sequencing, or images generated by reconstructing or reorganizing the images collected based on the sequencing. In some embodiments of the present application, these images are sometimes also referred to as original images. Thus, the subsequent training data generation module 142 processes such input image data, for example, extracts the positions and intensities of individual nucleic acid molecules in the images as feature values.
[0049] In some embodiments, the sequencing data is a part of the data determined by extracting from the sequencing images. The so-called sequencing data includes the positions and intensity values of multiple nucleic acid molecules in one or more surface regions where they are located. The so-called sequencing images are images of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of sequencing detection in one or more sequencing runs. In certain examples, the sequencing data is the information obtained after processing the images, including information representing or reflecting the target signal, that is, the nucleic acid molecules or amplicons to be detected that have undergone a specified biochemical reaction, such as the position information and intensity values of the amplicons. The size of the input sequencing data is smaller than the size of the corresponding sequencing image. Here, the so-called size is also referred to as the file size or dimension or data volume, and such sizes reflect the storage space required to store the data or the storage space occupied. Thus, compared with the case where the input data is the sequencing image, such input can be directly used as feature values or directly as part of the training data, making the model generated using this system 100 have significant advantages in terms of computing power consumption and calculation speed.
[0050] Specifically, the input sequencing data contains the position information and intensity information of the amplicons in the image, and the size of the input data is smaller than the size of the corresponding image. Therefore, from a certain perspective, it can be said that the input data is compressed image information or data obtained by extracting or screening the image information, which can be represented in a simplified image form or a non-image form. Preferably, it is represented in a data form that is convenient for a computer or a processor to calculate and process. In a preferred example, it is represented in a non-image data form, such as written in the form of a matrix, a multi-dimensional matrix, an array, a multi-dimensional array, etc., only recording or presenting the relative position information and intensity information of the amplicons in the image. Thus, compared with the input original image, the file size of the input data can be greatly reduced, the requirements for computing power and / or storage can be lowered, and it is also beneficial for the computer to quickly perform various operations or processes on it to quickly obtain the prediction result.
[0051] In some specific embodiments, the size of the input sequencing data is less than or equal to half of the size of its corresponding image. In some examples, especially for regular surfaces, the file size of the input data is less than or equal to one-third of the file size of its corresponding image. In some tests, especially for high-density regular surfaces, the size of the input data is even less than or equal to one-fourth or one-fifth of the file size of its corresponding image. Thus, compared with directly inputting image data, the storage or occupied space is significantly reduced, and the running speed is significantly increased.
[0052] Specifically, in some examples, the input sequencing data is represented as multiple groups of matrices. One group of matrices includes multiple matrices. One group of matrices reflects the position information and intensity information of amplicons on the images collected in one round of sequencing. One matrix contains multiple rows and multiple columns. One matrix reflects the information of all or a region of an amplicon on the image of one fluorescence channel in one round of sequencing. One element of the matrix corresponds to one amplicon. The row and column positions where the element of the matrix is located reflect the position information of the corresponding amplicon on the corresponding image. The value of the element of the matrix reflects the intensity information of the corresponding amplicon presented on the image of one fluorescence channel in this round of sequencing. In some embodiments, this matrix is also referred to as a fluorescence brightness matrix. Thus, compared with inputting images, the information amount or size of the input sequencing data is greatly reduced, which is beneficial to reducing the requirements for storage or computing power, and is also beneficial to improving the running calculation speed. Moreover, such input sequencing data can be directly used as eigenvalues or part of eigenvalues, reducing subsequent operations.
[0053] In a more specific example, one piece of input sequencing data can correspond to multiple amplicons in a region or field of view (FOV). One piece of sequencing data, such as 1 or 1 group of "fluorescence brightness matrices", can correspond to the position and intensity information of the amplicons that have undergone a specified reaction in the image collected in one round of sequencing for one FOV. Thus, compared with inputting the original image, the size of the input data is significantly reduced, and the requirements for computing power and / or storage can be reduced.
[0054] Regarding the number of matrices included in one group and the dimensions of the matrices, it can be understood that one group of matrices corresponding to a group of images of one FOV for multiple fluorescence channels in one round of sequencing, such as four fluorescence brightness matrices of amplicons in four fluorescence channels of one round of sequencing for one FOV from four-color sequencing, such as three fluorescence brightness matrices of amplicons in three fluorescence channels of one round of sequencing for one FOV from three-color sequencing, such as two fluorescence brightness matrices of amplicons in two fluorescence channels of one round of sequencing for one FOV from two-color high-throughput sequencing, can all be referred to as a three-dimensional matrix. Combining with common knowledge, those skilled in the art can understand the specific meaning of the description involving matrices, the number of matrices, or dimensions according to the specific embodiments including the context.
[0055] Further, in some examples, the advantages of the input matrix compared to the original input image are not only in the differences in image and matrix sizes. For example, among the four images of a field of view from a regular surface collected in one round of four-color fluorescence imaging sequencing, assuming that the size of each image is 2000*2000 pixels, while the size of the template point matrix (fluorescence brightness matrix) reflecting the positions and intensity information of amplicons on the image, written in matrix form, is 800*800, then the size is reduced to 0.4^2 times the original. Moreover, the brightness information corresponding to the 800*800 template points input into this matrix can be the brightness information corresponding to sub-pixel level positioning positions, such as the brightness obtained through bilinear interpolation or other interpolation methods. Therefore, in terms of information volume, the information of this 800*800*4 template brightness matrix reflects the information of the four 2000*2000*4 matrices of images plus the 800*800 position matrix. It can be said that the information volume of the input matrix reflects the position information and intensity information of amplicons determined directly or indirectly based on image information, including the information of processed images. Thus, the input data not only contains target information, its size is significantly reduced, and it also presents in a matrix form that is easy for machine operation and calculation, further facilitating the rapid prediction to obtain accurate and reliable base recognition results. Moreover, such input sequencing data can be directly used as eigenvalues, or as part of features or training data to reduce subsequent operations.
[0056] In some other specific examples, the format of the sequencing data is a matrix. One matrix reflects the information of amplicons on a block of an image in one fluorescence channel of one round of sequencing. The input sequencing data also includes a matrix reflecting the position information of each block in the image it comes from. Thus, the matrix corresponding to one fluorescence channel image is converted into a group of relatively small matrices, or in other words, one fluorescence channel image is converted into a group of blocks, and the information of amplicons in the simplified blocks is extracted in matrix form. By processing these relatively small matrices in parallel, in the same hardware and software operation environment, the processing speed of the input data can be further improved, and the model can be trained more quickly.
[0057] Please refer to Figure 2 , in a specific example, an image is converted into a group of block matrices. Different-sized block matrices are all zero-padded to the same size as the input. Therefore, the entire image is converted into a total of 3*3, that is, 9 blocks; and the length and width of the blocks may be different. Therefore, after splitting into 9 blocks, the edges of all blocks with sizes less than b*y can be filled with 0 values to a size of b*y; finally, an additional position encoding layer is added according to the position numbers of the blocks in the whole image. For example, the block in the lower right corner is numbered 9, and a layer of Pos(P) position encoding is filled, and all the values are 9, to obtain the final input matrix.
[0058] In some embodiments, the sequencing data can reflect no less than 10 million sequence reads (also referred to as reads), and / or no less than 1 billion base pairs (bp) of base reads. In this way, enough training data can be generated so that it is possible to generate a model with better prediction effect based on the system 100.
[0059] In some embodiments, in the machine learning layer 140, the training data generation module 142 processes user input from the input layer 120 including sequencing data to generate training data. The generated training data includes multiple pairs of associated features and labels. The so-called features describe the intensity corresponding to nucleic acid molecules at one or more positions. Such features are, for example, the sequencing data in any of the above embodiments, such as a sequencing image, a processed sequencing image, a matrix corresponding to the specific position information of the image (amplicon position and intensity information), or a block matrix group corresponding to the specific position information of the image. The so-called intensity, sometimes also referred to as brightness, can be represented by the pixel value / gray value, sub-pixel value, or sub-pixel value at the position of the nucleic acid molecule on the image. For example, if the feature of a nucleic acid molecule at a certain position occupies one or more pixels or multiple sub-pixels on the image, the intensity value of any pixel or sub-pixel where it is located can be used to represent the intensity information of the nucleic acid molecule at that position on the image. It is also possible to determine the center or centroid of the feature and use the pixel value or sub-pixel value at its center or centroid to represent the intensity information of the nucleic acid molecule at that position on the image. The intensity information of nucleic acid molecules at one or more positions on the image can be the original signal intensity value on the image, or the processed signal intensity value, such as the intensity value after background removal and / or corrected by at least one of chromatic aberration, crosstalk, and phase, or the intensity value after normalization or standardization processing, etc. Crosstalk correction, phase correction (phasing or prephasing correction), etc. can be performed by methods disclosed in, for example, EP3077943B1, CN113012757B, etc. The entire content of the cited relevant literature is incorporated herein by reference. The so-called labels include one selected from the four bases A, T, C, and G.
[0060] In some examples, the training data includes multiple pairs of associated features and labels. The features can be, for example, the sequencing data or the processed sequencing data or a part thereof formed by any relevant embodiment or combination of relevant embodiments in the context. The base or base combination category corresponding to the sequencing data (the corresponding ground truth base) can be used as a label.
[0061] The embodiments of the present application do not limit the generation method or manner of training data. Specifically, through conventional base recognition methods or software or tools, such as the base recognition software usually carried or provided by various sequencing platforms / sequencers, the bases incorporated into the template nucleic acid in each round can be recognized to obtain the base recognition results of each round or determine at least a part of the sequence of the nucleic acid molecule to be detected. In related examples, the so-called ground truth bases or reference answers, etc., come from user input or are determined by processing the sequencing data and tasks input by the user. In some specific examples, the training data generated by the training data generation module 142 uses the bases on the reference sequence corresponding to the base recognition results of the amplicons as labels or target values. The so-called labels or target values are sometimes also referred to as true results or correct answers or true labels or ground truth bases in the embodiments of the present application.
[0062] In a certain embodiment, the user input further includes a reference sequence. The reference sequence is a sequence with a known sequence, which can be a reference genome, a part of the reference genome, or a known sequence formed by recombining and reorganizing one or more reference genomes. Please refer to Figure 3 , the training data generation module 142 further includes an alignment sub-module 1422. The alignment sub-module 1422 is used to align the sequence reads or base reads reflected by the sequencing data to the reference sequence to generate an alignment result; based on the alignment result, determine the labels associated with the features. For example, use the bases corresponding to the sequence or base reads of the sequencing data aligned to the reference sequence as the corresponding labels to generate training data.
[0063] In a specific example, the alignment sub-module 1422 is connected to a sequencing system or a sequencing platform and is connected to the base recognition software provided or carried by the sequencing system or the sequencing platform itself.
[0064] In addition, during the process of the training data generation module 142 constructing training data, sequences with poor sequencing quality or unreliable or relatively low-confidence alignment positions can be excluded to improve the reliability of the generated training data. Training the model based on training data with high reliability is conducive to obtaining a model with more accurate and reliable predictions. For example, in a certain example, for sequences that fail to be successfully aligned to the reference genome or whose alignment results are unreliable or have low confidence, these sequences or bases are assigned another value of "N" or marked as "N", and the bases or sequences marked as N are not used as training data or do not participate in the subsequent model training process, which can improve the reliability of the generated training data or the model trained based on such training data.
[0065] In some embodiments, the so-called tag further includes a combination selected from any two, three, or four of the bases A, T, C, and G. In one example, the so-called tag further includes a combination formed by base recognition in the current round and the previous round, and the so-called tag includes a combination formed by base recognition in the current round and the next round, such as AA, AT, AC, AG, TT, TA, TC, TG, CC, CA, etc.; in another example, the so-called tag further includes a combination formed by base recognition in the current round and the previous and next rounds, such as AAA, ATG, ATT, ATC, ACG, ACT, ACA, TAT, CGA, GTA, etc.; in yet another example, the tag further includes a combination formed by base recognition in the current round, the previous round, and the next two rounds, such as AATT, AATG, AATC, AAAT, CGTA, GTAC, TGAC, ACTG, etc. In a specific example, the sequence corresponding to the sequencing data or the combination of base reads is aligned to the corresponding base combination on the reference sequence as such combinatorial tags. Training a model based on the training data of this type of tag usually requires relatively high computing power and memory configuration.
[0066] The so-called specified model can be any machine learning model including deep learning neural networks. In some embodiments, the specified model trained in the training module 144 is selected from at least one of Naive Bayes, Logistic Regression, Support Vector Machine, tree models such as lightGBT, and neural networks. In some specific examples, the specified model includes a neural network, such as a model capable of semantic segmentation, such as U-net or its variants, etc., and is a single model or a cascaded model, etc. In a specific example, the training of the specified model is implemented using the solution disclosed in CN119479831A, which is hereby incorporated herein by reference in its entirety.
[0067] Please refer to Figure 4 In some embodiments, the machine learning layer 140 further includes an evaluation module 146. The evaluation module 146 is connected to the training data generation module 142, the training module 144, and the output layer 160, and is used to evaluate the performance of the base recognition module using a test set before the output layer 160 outputs the base recognition model, so as to determine that the base recognition model meets the preset requirements and can be output. Among them, the test set is a part of the training data that does not overlap with the training set and the validation set. The so-called meeting the preset requirements means that at least one of the accuracy, precision, recall rate, and speed of the classification prediction result of the base recognition model for the test set meets the preset standard or can achieve the task input by the user.
[0068] Please refer to Figure 5, in some embodiments, the training data generation module 142 is the first training data generation module 142, the training data is the first training data, the training module 144 is the first training module 144, the designated model is the first designated model, the so-called feature is the first feature, the so-called label is the first label, and the machine learning layer 140 further includes a second training data generation module 152 and a second training module 154. Among them, the second training data generation module 152 is connected to the first training data generation module 142 and the first training module 144, and is used to process at least a part of the first prediction results generated by the base recognition model for the first training data, so as to generate second training data. The first prediction results include the probability and base category that the data is predicted to be recognized as the base category in the first label. The second training data includes multiple pairs of associated second features and second labels. The second features describe the intensity corresponding to the nucleic acid molecule at one or more positions and / or describe the probability of at least one base category (at least one base recognition probability). The second labels include two categories: whether the first prediction result is consistent or inconsistent with the corresponding ground truth base category. The second training module 154 is used to train the second designated model based on at least a part of the second training data, including performing multiple iterative trainings on the second designated model using the second training set and evaluating the performance of the second designated model after each iterative training using the second validation set, so as to obtain a quality score model. The quality score output by the so-called quality score model characterizes the accuracy, precision or error rate of the base recognition result predicted by the base recognition model from the first training module 144 or the output layer 160; the second training set and the second validation set are each independently a non-overlapping part of the second training data. In this way, the training data (second training data) for generating the quality score model is generated using the generated base recognition model and the prediction results of the model, and model learning is performed based on such training data to generate a quality score model.
[0069] In the related art, regarding the quality score or quality fraction (Q value) of the base recognition result / base classification detection, regardless of what method is used for calculating the Q value output by each sequencing platform of each company or what rules or custom rules are involved, the Q value can usually be regarded as another quantitative representation of the probability that the base detection result is incorrect or correct, and is an index characterizing the accuracy or reliability of the sequencing result. The higher the Q value, the higher the accuracy and reliability of the sequencing result. The Q value is usually calculated based on the Phred algorithm, and the formula can be written as Q = -10 * log 10(err), where err refers to the probability of base recognition error, that is, the error rate or error probability of the base detection result. For example, the common Q20, Q30, and Q40 respectively represent an error rate of one percent (i.e., a correct rate of 99%), an error rate of one-thousandth (i.e., a correct rate of 99.9%), and an error rate of one-ten-thousandth (a correct rate of 99.99%). The accuracy and reliability of the sequencing data can be intuitively understood through the Q value. In addition, in various practical applications of detecting by processing sequencing data or sequencing results, such as genetic variant detection like non-invasive prenatal screening, early screening of tumors by liquid biopsy, and pathogen detection, etc., it usually involves using the Q value to process sequencing data or setting quality standards to make the quality of the sequencing data relatively high or meet specific requirements to ensure the accuracy and reliability of the application detection results, etc.
[0070] In some examples, the features (first features) for training the base recognition model can be used as the second features. For example, the second features also describe the intensity corresponding to the nucleic acid molecules at one or more positions. Specifically, the second feature values can be determined with reference to the sequencing data or the associated (first) feature values in any of the above relevant embodiments. The second feature values can also be determined based on the prediction results of the base recognition model. For example, the second feature can be the probability of describing at least one base category (base recognition probability), such as the probability values of all base categories, the probability value of the base category with the highest probability, the probability values of the base categories with the highest and the second highest probabilities, etc. The second feature can also describe both the intensity corresponding to the nucleic acid molecules at one or more positions and the probability of at least one base category. The base categories referred to here are the same as the labels or the first labels in the above or relevant embodiments, and can be base categories formed by single bases such as A, T, C, G, or base categories formed by combinations of a single base with two or more bases such as AT, ACG, etc.
[0071] Specifically, in some examples, the second feature describes the probability of at least one base category in the first prediction result, that is, the second feature describes at least one base recognition probability. For example, the second feature can be the probability value of each base category including base combination categories recognized as the first label, or the probability value of the base category with the highest probability among them, the probability values of the base categories with the highest and the second highest probabilities among them, the probability values of the top three base categories sorted in descending order of numerical values among them, the probability values of the top four base categories sorted in descending order of numerical values among them, or any one or combination of the above or any one or combination of the processed or deformed probability value forms, etc. The probability or probability value referred to here and in any relevant embodiment can be an absolute value, a relative value, a ratio, a deformed numerical value after being normalized, etc., or a deformed numerical value enlarged or reduced based on a certain rule. In a specific example, the second feature describes the probabilities of multiple base categories in the first prediction result.
[0072] In some examples, the so-called corresponding ground truth base category is the base category determined to be correct. In a specific example, the corresponding ground truth base category comes from the sequences of human and non-human samples that have been detected and fully verified and characterized consistently by various sequencing instruments, sequencing chemistries, and sequencing protocols. Thus, it can be considered as the correct or true base category, or the accurate and reliable base category confirmed by the development of relevant detection methods and means up to now. In a certain example, the corresponding ground truth base category is the label or the first label (target value or the first target value) in the process of training to generate the base recognition model.
[0073] In some examples, the second specified model is selected from at least one of logistic regression, support vector machine, and tree model.
[0074] In some specific examples, training the second specified model in the second training module 154 further includes obtaining the classification probability of the base category of the first prediction result in the second training set; dividing the first prediction results into a groups based on these classification probabilities and / or base categories, and for each group, by comparing the base category in the first prediction result of this group with the corresponding ground truth base category, to determine the error rate of the base category of the first prediction result of this group, where a is a natural number and the value of a can achieve that the logarithmic difference of the error rates of the base categories of the first prediction results of adjacent groups is less than 2; and associating a predetermined quality score with the error rate of the base category. Thus, a corresponding relationship between each base recognition result / classification probability of the base category and the predetermined quality score is established, and this corresponding relationship is the so-called quality score model.
[0075] Specifically, the first prediction result can pass through the quality score model of any of the above related embodiments to obtain the so-called classification probability P value, denoted as P0, and P0 is a probability value with a numerical value between 0 and 1. Further, P0 can be divided into a segments, and the base categories in the first prediction result will also be correspondingly divided into these a segments or groups, where a is a natural number not less than 5. In some specific examples, P0 is divided into a segments in an evenly divided manner. For example, if P0 is evenly divided into 5 segments, then the ranges of these 5 segments of P0 are [0, 0.2), [0.2, 0.4), [0.4, 0.6), [0.6, 0.8), and [0.8, 1.0] in sequence.
[0076] Further, for each group, by comparing the base category in the first prediction result of each group with the corresponding ground truth base category, to determine the error rate or accuracy rate of the base category in the first prediction result of this group, including identifying or marking the bases that are consistent with the corresponding ground truth base category in this group as correct, and those that are inconsistent as incorrect. Thus, the accuracy rate or error rate of the base category of this group can be calculated.
[0077] In related examples, the Q value corresponding to the logarithm of the true error rate of a group or segment of bases is called the true Q value. For example, as mentioned above, the base categories in the first prediction result are divided into a segments according to the P0 value. Then, the total error rate of the bases in each segment can be calculated, and the corresponding Q value can be calculated based on this total error rate. This total error is the so-called true error rate, and the corresponding Q value is the so-called true Q value. Specifically, following the above example, find all the bases (base categories) with P0 in the range of [0, 0.2). Taking the total number of bases in this segment as the denominator and the number of bases with different or inconsistent base recognition results and reference answers (benchmark true value bases) as the numerator, calculate the total error rate of the base category corresponding to P0 in the range of [0, 0.2), and take the logarithm of this total error rate to convert it into a Q value. In this way, the true Q value of the base category corresponding to this segment is determined. Further, the true Q value can be assigned to each base in this segment. In this way, the true Q value or true error rate of each base category in this group or segment is determined.
[0078] In some preferred examples, the setting or adjustment of the size of a follows the following principle: after dividing P0 into a segments, the difference in the true quality values (true Q values) of the base categories corresponding to adjacent segments is not significant. Specifically, the difference in the true quality values of the base categories in adjacent segments is less than 2, which can be regarded as not significant. For example, any value within the range of [0.5, 2) can be regarded as "not significant". For example, the difference is less than or not greater than 0.5, 0.8, 1, 1.1, 1.2, 1.5, etc. More preferably, the difference in the true Q values of the bases detected in adjacent segments is less than 1, which is regarded as not significant. If the preset limit is exceeded, P0 can be re-segmented or further segmented, that is, a larger value of a is taken. Following this principle for multiple tests, the value of a is usually much greater than 10. The inventor comprehensively tests the quality requirements of sequencing data for various detection applications and the sequencing data obtained by sequencing multiple samples including human samples and non-human samples using sequencing platforms based on different sequencing principles or different manufacturers' sequencing platforms to establish such preferred principles or related setting or value-taking rules. In this way, it is beneficial to quickly and automatically generate a better model or a model that meets the user's input tasks. The system 100 can, for example, determine a more suitable a or an a that conforms to the principle through iterative operations.
[0079] Compared with the method of directly or first segmenting the classification probability P0 to directly determine the corresponding segmented base (base category) in the above-mentioned related examples, in some other specific examples, the base categories in the first prediction result are first classified. The bases can be classified according to the current (round) base recognition type, the first n (round) base recognition types, the last m (round) base recognition types, or any combination of any two or all three of them. m and n are each independently natural numbers. For example, classify according to the base categories such as A, C, G, and T recognized in the current (current round), or classify according to the base categories such as CC, AA, AT, CG, etc. recognized in the current (round) base and the base categories recognized in the previous round or the next round, or classify according to the base categories such as ATC, TAT, TTT, GCA, etc. recognized in the current (round) base and the previous and next (round) bases, or classify the base categories according to any combination of the above two or more, and so on.
[0080] In the related examples of first classifying the base categories in the prediction result, the classified base categories are often directly referred to as base categories. If there may be ambiguity or confusion, sometimes the classified base categories are also referred to as new base categories. The classification probability of the classified base categories can be the classification probability of the current (round) base category, for example, still P0 in the above-mentioned related embodiments, or the corrected P0. For example, correct or calibrate P0 based on the classification probability or base recognition probability of non-current bases such as previous bases and / or subsequent bases in the new base category. It is also possible to correct the mapping relationship between P0 and the Q value. Specifically, for example, the classification probabilities of non-current bases in this base classification are relatively low, such as all less than 0.8. To some extent, it can be regarded that the current base category is "dragged down" by one or some of the previous or subsequent and relatively unreliable bases, resulting in a decrease in the credibility of this base category. The P0 can be adjusted or the mapping relationship between P0 and the Q value can be adjusted. For example, multiply by a coefficient less than 1, such as 0.9, to update P0, or subtract 1 or 2 from the Q value corresponding to P0 or assign it to the adjacent lower-level segmented Q value to correct the classification probability of this base category or the mapping relationship between the classification probability of this base category and the Q value. Or, remove the bases with relatively low classification probability values (relatively unreliable bases) in this base category to update this base category, and replace this base category with the updated base category to enter the subsequent analysis process. Or, even directly remove this base category so that this base classification does not participate in the subsequent process. In this way, it helps to establish or generate a more accurate correspondence.
[0081] Furthermore, the subsequent processing of the method of classifying the base categories in the first prediction result can be similar to the methods or approaches in the previous examples. For example, segmenting based on the classification probability P0 of the classified base categories, calculating the true error rate and / or true Q-value of the bases in each segment, and counting or clarifying the number or distribution of the base categories of each true error rate and / or true Q-value, etc., to establish the correspondence between each classification probability and the predetermined quality score. The processes, methods, and technical effects of the foregoing related examples are equally applicable to the processing of such new base categories and associated data, and will not be elaborated here.
[0082] The so-called predetermined quality score refers to the requirements for segmented values of the predetermined Q-value. For example, it comes from user input, parameters that can be set and changed by the user of the system 100, or parameters that the system 100 will automatically select or adjust according to the tasks input by the user. Currently, two common types of Q-values calculated based on the Phred algorithm or output by various sequencing platforms are consecutive integers and non-consecutive integers. Specifically, the predetermined output Q-value is consecutive integers. For example, the Q-value takes all integers within the range of [0, 40], all integers within the range of [0, 45], or all integers within the range of [0, 50], etc.; the predetermined output Q-value is non-consecutive integers, also known as the Q-score binning method, which divides the continuous Q-values into several intervals (bins). For example, the Q-values form an arithmetic sequence or consist of non-consecutive integers without an obvious rule. For example, they are 0, 10, 20, 30, 40, and 50, or for example, they are composed of 0, 10, 17, 21, 28, 35, 40, and 50. Both continuous Q-values and non-continuous Q-values are widely used in current commercially available sequencing platforms. Relatively speaking, non-continuous Q-values are more common. Non-continuous Q-values can simplify data processing, improve robustness, and facilitate visualization, significantly enhancing the efficiency and stability of data analysis. However, non-continuous Q-values may also cause some information loss. In specific applications, which way of presenting the Q-value is suitable or expected generally requires weighing the pros and cons.
[0083] In one example, the predetermined quality score is consistent with the true Q-value in the above-related examples. That is, the true Q-value meets the requirements for the segmented values of the predetermined quality score. Thus, the correspondence between the classification probabilities of each base category and the predetermined quality score established is the so-called quality score model established.
[0084] In some other examples, the predetermined quality score is a set of value ranges different from the true Q value. Usually, it is necessary to further update such a corresponding relationship or mapping relationship based on the established corresponding relationship between the classification probability of each base or each group of bases and the true Q value, and in combination with the quantity of bases with each true Q value / true error rate, such as the quantity proportion or distribution, and the new Q value value range requirements, so as to establish or update the corresponding relationship, that is, to establish or update a quality score model that meets the user input tasks or requirements.
[0085] Specifically, in some specific examples where the predetermined quality score is different from the true Q value and the value range of the true Q value includes the value range of the predetermined quality score, the so-called association of the predetermined quality score with the error rate of the base category includes, based on the error rate of the base category of each grouped first prediction result, the quantity proportion or distribution of the bases in each group, and the predetermined quality score, merging the groups into new groups with fewer group numbers, adjusting the grouping of the base categories of adjacent new groups, and calculating the error rate of the base category of each new group, so that the error rate of the base category of the final new group corresponds to the predetermined quality score, in order to establish a mapping relationship between the classification probability of each base category and the predetermined quality score.
[0086] It can be understood that the so-called predetermined quality score, that is, the value range requirement of the predetermined Q value, is equivalent to predetermining a new number of segments and the error rate of each new segment. By merging one or more adjacent segments into new segments with fewer numbers, and allocating the base ownership of adjacent new segments (equivalent to adjusting the quantity proportion or distribution of the base categories in each segment) and calculating the error rate of the base category of this new segment, finally, the error rate of this segment calculated based on the P0 of the bases in each new segment is the predetermined error rate of this segment (the error rate or new Q value value range of this new segment). In this way, a mapping relationship between the classification probability of each base category and the predetermined quality score is established, and the so-called quality score model is established or generated.
[0087] Specifically, in an example, based on the statistics of the second training data, including the true Q value (true error rate) of each base (base category) in each original group or segment, and the distribution or distribution trend of the base quantity or quantity proportion in each group or segment, and this new quality score requirement (predetermined quality score), merge the bases in the original groups or segments and group or segment them again, and make the error rate corresponding to the bases in each segment after the re-grouping or re-segmenting (the true error rate of the bases in this segment after the re-segmenting) consistent with the predetermined quality score, so as to establish a corresponding relationship between the new groups of bases and the predetermined quality score, in order to update and generate the desired quality score model.
[0088] In system 100, the implementation of what is referred to herein as merging one or more adjacent segments into a smaller number of new segments, allocating the bases among the adjacent new segments, and calculating the error rate of the corresponding new segments can be achieved through iterative calculation. For example, if it is intended to divide segments with one integer per segment, such as consecutive integers 0 - 50, into four new segments 10, 20, 30, and 40, in the first division, the corresponding true Q for the new segment 10 may be 11, and the true Q value for the new segment 20 may be 19. This indicates that the new segments are not reasonable enough or do not meet the predetermined requirements. Therefore, the bases in the adjacent new segments are reallocated again. For example, the lowest score (corresponding to the base) that was assigned to the new segment 20 in the previous division is reallocated to the new segment 10 in the previous division, so as to reduce the true Q value corresponding to the new segment 10 and increase the true Q value corresponding to the new segment 20. Iterative updates are performed in this way to make the true Q value / true error rate corresponding to the final new segments consistent with the predetermined quality score.
[0089] It can be understood that the calculation or determination of the correct rate or accuracy rate, as referred to in the relevant embodiments, has the same meaning as the calculation or determination of the error rate. The correct rate and the error rate are corresponding and can be converted into each other. In addition, the correct rate and the correct probability or accuracy rate or accurate probability in the relevant embodiments, unless otherwise specified, generally have the same meaning; the corresponding error rate and error probability and error occurrence probability, etc., are also the same and can be replaced unless otherwise specified.
[0090] In some specific examples, it further includes merging and re - grouping the groups based on the error rate of the base categories in the first prediction results of each group and the proportion or distribution of the number of base categories in each group, so that the error rate of the base categories in the new groups corresponds to the first quality score, in order to establish a mapping relationship between the classification probability of each base category and the first quality score; and merging and re - grouping the new groups based on the error rate of the base categories in the new groups and the proportion or distribution of the number of base categories in each new group, so that the error rate of the base categories in the final groups corresponds to the predetermined quality score, in order to establish a mapping relationship between the classification probability of each base category and the predetermined quality score, where the values in the first quality score include all the values of the predetermined quality score. In this way, two mapping relationships can be directly generated, and two corresponding quality score models or a cascaded quality score model can be output for the user to compare and select.
[0091] The first quality score is a specified quality score, which can be referred to as the standard quality score. In one example, the standard quality score comes from parameters that can be selectively input by the user, such as parameters that the user can set, adjust, or input by themselves. In another example, the standard quality score or the corresponding mapping relationship is an optional parameter built into system 100 and can establish a mapping, which is a parameter or relationship that system 100 can judge whether to select or learn according to the task input by the user.
[0092] In a specific example, the value of the first quality score is a continuous integer, such as a continuous integer Q-value that adapts to or can cover the Q-values of most current sequencing platforms, such as the continuous integers in [0, 45] or [0, 50], while the predetermined quality score is a non-continuous integer, such as to meet or adapt to the tasks input by the user or the detection requirements of specified applications, etc. If the predetermined Q-value is the same, and the upstream machine learning base recognition model has changed or it is unknown whether the upstream machine learning base recognition model is the same, then in most cases, only the first mapping relationship, that is, the mapping relationship between the classification probability of the previous base category and the standard quality score, needs to be updated and adjusted, without adjusting the subsequent mapping relationship. In this way, a quality score model can be quickly generated, facilitating the conversion and evaluation of the actual quality of sequencing data with the same or different Q-values from different sequencing platforms, etc.
[0093] Moreover, it can also quickly and conveniently meet some special Q-value segmentation requirements. Specifically, the first mapping relationship (the previous mapping relationship) maps all bases to the standard segmentation interval (standard quality score) in any scenario. Therefore, these example solutions can provide standard intermediate results during the performance verification of the Q-value algorithm and the troubleshooting of algorithm problems, which is helpful for performance verification and problem troubleshooting. For example, based on the performance of the previous mapping relationship, it can be evaluated whether the mapping relationship from P0 to the continuous integer Q-value in [0, 50] is reasonable, and based on the performance of the subsequent mapping relationship, it can be evaluated whether the segmented mapping from the continuous integer Q-value in [0, 50] to the non-continuous integer Q-value is reasonable, etc.
[0094] In some examples, the output layer 160 is also used to output the quality score model in any of the above-related examples.
[0095] In certain examples, please refer to Figure 6 , the evaluation module 146 is also connected to the second training data generation module 152, the second training module 154, and the output layer 160, and is used to evaluate the performance of the quality score model using the second test set before outputting the quality score model in any of the above-related examples, so as to determine that these quality score models meet the preset requirements. Among them, the second test set is a part of the second training data that does not overlap with the second training set and the second validation set. The so-called meeting the preset requirements means that at least one of the accuracy rate, precision rate, recall rate, and speed of the classification prediction results of the quality score model for the second test set meets the preset standard or can achieve the tasks input by the user.
[0096] In certain embodiments, the specified model or the first specified model is selected from at least one of logistic regression, support vector machine, tree models such as lightGBT, and neural networks such as semantic segmentation models. The specified model or the first specified model can be a single model, or a cascade model or composite model containing multiple single models.
[0097] In some embodiments, please refer to Figure 7 , the machine learning layer 140 further includes a selection module 148 for determining a specified model, or a first specified model, or a second specified model based on user input. It can be that the user directly selects a certain specified model through user input, or the task system 100 automatically selects a specified model for the user according to the user input.
[0098] In addition, some embodiments of the present application provide a method for developing a base recognition model. The method includes using the system 100 for automatically generating a base recognition model in any of the above embodiments to access user input, and using at least one hardware processor to run the system 100 to present the best base recognition model to the user. The user input includes sequencing data and optionally one or more tasks. The so-called sequencing data is obtained by detecting nucleic acid molecules connected to multiple positions on the surface through surface imaging sequencing technology.
[0099] It can be understood that the descriptions of the technical features, additional technical features, and related technical effects of the system 100 for automatically generating a base recognition model in the above related embodiments also apply to the method for developing a base recognition model in this embodiment, and will not be repeated here.
[0100] For example, in some embodiments, the sequencing data includes images of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs.
[0101] In other embodiments, the sequencing data is a part of the data determined by extracting from the sequencing image. The sequencing data includes the positions and intensity values of multiple nucleic acid molecules in one or more surface regions where they are located. The sequencing image is an image of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs.
[0102] In some examples, the sequencing data can reflect no less than 10 million sequence reads and / or no less than 1 billion base reads.
[0103] In some embodiments, user input is received through a model generation framework or interface; and / or, the best base recognition model is presented to the user through a model generation framework or interface.
[0104] In some embodiments, the system 100 is run using two or more operating systems, and at least one of file sharing, remote access, network communication, virtual machines, containers, and virtualization technologies or tools is used to achieve the interaction between different operating systems, so as to quickly train and generate relevant models.
[0105] In some specific examples, a part of the system 100 is run using Windows and Linux respectively, and the interaction between different operating systems is implemented using at least one of WSL and / or SSH.
[0106] SSH (Secure Shell Protocol), which belongs to one of the "remote access" methods (for operating or accessing another operating system on one operating system), is applicable to the remote management of Linux and macOS and can be accessed from Windows through tools such as PuTTY and Termius. WSL (Windows Subsystem for Linux), which belongs to one of the "virtual file system and bridging tools" (bridging files or processes of different operating systems through virtualization technology), can implement running a Linux environment on Windows to achieve file and tool sharing.
[0107] In some embodiments, the system 100 is run using multi-threaded parallel computing. In this way, the speed of the generation model is improved. Specifically, in one example, the system 100, such as the training data generation module 142 and / or 152 in its machine learning layer 140, etc., runs interactively on multiple operating systems, and a multi-threaded parallel mechanism is set on any one of the operating systems.
[0108] Some embodiments of the present application also provide a base recognition model, which is generated using the system 100 for automatically generating a base recognition model in any of the above related embodiments.
[0109] Some other embodiments of the present application also provide a computer-readable storage medium, which stores one or more instructions for controlling a processor to execute the method for developing a base recognition model in any of the above embodiments. The computer-readable storage medium includes but is not limited to random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium well-known in the technical field.
[0110] An ordered list of executable instructions for implementing a logical function, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with such instruction execution systems, apparatuses, or devices. In related embodiments, the so-called computer-readable storage medium can be any device that can contain, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device or in combination with such instruction execution systems, apparatuses, or devices. More specific examples (non-exhaustive list) of computer-readable storage media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or other suitable media on which the aforementioned program can be printed, because the aforementioned program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then storing it in a computer memory. The various computer-readable storage media described in the related examples can represent one or more devices and / or other machine-readable storage media for storing information. The so-called "machine-readable storage medium" here can include, but is not limited to, wireless channels and various other media that can store, contain, and / or carry instructions and / or data.
[0111] Embodiments of the present application also provide a system for automatically generating a base recognition model, which is also referred to as a computer program product or a computing device. The system includes a memory for storing a program and one or more processors coupled to the memory. The processor runs the so-called program to develop a base recognition model in any of the related embodiments.
[0112] When using a software implementation method, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more instructions or programs. When these computer programs or instructions are loaded and executed on a computer, the processes or functions in the related embodiments of the present application are implemented in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0113] The so-called computer program product or computing device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The computing device may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components, the connections and associations between the components, and the functions of the components shown in the embodiments of the present application are only examples and are not intended to limit the description of the embodiments of the present application and / or related implementation manners.
[0114] Please refer Figure 8 , the system 100 or the electronic device 500 for automatically generating a base recognition model includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.
[0115] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0116] The processor or computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, the CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the modules, methods, or processes described in the relevant embodiments, such as implementing the functions of the machine learning layer 140, etc.
[0117] For example, in some examples, the method for developing a base recognition model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the methods in the above-mentioned relevant embodiments can be executed. In other embodiments, the computing unit 501 can be configured to execute the method for developing a base recognition model using an automatic training tool in any of the aforementioned relevant embodiments in any other suitable manner (such as by means of firmware).
[0118] It should be understood that each part of the system or method in the embodiments of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-mentioned relevant embodiments or examples, multiple layers, modules, sub-modules, or steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, any one or combination of the following related prior arts in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGA), field-programmable gate arrays (FPGA), etc.
[0119] Those of ordinary skill in the art can understand that all or part of the steps of running the system for automatically generating a model in any of the above embodiments or implementing the method in any of the above relevant embodiments can be completed by a program instructing the relevant hardware. The so-called program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the example method.
[0120] In addition, each functional unit in each embodiment or example may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0121] Public solutions for constructing a base recognition model based on machine learning are, for example, CN118429967A, CN118429965A, and CN118969086A, etc. Public solutions for establishing a Q-value model based on machine learning are, for example, CN118471340A, CN117912550A, CN117976042A, and CN118116469A, etc. The system 100 for automatically generating a base recognition model of the present application is applicable to the automatic development of any publicly disclosed base recognition model and / or Q-value algorithm based on machine learning, can automate the development process of machine learning models, and reduce the dependence on professional knowledge and experience backgrounds such as algorithm development. The establishment, deployment, and use of the system (automatic development tool) for automatically generating a base recognition model are described below by way of example in combination with one or more publicly disclosed solutions or related optimization and improvement solutions for constructing a base recognition model or Q-value algorithm based on machine learning.
[0122] The following further exemplarily describes the specific operation methods or implementation processes that some of the above embodiments or examples may include. However, this should not be construed as limiting the scope of the subject matter of the present application to the following further exemplary descriptions or embodiments.
[0123] Establishment and evaluation of a base recognition model based on machine learning and the corresponding Q-value algorithm (I) Establishment of a machine learning base recognition model
[0124] It mainly includes the following steps:
[0125] 1. Determine the usage scenario of the base recognition model to be built, and determine the corresponding platform, instrument, biochemical, fluid, etc. versions of the model. Here, take the multi-fluorescence channel based on surface imaging, such as the commonly referred to dual-color or four-color fluorescence high-throughput sequencing platform, as an example.
[0126] 2. Under the determined operating environment and conditions of 1, sequence the standard library samples of the known reference genome and save the original sequencing data, such as the original fluorescence image (sequencing image).
[0127] 3. Establishment of base recognition model training data (hereinafter referred to as training data I):
[0128] (1) Use the existing basic version of base recognition software, such as the base recognition software (which may or may not involve machine learning) built into each sequencing platform, to analyze the sequencing data and perform base recognition, and output a base recognition result file (fastq file). When using the basic version of base recognition software for base recognition, retain the cluster localization information during the base recognition process, as well as the fluorescence brightness information (intensity information) of each fluorescence channel for each round of each cluster (amplicon) for subsequent extraction of model feature values.
[0129] (2) Align the fastq file with the reference genome to obtain an alignment result file (sam or bam file).
[0130] (3) Based on the brightness information of each round of each cluster, optionally including the brightness information of the first N and last M rounds (N and M are independent integers respectively), and optionally other information such as the brightest base in the first N and last M rounds, the position information of the cluster, etc., make feature values.
[0131] (4) Make target values based on the information after sequence alignment of the clusters. For example, the target value is the reference base corresponding to each base after alignment, sometimes also called the ground truth.
[0132] (5) Correspond the feature values with the target values one by one to complete the production of training data I. When producing training data I, only the sequences with relatively high alignment confidence can be retained. In addition, preferably, polymorphic sites are not included in the scope of training data I. In short, when screening training data, ensure that the "target value" is accurate as much as possible, which is beneficial to training a model with more accurate prediction results. In an example, for the data with a relatively high probability of incorrect target values, either do not include it in training data I, or mark the target value as N, where N represents one of A, G, C, and T; in a scenario of model training using deep learning such as neural networks, the feature values and target values corresponding to all clusters are written into training data I.
[0133] 4. Take a part of training data I to train the base recognition model. For example, in this example, divide training data I into three parts, one for the development of the base recognition model (this part of training data I is briefly denoted as I.1), one for the development of the corresponding Q-value model (this part of training data I is briefly denoted as I.2), and one for verification (this part of training data I is briefly denoted as I.3).
[0134] 5. Establish a base recognition model: Randomly divide the training data I.1 into a training set and a test set, and establish a base recognition model based on the data set (the base recognition model is briefly denoted as model B0). The so-called "base recognition model" refers to the correspondence between the "feature values" and the "target values" mentioned above. When establishing the model, any machine learning or deep learning algorithm can be used. In this example, the model algorithm disclosed in CN119479831A is used, and its full text is incorporated herein by reference.
[0135] In paired-end (PE) sequencing, in some preferred examples, there are often some differences in the biochemical characteristics between R1 (read1) and R2 (read2). Therefore, it is preferably to extract features and train the model for each of them separately. Similarly, in some scenarios involving double-sided chip sequencing, the biochemical and optical characteristics of the two surfaces may be different. Therefore, it is preferably to establish models independently based on the detections or readings of each surface, which generally helps to obtain a model with better prediction effects. Here, taking the establishment of one of the models as an example, the other model is similar and will not be repeated hereinafter.
[0136] (II) Establishment of the Q-value algorithm corresponding to the base recognition model B0
[0137] 1. Establish Q-value algorithm training data based on training data I.2 (hereinafter referred to as training data II for the Q-value algorithm training data set)
[0138] (1) Use the established base recognition model B0 above to complete base recognition based on the feature values of training data I.2, and obtain the base recognition probability corresponding to the model, that is, the normalized recognition probabilities of A, G, C, and T output after the model prediction for each piece of feature data. The sum of the four base recognition probabilities for each piece of data is 1; and the base recognition result predicted by the model, usually the base corresponding to the maximum base recognition probability in each piece of data.
[0139] (2) Compare the above model base recognition results with the target values of training data I.2. Those that are consistent are marked as 1 or correct, and those that are inconsistent are marked as 0 or wrong. This correct / error information serves as the target value of training data II.
[0140] (3) The feature values of training data II can be exactly the same as those of training data I.2; or they can be the base recognition probabilities corresponding to each piece of data; or one or several of the four base recognition probabilities, for example, the feature value is the maximum base recognition probability, or the maximum and the second largest base recognition probabilities, etc.; or they can be the derivative calculation results based on the four base recognition probabilities, such as the ratio of the maximum and the second largest base recognition probabilities; or any combination of any part of the above various parameters.
[0141] 2. Establish the Q-value algorithm based on Training Data II (hereinafter referred to as the basic solution)
[0142] a) Conduct machine learning classification algorithm training based on Training Data II. Any suitable machine learning model can be selected, such as logistic regression, support vector machine, tree model, etc. The obtained model is denoted as Model Q0.
[0143] b) Apply the Q0 model to Training Data II, that is, use the Q0 model to obtain the model prediction result based on the eigenvalue of Training Data II. Here, the final model classification result may not be output, and only the classification probability P-value given during model prediction can be output, denoted as P0.
[0144] c) Usually, P0 is a probability value with a numerical value between 0 and 1. Divide P0 evenly into a segments, and the bases in Training Data II are correspondingly divided into a segments. The principle for choosing a is that after evenly dividing P0 into a segments, the difference in the true Q-values of the bases corresponding to adjacent segments is not significant. In the actual operation of the example, it is generally required that the difference in the true Q-values of adjacent segments is less than 1.
[0145] Here, some explanations are made for the "true Q-value". In fact, the Q-value is another representation of the error rate. The relationship between Q and the error rate can be written as Q = -10 * log10(err). Here, Q refers to the sequencing quality score of the base, that is, the Q-value; err refers to the error probability or error rate of this base. The "true Q-value" refers to the Q-value calculated based on the true error rate (err) of a group of bases. For example, as mentioned above, when dividing the bases into a segments according to the P0 value, the total error rate of the bases in each segment can be calculated. For example, calculate the error rate of this segment with the number of inconsistent bases as the numerator and the total number of bases in this segment as the denominator, and calculate the corresponding Q-value based on this error rate. This Q-value is the "true Q-value" defined here.
[0146] d) Write out the true Q-values corresponding to the a segments of bases respectively.
[0147] e) Define the segmentation requirements for the Q-value according to the algorithm development requirements. For example, in some scenarios, it is required that the value range of Q is all integers within [0, 45], and in some scenarios, it is required that the value of Q is 5 specific values within [0, 45], etc.
[0148] f) According to the above requirements for the value of Q, complete the segmentation of Q values. In the above steps, based on the statistics of training data II, the true Q values (true error rates) corresponding to the bases in each of the a segments are known; at the same time, the distribution trend of the proportion of the number of bases in different segments is also known. Therefore, based on the error rates and the proportion of the number of bases corresponding to each of the a segments, as well as the error rate requirements in the target Q value segmentation (i.e., the value of Q mentioned above), the bases in the above a segments can be merged and segmented again. After the merging and segmentation, the total error rate corresponding to the bases in each segment is consistent with the pre - required Q value.
[0149] g) Record the corresponding relationship between the final P0 values and Q values (segmented Q values). This relationship is the Q value algorithm corresponding to the machine learning algorithm.
[0150] 3. Establish a Q value algorithm based on training data II (hereinafter referred to as Supplementary Scheme 1)
[0151] a) The first few steps are the same as a) and b) of the basic scheme.
[0152] b) Corresponding to step c) of the basic scheme, when segmenting bases according to P0 values, the categories of base recognition are also considered to distinguish the bases. For example, in the basic scheme, all bases are grouped together and divided into a segments according to P0 values. In Supplementary Scheme 1, the bases are first classified according to the final recognition type. For example, they can be classified according to the categories A, C, G, and T of the current base recognition; or they can be classified according to the categories of the current base and the front and back bases, such as ATC, TAT..., and so on. In short, the bases can be classified according to the current base recognition type, the recognition types of the first n bases, the recognition types of the last m bases, or the pairwise combinations or the combination of the three of the above, where n and m are integers independently.
[0153] c) After completing the above classification, perform operations on each type of classified base according to d) - g) of the basic scheme.
[0154] d) Record the corresponding relationship between the various P0 values and Q values (segmented Q values) corresponding to each base classification finally. This corresponding relationship is the so - called corresponding Q value algorithm.
[0155] Supplementary Scheme 1, compared with the basic scheme, takes into account the potential P0 value preference problems that may exist between different base combinations and corrects them when mapping Q values. The preference of P0 values for base combinations mainly comes from biochemical preference characteristics, that is, the error rates corresponding to bases with the same luminance signal characteristics but different base combinations may vary greatly.
[0156] 4. Establish a Q value algorithm based on training data II (hereinafter referred to as Supplementary Scheme 2)
[0157] Either the basic solution or Supplemental Solution 1 can be superimposed with Supplemental Solution 2. One of the features of Supplemental Solution 2 is that the segmentation of the Q value is divided into two steps.
[0158] Specifically, in steps e)-f) of the basic solution or the corresponding steps of Supplemental Solution 1, the bases segmented according to the P0 value are directly mapped to the final Q value segmentation. In Supplemental Solution 2, all bases are first mapped to the Q value segmentation of all integer values within the range [0, 50].
[0159] Specifically, in the example basic solution, the bases segmented according to the P0 value are directly mapped to the Q value segmentation of all integer values within the range [0, 50], for example. In Supplemental Solution 1, the bases in different groups are first mapped to the Q value segmentation of all integer values within the range [0, 50], and then all base categories are combined according to the Q value selection or segmentation requirements. Supplemental Solution 2, in short, is to obtain the mapping relationship between the P0 values corresponding to different categories of bases and the Q value segmentation of all integer values within the range [0, 50], which is simply referred to as Mapping Relationship 2.1 here.
[0160] After obtaining Mapping Relationship 2.1, it is then mapped from the Q value segmentation of all integer values within the range [0, 50] to the final required Q value segmentation, which is simply referred to as Mapping Relationship 2.2 here. Mapping Relationship 2.1 and Mapping Relationship 2.2 together constitute the corresponding Q value algorithm.
[0161] One of the advantages of Supplemental Solution 2 is that Mapping Relationship 2.1 is implemented with the same logic in any scenario, which is beneficial to the implementation of automatic modeling software. If there are different Q value segmentation requirements in different scenarios, generally, only Mapping Relationship 2.2 needs to be adjusted. If the Q value segmentation requirements are the same, and the upstream machine learning base recognition model has changed, then in most cases, only Mapping Relationship 2.1 needs to be adjusted, and usually Mapping Relationship 2.2 does not need to be adjusted.
[0162] Since Mapping Relationship 2.1 maps all bases to the standard segmentation interval in any scenario, Solution 2 can provide standard intermediate results during the performance verification of the Q value algorithm and the troubleshooting of algorithm problems, which is helpful for performance verification and problem troubleshooting. For example, based on the performance of Mapping Relationship 2.1, it can be evaluated whether the mapping relationship from P0 to the integer Q value within the range [0, 50] is reasonable, and based on the performance of Mapping Relationship 2.2, it can be evaluated whether the segmentation mapping from the integer Q value within the range [0, 50] to the large-segment Q value is reasonable.
[0163] The advantage of Supplementary Solution 2 also lies in its ability to conveniently meet some special Q-value segmentation requirements. The bases corresponding to the subdivided Q-value segments (for example, the bases corresponding to integer Q-values in the range [0, 50], and also for example, the mapping association method from the bases directly segmented according to P0 in the basic solution to the final target Q-value segmentation is often not unique).
[0164] First of all, for example, in the product development stage, the final target Q-value is often given in the form of an interval. For example, it is required to "determine a value as the Q-value within the intervals (0, 15), (15, 20), (20, 30), (30, 40), and (40, 50) respectively", and the specific values within each interval are often not pre-defined. It can be determined according to the specific base classification situation during the development process, so there can be many Q-value segmentations that meet the requirements. Secondly, even if the specific value of the final Q-value is clear, there is more than one way to divide the bases that meet the requirements. In this case, Mapping Relationship 2.1 of Supplementary Solution 2 maps the bases to the standard Q-value segments, which is more conducive to Mapping Relationship 2.2 to meet special Q-value segmentation requirements, such as base division requirements like "the Q35 segment cannot contain bases with Q-values lower than 30".
[0165] (III) Model Performance Evaluation
[0166] a) Use the currently established base recognition model B0 and the corresponding Q-value algorithm to conduct tests on the validation set I.3.
[0167] b) Evaluation of the base recognition model B0
[0168] Based on the current round of brightness information in the I.3 eigenvalue, the base recognition result under the traditional algorithm (not based on machine learning model prediction) can be obtained based on the brightest base, and this result is denoted as the base recognition result 1 of I.3. Based on the eigenvalue of I.3, use the base recognition model B0 for prediction, and the obtained result is denoted as the base recognition result 2. Compare the base recognition result 1 and the base recognition result 2 with the target value of I.3 respectively, and calculate the error rate respectively (that is, the number of bases where the base recognition result is different from the target value divided by the total number of entries in the I.3 data). The reduction in the error rate of the base recognition result 2 compared to the base recognition result 1 is the improvement in the base recognition accuracy brought by the base recognition model B0.
[0169] c) Performance evaluation of the Q-value algorithm
[0170] Compare the base recognition result 2 with the target value of I.3 to obtain the correct / incorrect result of each data entry under the base recognition model B0, denoted as the base correct / incorrect result 2.
[0171] For the basic solution, based on the base correctness result 2, output the error rate corresponding to each Q-value segmented base, and convert the actual Q-value based on the error rate. Determine whether the Q-value algorithm meets the expectation by comparing the difference between the Q-value assigned by the algorithm and the actual Q-value.
[0172] For supplementary solution 1, compare the true Q-values corresponding to the bases in different Q-value segments, and also compare the true Q-values corresponding to the bases in different Q-value segments corresponding to different base recognition types, so as to evaluate whether the total base Q-value assignment is reasonable and whether the Q-value assignment of specific types of bases is reasonable / meets the expectation.
[0173] For supplementary solution 2, compare the true Q-values corresponding to the bases in the final Q-value segment, and also compare the true Q-values calculated for the bases corresponding to the integer Q-values in the range [0, 50] generated by mapping relationship 2.1.
[0174] In addition, before developing the base recognition model and Q-value algorithm, it generally involves presetting the passing criteria and qualified thresholds for the performance of the base recognition model and Q-value algorithm. For example, at least one of the accuracy, precision, recall, and speed of the classification prediction results of the base recognition model for the validation set I.3 meets the preset standard or can achieve the tasks input by the user, etc.
[0175] When the performance of the base recognition model or Q-value algorithm does not meet the threshold requirements, it can trigger the mechanism for redeveloping the algorithm or trigger an alarm for manual intervention. For example, remind the user that they can increase user input, such as increasing sequencing data, in order to generate more training data for better model training.
[0176] Finally, output the base recognition model and / or Q-value algorithm that passes or meets the preset metrics or expected requirements after model performance evaluation, in order to complete machine learning modeling.
[0177] Automation of the model development process
[0178] 1. For the above processes of the clear data processing methods and model development logics described above, they can be implemented in any programming language on a Linux system or a Windows system.
[0179] 2. When completing the main process of machine learning model training using the Windows system, system interaction requirements between Windows and Linux may be involved. For example, since most of the current mainstream bioinformatics analysis software is developed on the Linux system, even if there is a Windows version, its running speed is often much slower than the version under the Linux system. In the process of building a base recognition model in the above example, bioinformatics analysis steps such as alignment and error rate statistics are included. If the main process of model development is carried out under Windows, then completing this part of the bioinformatics process under Linux will help improve the efficiency of the automatic training tool.
[0180] Here, the interaction between Windows and Linux systems is achieved through the following two methods respectively.
[0181] (1) WSL2
[0182] Under the Windows system, install the Linux subsystem through the WSL mechanism, and the interaction between Windows and Linux is achieved using WSL. Advantages: The Linux subsystem starts quickly, occupies less resources, and has almost no impact on the performance of the Windows system; WSL2 has a complete Linux kernel and supports the installation of the vast majority of Linux software; WSL2 supports container technologies such as docker / singularity, and software deployment is convenient and fast. Disadvantages: The file mutual access between Windows and Linux is output through the network, and the read and write speed is slower than that of the native system.
[0183] (2) OpenSSH
[0184] Implementation method: Install OpenSSH software on Windows and Linux, and interact through the SSH protocol. Advantages: Based on the SSH protocol, the security is greatly improved and the connection is stable; it is convenient to integrate. By configuring SSH passwordless access, operations such as remote connection, task submission, and task status determination can be performed in a python script; independent Windows and Linux nodes can make the most of the computing power of each node. Disadvantages: The installation and deployment are relatively complex, and the OpenSSH software needs to be deployed on win and Linux respectively, and passwordless access needs to be configured.
[0185] 3. In the model training stage, the running speed of the training tool can be improved through parallel computing.
[0186] For example, during the process of base recognition using the aforementioned (1) 3 (1) utilization basis or traditional version of non-machine learning-based base recognition software, multiple threads can be called through code for simultaneous analysis. When developing an automatic training tool, it is generally necessary to pre-analyze the computing power and memory occupied by the base recognition software in the basic version and write the calculation formula for the maximum number of threads corresponding to different computing powers and memories. When using the automatic modeling tool (automatic training tool), the tool will automatically read the computing power and memory of the system used for modeling and automatically calculate the maximum number of threads that can run in parallel according to the maximum number of threads calculation formula.
[0187] Regarding calculating the maximum number of threads or the optimal number of threads corresponding to specific computing power and memory, those skilled in the art can reasonably calculate and determine by comprehensively considering hardware specifications (such as the number of CPU cores, hardware characteristics of the GPU such as the number of multiprocessors, thread block size, etc., memory capacity and bandwidth) and task types (such as CPU-intensive or IO-intensive). For example, calculate based on a processor such as a CPU and / or GPU, based on memory and / or based on task type, or comprehensively consider thread number calculation in order to reasonably calculate and optimize the number of threads.
[0188] (1) Whether the aforementioned (1) 3 (2) is carried out on a Windows system or a Linux system, a multi-thread parallel mechanism can be set. Similarly, pre-enter the calculation formula for the optimal number of parallel threads, and automatically set the number of parallel threads during the actual operation process.
[0189] (2) The production of training data in the aforementioned (1) 3 (3)-(5) can also set multi-thread parallel calculation according to the above idea.
[0190] (3) The machine learning training step of the aforementioned (1) 5 base recognition model can also set parallel operations. In addition, in the scenario of separately training models for R1 and R2 data based on paired-end (PE) sequencing, the models of R1 and R2 can also be trained in parallel when the computing power and memory permit. Similarly, in some scenarios of separately training models for data from multiple sides such as double-sided chip sequencing, parallel calculation is set for each model training.
[0191] 4. The form of the eigenvalue of the model can be determined before training to improve the data extraction speed. And when performing base recognition using the traditional or basic version algorithm in step (1) 3 (1), directly calculate and output the eigenvalue using the cached data during the base recognition process without retaining other intermediate parameters that are not needed, thereby improving the efficiency of data calculation and reducing the time for data writing.
[0192] Using the automatic training tool
[0193] The following example shows how an automatic training tool outputs a model file based on user input such as sequencing images in a specific scenario.
[0194] Suppose a user needs to use a machine learning model in a certain special application scenario. This scenario may require special sample features, special biochemical and instrument configurations, etc., and no suitable machine learning model has been established in this scenario. Under such a premise, the user can use the system (automatic training process) for automatically generating models by themselves to build a corresponding or suitable model.
[0195] 1. Generally, the user first clarifies the sequencing conditions corresponding to the model to be established, such as biochemistry, instrument configuration, and the characteristics of the library sample, etc. Then the user performs sequencing under the above clarified sequencing conditions to obtain sequencing data for generating training data. When sequencing, for example, a library sample with a known reference sequence is used. If the user does not specifically specify a special library sample in this sequencing scenario, a library sample with a larger genome such as a human genome sample, etc., can be used as the sequencing sample. A sample with a larger genome can make the training data contain more diverse sequence feature information, which usually helps to improve the generalization of the model. If the "special sequencing scenario" required by the user includes special library features, such as a library with extremely unbalanced bases, etc., then the sequencing library here is preferably a library with the same or similar features as the target sequencing library to help the machine learning model learn the sequencing preferences introduced by the library features.
[0196] 2. When the user performs the above sequencing, the machine learning automatic training tool is called. In current mainstream commercially available sequencing platforms that implement sequencing based on surface presentation or common sequencing processes, the sequencer often processes data including base recognition while sequencing. The software in the sequencer for implementing this data processing is the so-called traditional or basic version base recognition algorithm. The so-called traditional or basic version base recognition algorithm or tool may be a machine learning model or a software that does not involve machine learning. Here, the user can directly use the automatic training tool during the sequencing process. The automatic training tool will automatically call the traditional or basic version base recognition algorithm during the sequencing step of the sequencer, and during the process of using the traditional or basic version algorithm to complete base recognition, it can simultaneously generate the feature values required by the machine learning model based on sequencing intermediate parameters such as the brightness of the original image and the base recognition information of the context, etc., and save them in the specified path. It can be understood that during the conventional sequencing process (without calling the automatic training tool), the so-called sequencing intermediate parameters are often the parameter data that are required or usually generated during the base recognition process of the sequencer but generally do not need to be output or output to the user.
[0197] 3. The automatic training process can automatically identify the sequencing progress. For example, after the sequencing is completed, the automatic training tool can automatically call the alignment tool, align the obtained fastq file from sequencing with the reference genome (such as the one input by the user), and output the alignment result to the specified path.
[0198] 4. After the alignment is completed, the automatic training tool will start to automatically generate training data based on the stored feature values and the alignment result. The feature values of the training data have been all output during the sequencing process, and the target values will be extracted based on the alignment result. For details, please refer to part (1) 3 of the foregoing example. The training data will be automatically stored in the specified path.
[0199] 5. After the training dataset is generated, the automatic training tool can, according to the processes described in the foregoing (1) 4 and 5 and (2), automatically divide the training data, automatically complete the modeling of the machine learning model, and automatically complete the development of the Q-value algorithm.
[0200] 6. After the machine learning model is modeled and the Q-value algorithm is developed, the automatic training tool will automatically save the model file and the algorithm file in the specified path. The model files of the machine learning and the Q-value algorithm can appear as the configuration files of the machine learning base recognition tool. As long as the configuration files are stored in the specific folder called by the base recognition tool, the next time the machine learning base recognition tool is used for sequencing, this configuration will be automatically used for base recognition and Q-value calculation. Therefore, after the modeling is completed, the user can directly copy the configuration file to the path of the corresponding configuration file of the base recognition tool on the sequencer, and the model can be directly used for the next sequencing.
[0201] In the description of this specification, the descriptions such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "certain examples", "specific examples" or "embodiments" mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0202] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A system for automatically generating a base recognition model, characterized in that: include: An input layer, for receiving user input including sequencing data and optionally one or more tasks from a user, wherein the sequencing data is obtained by detecting nucleic acid molecules attached to a plurality of positions on a surface using a surface imaging sequencing technology; The machine learning layer includes a training data generation module and a training module. The training data generation module is used to process the user input from the input layer to generate training data, wherein the training data includes a plurality of pairs of associated features and labels, wherein the features describe the intensity corresponding to the nucleic acid molecules at one or more positions, and the labels include one selected from the bases A, T, C and G. The training module is used to train a specified model based on at least a portion of the training data, including performing multiple iterations of training on the specified model using a training set and evaluating the performance of the specified model after each iteration of training using a validation set, so as to obtain a base recognition model, wherein the training set and the validation set are each independently a non-overlapping portion of the training data; as well as The output layer is used to output the base recognition model.
2. The system according to claim 1, characterized in that The sequencing data includes images of nucleic acid molecules at multiple locations on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs; Alternatively, the sequencing data is a portion of data extracted and determined from a sequencing image, the sequencing data includes positions and intensity values of multiple nucleic acid molecules in one or more surface regions where they are located, and the sequencing image is an image of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs; Optionally, the sequencing data can reflect no less than 10 million sequence reads, and / or no less than 1 billion base pairs of base reads; Optionally, the tag further comprises a combination of any two, three or four of the bases selected from A, T, C and G; Optionally, the user input further includes a reference sequence, and the training data generation module further includes an alignment submodule, the alignment submodule being used to align the sequence reads or base reads reflected by the sequencing data to the reference sequence to generate an alignment result; and determining the label based on the alignment result to generate the training data; Optionally, the machine learning layer further includes an evaluation module, which is connected to the training module and the result output layer and is used to evaluate the performance of the base recognition model using a test set before outputting the base recognition model, so as to determine whether the base recognition model meets preset requirements, wherein: The test set is a part of the training data that does not overlap with the training set and the validation set. Meeting the preset requirements means that at least one indicator of the accuracy, precision, recall and speed of the classification prediction results of the test set by the base recognition model meets the preset standards or can achieve the task input by the user.
3. The system according to claim 1 or 2, characterized in that: The training data generation module is a first training data generation module, the training data is the first training data, the training module is the first training module, the specified model is the first specified model, the feature is the first feature, the label is the first label, and the machine learning layer also includes a second training data generation module and a second training module, wherein The second training data generation module is connected to the first training data generation module and the first training module, and is used to process at least a part of the first training data generated by the first prediction result of the base recognition model to generate second training data, wherein the first prediction result includes the probability and base category of the data being predicted and recognized as various base categories in the first label, and the second training data includes a plurality of pairs of associated second features and second labels, wherein the second feature describes the intensity corresponding to the nucleic acid molecule at one or more positions and / or describes the probability of at least one base category in the first prediction result, and the second label includes two categories: the first prediction result is consistent with or inconsistent with the corresponding reference true value base, The second training module is used to train a second designated model based on at least a portion of the second training data, including performing multiple iterations of training on the second designated model using a second training set and evaluating the performance of the second designated model after each iteration of training using a second validation set, so as to obtain a quality score model, wherein a quality score output by the quality score model reflects the accuracy of the first prediction result, and the second training set and the second validation set are each independently a non-overlapping portion of the second training data; Optionally, the second feature describes the probabilities of multiple base categories in the first prediction result.
4. The system according to claim 3, characterized in that Training the second specified model in the second training module also includes, Obtaining the classification probability of the base category of the first prediction result in the second training set; Divide the first prediction results into a groups based on the classification probability and / or the base categories, and for each group, compare the base categories in the first prediction results of the group with the corresponding reference true value base categories to determine the error rate of the base category of the first prediction results of the group, where a is a natural number and the value of a can achieve a logarithm difference of the error rates of the base categories of the first prediction results of adjacent groups less than 2; as well as associating a predetermined quality score with an error rate for the base class; Optionally, further comprising, Based on the error rate of the base category of the first prediction result of each group, the proportion or distribution of the number of base categories of each group, and the predetermined quality score, the groups are merged into new groups with fewer groups, the attribution of the base categories of adjacent new groups is adjusted, and the error rate of the base category of each new group is calculated, so that the error rate of the base category of the final new group corresponds to the predetermined quality score, so as to establish a mapping relationship between the classification probability of each base category and the predetermined quality score; Optionally, further comprising, Based on the error rate of the base category of the first prediction result of each group and the proportion or distribution of the number of base categories of each group, the groups are merged and regrouped so that the error rate of the base category of the new group corresponds to the first quality score, so as to establish a mapping relationship between the classification probability of each base category and the first quality score; Based on the error rates of the base categories of the new groups and the proportion or distribution of the number of base categories of each new group, the new groups are merged and regrouped so that the error rates of the base categories of the final groups correspond to the predetermined quality scores, so as to establish a mapping relationship between the classification probability of each base category and the predetermined quality score, wherein the values in the first quality score include all values of the predetermined quality score; Optionally, the value of the first quality score is a continuous integer, and the predetermined quality score is a non-continuous integer; Optionally, the user input includes the predetermined quality score; Optionally, the user input includes the corresponding benchmark truth base category; Optionally, the output layer is further used to output the quality score model; Optionally, the evaluation module is further connected to the second training module and the result output layer, and is used to evaluate the performance of the quality score model using a second test set before outputting the quality score model, so as to determine whether the quality score model meets preset requirements, wherein: The second test set is a part of the second training data that does not overlap with the second training set and the second validation set, and the satisfying the preset requirement is that at least one indicator of the accuracy, precision, recall and speed of the classification prediction result of the second test set by the quality score model meets the preset standard or can achieve the task input by the user; Optionally, the designated model or the first designated model or the second designated model is selected from at least one of logistic regression, support vector machine, tree model and neural network; Optionally, the machine learning layer further comprises a selection module for determining the designated model or the first designated model or the second designated model based on the user input.
5. A method for developing a base recognition model, characterized in that The method comprises accessing user input using the system for automatically generating a base recognition model according to any one of claims 1 to 4, and operating the system using at least one hardware processor to present an optimal base recognition model to the user, wherein the user input comprises sequencing data and optionally one or more tasks, wherein the sequencing data is obtained by detecting nucleic acid molecules connected to multiple positions on a surface using surface imaging sequencing technology.
6. The method according to claim 5, characterized in that The sequencing data includes images of nucleic acid molecules at multiple locations on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs; Optionally, the sequencing data is a portion of data extracted and determined from a sequencing image, the sequencing data includes positions and intensity values of multiple nucleic acid molecules in one or more surface regions where they are located, and the sequencing image is an image of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs; Optionally, the sequencing data can reflect no less than 10 million sequence reads, and / or no less than 1 billion base pairs of base reads; Optionally, receiving the user input via a model generation framework or interface; and / or presenting the optimal base recognition model to a user via the model generation framework or interface; Optionally, the system is operated using two or more operating systems, and interaction between the different operating systems is achieved using at least one of file sharing, remote access, network communication, virtual machines, containers, and virtualization technologies or tools; Optionally, using Windows and Linux to respectively run a part of the system, and using at least one of WSL and / or SSH to implement interaction between different operating systems; Optionally, the system is run using multi-threaded parallel computing.
7. A base recognition model, characterized in that: Produced using the system for automatically generating base recognition models as described in any one of claims 1-4.
8. A computer-readable storage medium storing a plurality of instructions for controlling a processor to execute the method of claim 5 or 6.
9. A system for automatically generating a base recognition model, characterized in that: The system comprises a memory storing a program and one or more processors coupled to the memory, wherein the processors execute the program to implement the method of claim 5 or 6.
Citation Information
Patent Citations
Self-learning base detector trained with oligonucleotide sequences
CN117546249A
Gene sequencing base quality evaluation method based on deep learning, product, equipment and medium
CN117726621A
Method for determining mass fraction of read segment, sequencing method and device
CN117976042A
Basecaller for DNA sequencing using machine learning
US20150169824A1
Base calling model training method and system, base calling model identification method and system, and device and medium
WO2024124453A1