System and method for automatically generating machine learning model
Through a system that automatically generates machine learning models, automatically process sequencing results and generates quality score models, the problems of manual operation intensive and inefficient in the existing technology are solved, and a simplified model development process and efficient evaluation of base recognition results are achieved.
Patent Information
- Application Number
- CN202510257654.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-03
AI Technical Summary
In the prior art, machine learning models require a lot of manual operation during the development of sequencing, resulting in inefficiency and error-prone, and require high professional knowledge and experience.
It provides a system that automatically generates machine learning models, receives sequencing results and user input through the input layer, and uses training data generation modules and training modules in the machine learning layer to automatically generate quality score models to evaluate the accuracy of base recognition results.
The automated machine learning process is implemented, the modeling process is simplified, and the dependence on professionals is reduced, so that non-machine learning experts can also develop quality score models to evaluate the accuracy and reliability of base detection results.
Smart Images

Figure CN120220826A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing based on artificial intelligence. Specifically, it relates to systems and methods for automatically generating machine learning models, and more specifically, to systems and methods for generating a quality score model or generating a base calling model and a corresponding quality score model based on an automatic training tool. Background Art
[0002] The subject matter discussed in this section should not be regarded as prior art merely because it is mentioned in this section. Similarly, the technical problems mentioned in this section or associated with the subject matter provided as background art should not be regarded as having been recognized in the prior art. The subject matter in this section only represents different methods, which may themselves also correspond to specific embodiments of the technical solutions included in the claims.
[0003] In the related art, sequencing generally includes detecting relevant signals to identify the type of nucleotide or base linked to one or more nucleotide positions in a nucleic acid molecule, that is, determining at least a part of the nucleotide sequence of the nucleic acid molecule through base calling. The change in the signal and / or signal intensity corresponding to a specific position of the nucleic acid molecule to be detected can indicate the type of base at that position on the nucleic acid molecule. For example, different nucleotides can be attached with different fluorescent molecules or optically detectable distinguishable luminescent markers, and in one round of sequencing, multiple nucleotides are allowed to contact the nucleic acid molecule to be detected, and then the luminescent signal is detected to distinguish the type of nucleotide or base linked to the template in this round of sequencing. Typical examples include the sequencing platforms of Illumina or MGI. It is also possible to separate the optical signal detection and nucleotide incorporation or ligation into independent steps to achieve one round of sequencing, such as the sequencing platform of Element Bioscience; there is also, for example, the Ion Torrent sequencing platform of ThermoFisher that does not rely on fluorescence signal or optical signal detection, but uses an electrochemical sensor to detect the pH change caused by the release of H+ ions when nucleotides are incorporated into the nucleic acid molecule to be detected to identify the type of incorporated nucleotide or base; in addition, there are also single molecule sequencing platforms based on various nanopores such as PacBio and Oxford Nanopore. Compared with the above-mentioned sequencing rounds, the sequencing of such platforms is continuous. For example, the fluorescence signal emitted during the process of observing the incorporation of modified nucleotides with fluorescently labeled phosphate groups into a single molecule template (nucleic acid molecule to be detected) using a zero-mode waveguide (ZMW) nanopore is used to identify the type of incorporated base in real time, or the base sequence of the nucleic acid molecule is directly determined by physically detecting the electrical signals generated when different bases pass through the nanopore without involving biochemical enzyme-catalyzed processes such as polymerization reactions.
[0004] Each sequencing platform (also known as a sequencing system or sequencer) commercially available from various companies currently includes base recognition software adapted to that platform, model, or detection principle to identify the base types in the nucleic acid molecule sequence to be tested based on the detection and analysis of relevant signals. Moreover, each sequencing platform generally also includes quality score software to quantitatively evaluate the accuracy or reliability of the bases or base combinations identified and determined through one or more rounds of sequencing or one or more sequencing runs. Related techniques include, for example, the widely accepted Phred quality Q score (also known as the quality fraction or Q value) [Ewing B, Green P. "Base-calling of automated sequencer traces using phred. I. Accuracy assessment." Genome Research. 1998; 8(3): 175-185.], scores such as Q20, Q30, or Q40, where Q = -10log 10 P is logarithmically related to the base classification error probability. The Q value is used to characterize the quality of bases or DNA sequences and can be used to compare the effectiveness of detection results obtained by different sequencing methods or different sequencing platforms; moreover, during base recognition or detection, the advantage of real-time calculation of quality scores also includes the ability to terminate defective sequencing runs early. Early termination of defective or unqualified sequencing runs can significantly save resources.
[0005] In related technologies, machine learning includes various algorithms, such as traditional machine learning algorithms like linear regression, decision trees, support vector machines, and deep learning of neural networks. Applying machine learning to the field of sequencing or gene detection, for example, using machine learning to develop base recognition algorithms, base recognition result classification scores, or quality score algorithms, etc., there have been some related disclosures, such as Wei-Chun Kao, et al.: "naiveBayesCall: An Efficient Model-Based Base-Calling Algorithm for High-Throughput Sequencing", 2010-04-25, RESEARCH IN COMPUTATIONAL MOLECULAR BIOLOGY, SPRINGER BERLIN HEIDELBERG, BERLIN, HEIDELBERG, pages 233-247, XP019141683, ISBN: 978-3-642-12682-6; CN118942549A, CN112789680, CN118429967A, CN118429965A, CN118116469A, CN117976042A, and US10068053B, etc.
[0006] Machine learning models for common application directions in the field of sequencing or sequencing-related fields, including the model training, development, or usage processes involved in the publicly disclosed solutions listed above, often involve a large number of manual operations or interventions. For example, in the early stage, it is often necessary to manually collect and preprocess data, select models, adjust hyperparameters, train models, and evaluate models. Inevitably, manual intervention leads to the need to configure and optimize machine learning models for different real-world scenarios, making manual operations intensive. Generally, intensive manual operations are prone to errors, low efficiency, or difficult management. Moreover, this process usually requires the workers or operators to have rich relevant professional knowledge and experience, such as algorithm development experience or professional background, especially when the machine learning model is complex and involves configuring and optimizing different types of algorithms. Summary of the Invention
[0007] The embodiments of the present application aim to at least partly solve at least one of the above technical problems or at least provide a practical commercial option.
[0008] The embodiments of the present application provide a system for automatically generating a machine learning model. The machine learning model generated by this system outputs a quality score to evaluate the accuracy of base recognition results. The system includes: an input layer for receiving user input from a user, including sequencing results and optionally one or more tasks. The sequencing results are obtained by detecting nucleic acid molecules connected to multiple positions on the surface through surface imaging sequencing technology. The sequencing results include base recognition results and corresponding base recognition probabilities; a machine learning layer including a training data generation module and a training module. The training data generation module is used to process the user input from the input layer to generate training data. The training data includes multiple pairs of associated features and labels. The features describe one or more base recognition probabilities corresponding to base recognition and / or intensities corresponding to nucleic acid molecules at one or more positions. The labels include two categories: whether the base recognition result is consistent or inconsistent with the corresponding ground truth base. The training module is used to train a specified model based on a part of the training data to obtain a quality score model, including performing multiple iterative trainings on the specified model using a training set and evaluating the performance of the specified model after each iterative training using a validation set to obtain a quality score model. The training set and the validation set are each independently a non-overlapping part of the training data; and an output layer for outputting a quality score model that meets preset requirements.
[0009] An embodiment of the present application also provides a method for developing a quality score model. The method includes accessing user input using the system for automatically generating a machine learning model in any of the embodiments, and running the system using at least one hardware processor to present the best quality score model to the user. The so-called user input includes sequencing results and optionally one or more tasks. The sequencing results are obtained by detecting nucleic acid molecules connected to multiple positions on the surface through surface imaging sequencing technology, and the sequencing results include base identification results and base identification probabilities.
[0010] An embodiment of the present application also provides a quality score model, which is generated using the system for automatically generating a machine learning model in any of the embodiments.
[0011] An embodiment of the present application also provides a computer-readable storage medium, which stores a plurality of instructions for controlling a processor to execute the method for developing a quality score model in any of the embodiments.
[0012] An embodiment of the present application also provides a system for automatically generating a machine learning model. The system includes a memory storing a program and one or more processors coupled to the memory. The processor runs the program to implement the method for developing a quality score model in any of the embodiments.
[0013] An embodiment of the present application also provides a sequencing system or sequencing platform, which includes the system for automatically generating a machine learning model in any of the above embodiments or examples.
[0014] Through the system, method or related system platform for automatically generating a machine learning model in any of the above embodiments, an automated machine learning process is realized, which simplifies the modeling process. Moreover, it reduces the dependence on professionals such as algorithm developers, enabling non-machine learning experts, such as users of sequencing platforms without relevant algorithm development backgrounds or experiences, to develop quality score models by themselves to evaluate or compare the accuracy or reliability of base detection results, and better meet various application detection needs or requirements of users based on sequencing.
[0015] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and / or additional aspects and advantages of the embodiments of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:
[0017] Figure 1 is a schematic structural diagram of the system for automatically generating a machine learning model according to an embodiment of the present application;
[0018] Figure 2 It is a schematic diagram showing that in an embodiment of the present application, an entire image is converted into a set of blocks, and the set of blocks is written as a set of matrices as input data;
[0019] Figure 3 It is a schematic structural diagram of a system for automatically generating a machine learning model according to an embodiment of the present application;
[0020] Figure 4 It is a schematic structural diagram of a system for automatically generating a machine learning model according to an embodiment of the present application;
[0021] Figure 5 It is a schematic structural diagram of a system for automatically generating a machine learning model according to an embodiment of the present application;
[0022] Figure 6 It is a schematic structural diagram of a system for automatically generating a machine learning model according to an embodiment of the present application;
[0023] Figure 7 It is a schematic structural diagram of a system for automatically generating a base recognition model according to an embodiment of the present application;
[0024] Figure 8 It is a schematic structural diagram of a system for automatically generating a base recognition model according to an embodiment of the present application;
[0025] Figure 9 It is a schematic structural diagram of a system for automatically generating a base recognition model according to an embodiment of the present application;
[0026] Figure 10 It is a schematic structural diagram of a system for automatically generating a base recognition model according to an embodiment of the present application;
[0027] Figure 11 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0028] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.
[0029] In this article, unless otherwise specified, the singular forms "a", "an", etc. used to limit include their plural referents (one or more than one). "A set" or "a plurality" means two or more than two.
[0030] In this document, terms such as "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity or order of the indicated technical features. In the description of this application, the meaning of "a plurality" is two or more, unless otherwise specifically stated.
[0031] Unless otherwise specified, the terms "connected" and "linked" in this document should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection, or a connection that enables mutual communication; it can be a direct connection or an indirect connection through an intermediate medium, and can be the internal communication of two components or the interaction relationship between two components; it can be a connection through physical adsorption or other forces, or a chemical connection through chemical bonds such as incorporation through a polymerization reaction. Those skilled in the art can understand the specific meaning of this term in the corresponding examples according to the specific embodiments described, including the context and common knowledge.
[0032] In this document, "sequencing" refers to nucleic acid sequence determination, which is the same as "nucleic acid sequencing" or "gene sequencing", and refers to determining the base sequence of the primary structure of a nucleic acid molecule. It can be achieved by methods such as sequencing by synthesis (SBS), sequencing by ligation (SBL), or sequencing by hybridization (SBH). Unless otherwise specified, the so-called sequencing by synthesis in this document, in addition to including the commonly understood SBS that utilizes polymerase to catalyze the incorporation of nucleotides into the nucleic acid molecule to be tested (polymerization reaction) and detects the corresponding reaction signals to identify the types of incorporated nucleotides (the typical ILLUMINA / Solexa technology), also includes sequencing methods similar to SBS that utilize polymerase or non-polymerase to controllably introduce or link nucleotides to the nucleic acid molecule to be tested, directly or indirectly, simultaneously or sequentially detect the corresponding signals to determine the types of one or more nucleotides that are linked, such as SBL, SBH, SBB (sequencing by binding), or SBE (sequencing by expansion) that achieve sequencing based on surface fluorescence imaging detection.
[0033] In the process of sequencing based on solid-phase surface signal detection, sequencing can also be said to measure the intensity values corresponding to the positions of one or more nucleic acid molecules. The intensity values here are sometimes simply referred to as intensities, and can be any signal such as the intensity of electrical or electromagnetic radiation such as visible light. A base can have one intensity value, or multiple intensity values, or multiple bases can have one or fewer intensity values than the number of bases. In addition, the intensity values can be for specific positions on the solid-phase surface or for multiple positions of nucleic acid molecules. The intensity values can be converted or mapped to pre-determined numerical values (such as integers in binary or decimal number systems) or continuous numerical values or a certain range.
[0034] Sequencing can be carried out through a sequencing platform. According to the embodiments of the present application, the selectable sequencing platforms include but are not limited to the Hiseq, Miseq, Nextseq, and Novaseq series sequencing platforms of Illumina, the Ion Torrent platform of Thermo Fisher / Life Technologies, the BGISEQ and MGISEQ / DNBSEQ platforms of BGI Genomics, the AVITI of Element Bioscience, and single-molecule sequencing platforms; the sequencing method can be single-end sequencing, paired-end sequencing, or the sequencing methods supported by the selected automated sequencing platform, etc.
[0035] The "sequencing process" or "sequencing run" in the embodiments of the present application refers to measuring the intensity values corresponding to the positions of one or more nucleic acids in batches. For example, in a scenario where sequencing involves multiple cycles (also called rounds or sequencing rounds) of imaging nucleic acid molecules on a substrate, such as the surface of a solid-phase substrate, after a specified biochemical reaction, a series of intensity values are obtained during the same sequencing run. Generally, the nucleic acid molecules in the first sequencing run will not participate in the second sequencing run independent of the first sequencing run, such as not being included in an image collected at a certain time point (a certain round) of the second sequencing run.
[0036] In some examples, sequencing by synthesis is used for multiple rounds of sequencing to obtain a sequencing sequence or read. For example, the nucleic acid molecule to be tested is contacted with a polymerase and modified nucleotides and placed under conditions suitable for polymerization reaction, and the modified nucleotides are controllably incorporated into the nucleic acid molecule to be tested, or in other words, single-base extension is controllably achieved, and the corresponding reaction signal is detected. Based on this signal, the type of nucleotide incorporated into the nucleic acid molecule to be tested in this reaction is determined. In this way, multiple controllable single-base extensions and detections of corresponding signals are carried out, so as to detect the types of nucleotides or bases incorporated into the nucleic acid molecule to be tested in multiple or multiple rounds of reactions according to the reaction signal information, in order to sequence a part of the nucleic acid molecule to be tested.
[0037] The so-called nucleic acid molecule to be detected, also known as nucleic acid template or template, can be an unamplified single molecule, or a molecule cluster or group or long chain containing multiple identical polynucleotide molecules after amplification, such as the fluorescence clusters (clusters) or DNA nanoballs (DNBs) formed by bridge amplification or rolling circle amplification adopted by current mainstream commercially available sequencing platforms. The nucleic acid molecule to be detected can be presented in the form of single-stranded, double-stranded, and / or triple-stranded or more-stranded complexes hybridized with probes or primers.
[0038] The corresponding reaction signals can be, for example, fluorescence signals, and can be manifested as image data (such as color images or grayscale images) formed by collecting these fluorescence signals. Thus, these image data are processed and analyzed to detect the nucleotides incorporated into the nucleic acid molecule to be detected in each reaction or each round of reaction, so as to determine a part of the base sequence of the nucleic acid molecule to be detected.
[0039] Specifically, in some examples, sequencing is achieved based on surface fluorescence imaging detection. The nucleic acid molecule to be detected is connected to a solid surface. For example, nucleotides can be modified to carry or be capable of binding fluorescence labels, and carry excisable inhibitory groups that can prevent other nucleotides from polymerizing and connecting to the next position of the nucleic acid molecule to be detected (such modified nucleotides are also called reversible terminators). And after each polymerization reaction or single-base extension reaction is completed, it is excited to make the fluorescence label emit light, and these emission signals are collected to obtain an image of the nucleic acid molecule to be detected where single-base extension reaction occurs at a specified surface position; then, the inhibitory group and fluorescence label, etc. are removed to perform the next or the next round of polymerization reaction and signal collection (taking pictures). Repeating such polymerization reaction - taking pictures - excision multiple times or multiple rounds to obtain image information related to the nucleotides connected to the nucleic acid molecule to be detected in each single-base extension reaction.
[0040] It can be understood that when the nucleic acid molecule to be detected at a specified surface position undergoes a polymerization reaction and emits fluorescence, it generally appears as bright spots or bright patches (spots) with a signal intensity higher than the background at the corresponding position in the image collected in this round of reaction. Therefore, according to the information contained in these image sets corresponding to specific chemical characteristics (the nucleic acid molecule to be detected undergoing a polymerization reaction), such as intensity and / or morphology, etc., it can be judged whether there is a nucleic acid molecule to be detected at the specified position and whether the nucleic acid molecule to be detected has undergone a polymerization reaction. Further, in combination with the preset corresponding relationship between fluorescence emission signals and nucleotide types, the type of nucleotide connected to the nucleic acid molecule to be detected through a biochemical reaction can be detected, thereby determining at least a part of the sequence of the nucleic acid molecule to be detected and obtaining the so-called read segment.
[0041] Determining a set of positions corresponding to chemical features (nucleic acid molecules to be detected) on a surface, that is, obtaining a template, can be determined by identifying features corresponding to chemical features on a detection image, such as the positions of bright spots, or can also be determined by identifying other features on the image that have a specific association relationship with the spatial position relationship of chemical features. For example, for a regular array surface containing a labeled region (also often called a tracer region), generally the spatial position relationship between the labeled region and the reaction region (the region where nucleic acid molecules to be detected are located) is preset and known, and the signal from the labeled region and the signal from the reaction region are clearly distinguishable or the signal from the labeled region presents as a detectable feature on the image. By identifying the position of the labeled region on the image, the positions of each amplicon (nucleic acid molecules to be detected) in the reaction region can be determined.
[0042] It should be noted that the nucleotides referred to in this article include ribonucleic acid or deoxyribonucleic acid, including natural nucleotides or their derivatives or modified forms (also called modified nucleotides or modified nucleotides, etc.). In this article, sometimes the base contained in a nucleotide is also used to refer to the nucleotide, and those skilled in the art can clearly understand according to conventional knowledge and / or context. In addition, the nucleic acid sequences, nucleotide sequences, oligonucleotide sequences, or base sequences referred to in this article can be used interchangeably unless otherwise stated, and sometimes are directly abbreviated as sequences. In addition, sometimes one or more nucleotides or bases are also referred to as sequences, and those skilled in the art can clearly understand according to conventional knowledge and / or context.
[0043] In some embodiments, a round of sequencing can include a single base extension reaction (one repeat). For example, four different nucleotides (dATP, dTTP, dGTP, and dCTP) can be placed in the same polymerization reaction system with multiple nucleic acid molecules to be detected, so that each nucleotide can be excited to emit a signal distinguishable from other types of nucleotides, thereby determining the type of nucleotide incorporated or introduced at a position of any nucleic acid molecule to be detected through the information obtained from one repeat reaction. For example, the four nucleotides are respectively labeled with fluorescent markers of four different emission bands for four-color or four-channel single-molecule sequencing or high-throughput sequencing; or, for example, the four nucleotides are respectively labeled with fluorescent markers of three different emission bands and a non-fluorescent label (cold nucleotide) for two-color or three-color high-throughput sequencing. One round of sequencing includes one repeat, and one round of sequencing can detect the type of base at a position of any nucleic acid template.
[0044] The so-called amplicon refers to the nucleic acid molecule to be tested after amplification. An amplicon is a cluster or strand or mass containing multiple identical polynucleotide sequences; the so-called fluorescent cluster, mass or sphere refers to an amplicon that can be excited to emit light, such as fluorescence, after a specified biochemical reaction. The amplicon can emit light and be imaged and detected during the sequencing process. By processing the signals collected by imaging and / or the corresponding image information, the bases incorporated or linked to the nucleic acid molecule to be tested after the specified biochemical reaction can be detected, thereby determining at least a part of the sequence of the nucleic acid molecule to be tested.
[0045] In the embodiments of the present application, unless otherwise specified, the so-called amplicon, fluorescent cluster or mass or sphere, nucleic acid molecule cluster, clone mass, nucleic acid molecule to be tested, template, nucleic acid template, template spot, bright spot corresponding to chemical features or fluorescent clusters on the surface, etc. can generally be used interchangeably. The so-called signal intensity, intensity, brightness, fluorescence brightness, etc. can generally also be used interchangeably. The so-called surface, chip, array, etc. can generally also be used interchangeably in the embodiments of the present application.
[0046] The machine learning model referred to in the embodiments of the present application, also simply referred to as the model, refers to a determination technique for predicting whether the type of the output base or base combination is correct or incorrect or the quality score based on known results (training data); in some embodiments, it also includes a determination technique for predicting the type of the output base or base combination. The so-called quality score characterizes the accuracy or error rate or reliability of the predicted result of the base or base combination. The known result can be a sequence or base considered or assumed to be accurate, that is, it is assumed that the sequence or base is correct. In some embodiments, the sequence or base assumed to be accurate is the result expected to be predicted by the model, and sometimes it is also called the corresponding ground truth base. The model can be supervised learning using training data.
[0047] In the embodiments of the present application, the concepts of labels or target values involved in machine learning are usually synonymous and refer to the results that the model needs to predict. The bases or base sequences involved in the labels or target values in machine learning are assumed to be accurate sequences (bases or base combinations), which are recognized as accurate sequences and are sometimes also called the corresponding ground truth bases. When measuring or generating training data, the corresponding bases in the labels therein may be inaccurate, but they are assumed to be accurate during model training. The assumed sequences can be measured in various ways. As described in the embodiments of the present application, the reference sequence, for example, the bases or base combinations corresponding to the positions where the sequencing results are aligned to the reference sequence, are used as the assumed sequences. The sequences measured or verified by detection methods or means generally considered to have higher accuracy, such as Sanger sequencing and / or DNA synthesis, are used as the assumed sequences. The consistent results measured by various detection methods or means are used as the assumed sequences. Or the sequences or base combinations with a high probability of being accurate, such as bases with a quality score reaching or exceeding Q40, or sequences with Q40 reaching or exceeding 80 or 85 or 90, are used as the assumed sequences, etc. Here, Q40 represents an error rate or error probability of one in ten thousand. If the training model outputting highly accurate or more reliable prediction results is not the desired goal, the measurement results with low accuracy can also be used as the assumed sequences and as training data. However, it can be understood that the accuracy of the prediction results of the model learned accordingly is generally not high either.
[0048] The embodiments of the present application relate to the learning of quality score algorithms or quality score models or related models, and based on the quality score model generated by training, the accuracy or reliability of the base determination results is scored to quantitatively evaluate or assess the accuracy or credibility of these base determination results. It can be used to compare the sequencing quality of different sequencing methods or sequencing platforms, and to determine whether the accuracy of the bases or sequences identified in a certain round or multiple rounds of sequencing or one or multiple sequencing runs meets the expectations and whether to continue this sequencing run, etc.
[0049] The base determination referred to in the related embodiments, also known as base identification, base class identification, base detection, base class detection, base reading, or base interpretation, etc., refers to the determination of bases or base combinations at designated or non-designated positions in a nucleic acid molecule. The base determination can be carried out independently or as part of a designated base such as A or T or a designated base combination such as ACT or GTCA, etc. The base determination can be for the same genomic position, for example, the scores or probabilities of detecting multiple bases at this position are close to each other, or for different positions. The score output from a machine learning model can be used to determine the base type at a designated or non-designated position. For example, scores can be provided or assigned to each identified base or base combination, and determining the base based on these scores can be part of the model prediction. In addition, some models can provide scores, and subsequent processes can use these scores. The score or value can reflect probability or likelihood. The probability scores of each base or base combination often sum to a fixed value, such as 1; while the scores reflecting likelihood generally do not need to sum to a fixed value, of course, they can also sum to a fixed value, for example, limiting the likelihood score of each to between 0 and 1, so that the likelihood scores can also sum to 1. In the embodiments of the present application, the values or scores reflecting the probability or likelihood size provided or assigned to each base or base combination are sometimes collectively referred to as base identification probabilities. In the related embodiments, the base class in the base identification result is often simply referred to as a base.
[0050] Please refer to Figure 1, some embodiments provide a system 100 for automatically generating a machine learning model. The machine learning model generated by the system 100 is used to output a quality score to evaluate the accuracy of base calling results. The system 100 includes: an input layer 120 for receiving user input from a user, the user input including sequencing results and optionally one or more tasks. The so-called sequencing results are obtained by sequencing nucleic acid molecules attached to multiple positions on a surface through surface imaging sequencing technology, and the sequencing results include base calling results and corresponding base calling probabilities; a machine learning layer 140 including a training data generation module 142 and a training module 144. The training data generation module 142 is configured to process the user input from the input layer 120 to generate training data. The so-called training data includes multiple pairs of associated features and labels. The so-called features describe the intensity corresponding to nucleic acid molecules at one or more positions and / or one or more base calling probabilities. The so-called labels include two categories: whether the base calling result is consistent or inconsistent with the corresponding ground truth base. The training module 144 is configured to train a specified model based on a portion of the training data to obtain a quality score model, including performing multiple iterative trainings on the specified model using a training set and evaluating the performance of the specified model after each iterative training using a validation set to obtain a quality score model. The training set and the validation set are each independently a non-overlapping portion of the training data; and an output layer 160 for outputting a quality score model that meets preset requirements. In related embodiments, the system 100 for automatically generating machine learning is also referred to as an automatic training tool.
[0051] In related technologies, regarding the quality score or quality fraction (Q value) of base calling results / base classification detections, regardless of what method is used for calculating the Q value output by each sequencing platform of each company, or what rules or custom rules are involved, the Q value can generally be regarded as another quantitative representation of the probability that the base detection result is incorrect or correct, and is an indicator characterizing the accuracy or reliability of the sequencing results. The higher the Q value, the higher the accuracy and reliability of the sequencing results. The Q value is usually calculated based on the Phred algorithm, and the formula can be written as Q = -10 * log 10(err), where err refers to the probability of base recognition error, that is, the error rate or error probability of the base detection result. For example, the common Q20, Q30, and Q40 respectively represent an error rate of one percent (i.e., a correct rate of 99%), an error rate of one-thousandth (i.e., a correct rate of 99.9%), and an error rate of one-ten-thousandth (a correct rate of 99.99%). The accuracy and reliability of the sequencing data can be intuitively understood through the Q value. In addition, in various practical applications of detecting by processing sequencing data or sequencing results, such as genetic variant detection in non-invasive prenatal screening, early screening of tumors in liquid biopsy, and pathogen detection, etc., it usually involves using the Q value to process sequencing data or setting quality standards to make the quality of the sequencing data relatively high or meet specific requirements to ensure the accurate and reliable application detection results, etc.
[0052] The embodiments of the present application do not limit the acquisition method of the sequencing results included in the user input received by the input layer 120. The so-called sequencing results are obtained by sequencing and detecting nucleic acid molecules connected to multiple positions on the surface through surface imaging sequencing technology, including base recognition results and corresponding base recognition probabilities. In some embodiments, the sequencing data is derived from a sequencing platform that realizes sequencing based on surface imaging detection. Most of the current mainstream high-throughput sequencing platforms, such as the sequencers of companies like Illumina, BGI Genomics, and Element Biosciences, etc., are platforms that realize the determination of the sequences of nucleic acid molecules on the surface based on surface imaging (such as surface fluorescence imaging) detection. By converting relevant biochemical reaction signals into electrical signals, such as image data, and then determining the base arrangement order (base recognition result) of the nucleic acid molecules to be detected based on the processing and analysis of the image data; in addition, the sequencing data publicly disclosed or provided by relevant manufacturers or user institutions of the sequencing platform also mostly originates from sequencing platforms based on this principle. The sequencing results can be one or more nucleic acid samples, the off-machine data obtained through one or more library constructions, one or more sequencing runs on one or more sequencing platforms, and optionally include some intermediate data, such as detected sequences, detected bases, the magnitudes of the probabilities of being recognized as various bases, and optionally related image data or other forms or formats of data, etc. The sequencing results can be obtained by the user's own sequencing, or provided by others, or downloadable or accessible from relevant data platforms, or can be the sequencing results formed by the mixture or combination of the sequencing data generated by oneself and others for one or more samples, one or more sequencing platforms for one or more runs.
[0053] In some embodiments, the sequencing results provided by the user input can be in the form of an access link or path, to reduce storage requirements and also facilitate the system 100 to receive and connect to read the data. The model output by the output layer 160 can also be in the form of an access path or link, for example, reserved or saved under a specified path to facilitate calling and running.
[0054] In some examples, the so-called user input or sequencing result further includes the alignment result of the sequence or base detected in the sequencing result to the reference sequence, for example, including the base or base combination corresponding to a specific position on the reference sequence that is aligned. The reference sequence is a sequence with a known sequence, which can be a reference genome, a part of the reference genome, or a known sequence formed by recombining and reorganizing one or more reference genomes. In some related examples, the base or base combination on the reference sequence corresponding to the alignment position serves as the corresponding ground truth base in the label of the training data.
[0055] In some embodiments, the sequencing result includes no less than 10 million sequence reads, and / or no less than 1 billion base pairs of base reads. In this way, sufficient training data can be generated so that it is possible to generate a model with better prediction effects based on the system 100.
[0056] The training data generation module 142 is used to process the user input from the input layer 120 to generate training data. The so-called training data includes multiple pairs of associated features and labels for supervised training of a specified model based on the training data.
[0057] In the process of determining and identifying the base or base sequence at a certain position of a nucleic acid molecule to be tested based on a method for sequencing implemented by surface imaging detection or a related sequencing platform, a base recognition algorithm or software or tool is involved. Various base recognition algorithms, software or tools usually involve comparing the possibility / probability of identifying the four bases A, T, C, and G at this position or determining the base recognition result based on this. For example, according to the analysis of the signal intensity and the like at the position of the corresponding biochemical feature on the image, the possibility / probability that this position may be identified as A is 10%, as G is 10%, as C is 10%, and as T is 70%. Then, based on this, the base at this position is very likely to be identified as T, provided that the relative possibility / probability of these four bases at this position is the final result (the signals corresponding to various bases have been corrected or will not be corrected and adjusted, etc.). Another example is that after processing the image information including crosstalk correction and / or phase correction of relevant chemical features, the possibility / probability that a certain position may be identified as A is 22%, as G is 28%, as C is 24%, and as T is 26%. Then, in view of the fact that there is no obvious difference or the difference is less than the preset level in identifying any base at this position, the base at this position is very likely to be identified as N or assigned as N, where N is selected from one of A, T, C, and G, or N refers to an uncertain one of ATCG.
[0058] In some embodiments, the so-called features describe one or more base recognition probabilities. For example, the probability values of all base categories, the probability value of the base category with the highest probability, the probability values of the base categories with the highest and the second highest probabilities, etc. The base categories referred to herein can be the base categories formed by individual bases A, T, C, G in each base recognition result, or the base categories formed by a combination of a single base and two or more bases such as AT, ACG, etc.
[0059] Specifically, for example, the feature can be the probability value of identifying each base category including base combination categories (probabilities of all types of base recognition), or the probability value of the base category with the maximum probability among them, the probability values of the base categories with the maximum and the second maximum probabilities among them, the probability values of the base categories ranked top three in descending order of numerical value among them, the probability values of the base categories ranked top four in descending order of numerical value among them, or any one or combination of the above or the processed or transformed probability value forms of any one or combination, etc. The probabilities or probability values referred to in related embodiments can be absolute values, relative values, ratios, deformed numerical values after being normalized or standardized, etc., or enlarged or reduced based on certain rules. In a specific example, the feature describes the probabilities of multiple base categories corresponding to each base recognition result.
[0060] In some other embodiments, the user input or the sequencing result input by the user further includes intermediate sequencing data or parameters such as sequencing images, including the processed sequencing images or part of the information of the sequencing images. The so-called features describe the intensities corresponding to nucleic acid molecules at one or more positions. Alternatively, the so-called features include both the intensities corresponding to nucleic acid molecules at one or more positions and the probabilities of at least one base category (base recognition probabilities).
[0061] Specifically, in some examples, the intermediate sequencing data includes images of nucleic acid molecules at multiple positions on one or more surfaces collected from one or more rounds of sequencing detections in one or more sequencing runs. The so-called images or sequencing images include the images directly collected during the sequencing process, the images collected after preprocessing of the sequencing, or the images reconstructed or recombined based on the images collected during the sequencing. In the embodiments of the present application, these images are sometimes also referred to as original images. Thus, the subsequent training data generation module 142 processes such image data, for example, extracts partial information in the images such as the positions and intensities of each nucleic acid molecule, etc., as feature values.
[0062] In some other examples, the intermediate sequencing data includes a portion of data determined by extracting from the sequencing images, including the positions and intensity values of multiple nucleic acid molecules in one or more surface regions where they are located. The so-called sequencing images are images of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of sequencing detections in one or more sequencing runs. Specifically, the user input includes the information obtained after image processing, including information representing or reflecting the target signal in the image, that is, the nucleic acid molecules or amplicons to be tested in which a specified biochemical reaction has occurred. For example, the position information and intensity values of the amplicons. The size of the input data is smaller than the size of the corresponding sequencing image. Here, the so-called size is also called the file size or dimension or data volume, and such sizes reflect the storage space required to store the data or the storage space occupied. Thus, compared with the case of inputting the sequencing image, such inputs can be directly used as feature values or directly as part of the training data, making the system 100 use this to generate a model have significant advantages in terms of computing power consumption and computing speed.
[0063] Specifically, the intermediate sequencing data or parameters include the position information and intensity information of the amplicons on the image, and the size of such data is smaller than the size of the corresponding image. Therefore, from a certain angle, it can be said that the input data is compressed image information or data obtained by extracting or screening the image information, which can be represented in a simplified image form or a non-image form. Preferably, it is represented in a data form convenient for a computer or a processor to calculate and process. In a preferred example, it is represented in a non-image data form, such as written in the form of a matrix, a multi-dimensional matrix, an array, a multi-dimensional array, etc., only recording or presenting the relative position information and intensity information of the amplicons on the image. Thus, compared with inputting the original image, the file size of the input data can be greatly reduced, the requirements for computing power and / or storage can be reduced, and it is also beneficial for the computer to quickly perform various operations or processes on it to quickly obtain the prediction result.
[0064] More specifically, in some examples, the size of the input data is less than or equal to half of the size of the corresponding image. In a certain example, especially for a regular surface, the file size of the input data is less than or equal to one-third of the file size of the corresponding image. In some tests, especially for a high-density regular surface, the size of the input data is even less than or equal to one-fourth or one-fifth of the file size of the corresponding image. Thus, compared with directly inputting image data, the storage or occupied space is significantly reduced, and the running speed is significantly improved.
[0065] Specifically, for example, the intermediate sequencing data is presented as multiple groups of matrices. One group of matrices includes multiple matrices. One group of matrices reflects the position information and intensity information of amplicons (nucleic acid molecules to be measured) on the images collected in one round of sequencing. One matrix contains multiple rows and multiple columns. One matrix reflects the information of amplicons on all or a region of the image of one fluorescence channel in one round of sequencing. One element of the matrix corresponds to one amplicon. The row and column positions where the element of the matrix is located reflect the position information of the corresponding amplicon on the corresponding image. The value of the element of the matrix reflects the intensity information of the corresponding amplicon presented on the image of one fluorescence channel in this round of sequencing. Sometimes this matrix is also referred to as the fluorescence brightness matrix. In this way, compared with the input image, the amount of information or the size of the input data is greatly reduced, which is beneficial to reducing the requirements for storage or computing power, and is also beneficial to improving the operation calculation speed. Moreover, such input data can be directly used as eigenvalues or part of eigenvalues, reducing subsequent operations.
[0066] For another example, one piece of intermediate sequencing data can correspond to multiple amplicons in a region or field of view (FOV). One piece of data, such as 1 or 1 group of "fluorescence brightness matrices", can correspond to the position and intensity information of the amplicons that have undergone a specified reaction in the images collected in one round of sequencing for one FOV. In this way, compared with the input original image, the size of the input data is significantly reduced, and the requirements for computing power and / or storage can be reduced.
[0067] Regarding the number of matrices included in one group of matrices and the matrix dimensions involved, it can be understood that one group of matrices corresponding to a group of images of one FOV for multiple fluorescence channels in one round of sequencing, such as four fluorescence brightness matrices of amplicons from four fluorescence channels in one round of sequencing for one FOV in four-color sequencing, such as three fluorescence brightness matrices of amplicons from three fluorescence channels in one round of sequencing for one FOV in three-color sequencing, such as two fluorescence brightness matrices of amplicons from two fluorescence channels in one round of sequencing for one FOV in two-color high-throughput sequencing, can all be referred to as a three-dimensional matrix. Combining with the common understanding, those skilled in the art can understand the specific meaning of the description involving matrices, the number of matrices or dimensions according to the specific embodiments including the context.
[0068] Furthermore, in some specific examples, the advantages of the input matrix compared to the original input image are not only in the differences in image and matrix sizes. For example, among the four images of a field of view from a regular surface collected in one round of sequencing for four-color fluorescence imaging sequencing, assuming that the size of each image is 2000*2000 pixels, while the size of the template point matrix (fluorescence brightness matrix) reflecting the positions and intensity information of the amplicons on the image, written in matrix form, is 800*800, then the size is reduced to 0.4^2 times the original. Moreover, the brightness information corresponding to the 800*800 template points input into this matrix can be the brightness information corresponding to sub-pixel level positioning positions, such as the brightness obtained through bilinear interpolation or other interpolation methods. Therefore, in terms of the amount of information, the information of the 800*800*4 template brightness matrix reflects the information of the four images of 2000*2000*4 plus the 800*800 position matrix. It can be said that the amount of information of the input matrix reflects the position information and intensity information of the amplicons determined directly or indirectly based on the image information, including the information of the processed image. In this way, the input data not only contains the target information, its size is significantly reduced, but also presents in the form of a matrix that is easy for machine operation and calculation, which is further beneficial for quickly training the model. Moreover, such input data can be directly used as eigenvalues, or as part of the features or training data to reduce subsequent operations.
[0069] In some other specific examples, the format of the intermediate sequencing data or sequencing results input by the user is also a matrix. One matrix reflects the information of the amplicons on a block of a fluorescence channel image in one round of sequencing, and the input data also includes a matrix reflecting the position information of each block in the image it comes from. In this way, the matrix corresponding to a fluorescence channel image is converted into a group of relatively small matrices, or in other words, a fluorescence channel image is converted into a group of blocks, and the information of the amplicons in the simplified blocks is extracted in matrix form. By processing these relatively small matrices in parallel, in the same hardware and software operation environment, the processing speed of the input data can be further improved, and the model can be trained more quickly.
[0070] Please refer to Figure 2 , in a specific example, an image is converted into a group of block matrices. After all block matrices of different sizes are zero-padded to the same size as the input, the entire image is thus converted into a total of 3*3, that is, 9 blocks; and the length and width of the blocks may be different. Therefore, after splitting into 9 blocks, the edges of all blocks with sizes less than b*y can be filled with 0 values to a size of b*y; finally, an additional position encoding layer is added according to the position numbers of the blocks in the whole image. For example, the block in the lower right corner is numbered 9, and a layer of Pos(P) position encoding is filled, with all values being 9, to obtain the final input matrix.
[0071] The so-called labels include two categories: those where the base recognition result is consistent with the corresponding ground truth base and those where they are inconsistent. Among them, as mentioned in the relevant embodiments, the base recognition result can be a base category formed by a single base such as A, T, C, and G or A, T, C, G, and N, or a base category formed by a combination of two or more bases such as AT, ACG, etc., or a base category formed by a combination of a single base and two or more bases.
[0072] The corresponding ground truth base is the base category determined to be correct. In some examples, the corresponding ground truth base category comes from the sequences of human and non-human samples that have been detected and fully verified to be consistent by various sequencing instruments, sequencing chemistries, and sequencing protocols. Thus, it can be considered as the correct or true base category, or the accurate and reliable base category confirmed by the development of relevant detection methods and means up to now. In other examples, the corresponding ground truth base is the base or base combination at the corresponding position of the reference sequence when the sequencing result is aligned to the reference sequence.
[0073] In one example, the user input includes the corresponding ground truth base. In other examples, the user input includes a reference sequence, such as providing an access or link address to the reference sequence, etc., so that the training data generation module 142 processes the user input to generate training data, including aligning the sequencing result to the reference sequence and using the base or base combination on the corresponding reference sequence as the corresponding ground truth base category.
[0074] In some embodiments, training the specified model in the training module 144 further includes obtaining the classification probability of the base recognition result in the training set; dividing the base recognition results into a groups based on the classification probability and / or the base category in the base recognition result, where a is a natural number not less than 5; for each group of classification probabilities and / or the corresponding base recognition results, determining the error rate or accuracy rate of the base category in the group of base recognition results by comparing the base category in the group of base recognition results with the corresponding ground truth base; and associating a predetermined quality score with the error rate or accuracy rate of the base category.
[0075] Specifically, the base recognition result can pass through the quality score model or algorithm of any of the above related embodiments to obtain the so-called classification probability P value, denoted as P0. P0 is a probability value with a numerical value between 0 and 1. Further, P0 can be divided into a segments, and the base categories in the base recognition result will also be respectively divided into these a segments or groups accordingly. a is a natural number not less than 5. In some specific examples, P0 is divided into a segments in an evenly divided manner. For example, if P0 is evenly divided into 5 segments, then the ranges of these 5 segments of P0 are [0, 0.2), [0.2, 0.4), [0.4, 0.6), [0.6, 0.8), and [0.8, 1.0] in sequence.
[0076] Further, for each group, compare the base category in the base recognition result of each group with the corresponding ground truth base to determine the error rate or correct rate of the base category in the base recognition result of this group, including identifying or marking the bases that are the same as the corresponding ground truth base in this group as correct, and those that are inconsistent as incorrect. In this way, the correct rate or error rate of the base category of this group can be calculated.
[0077] In related examples, the Q value corresponding to the logarithm of the true error rate of a group or a segment of bases is called the true Q value. For example, as mentioned in the above embodiment, the base categories in the base recognition result are divided into a segments according to the P0 value. Then, the total error rate of the bases in each segment can be calculated, and the corresponding Q value can be calculated based on this total error rate. This total error is the so-called true error rate, and the corresponding Q value is the so-called true Q value. Specifically, following the above example, find all the bases (base categories) with P0 in the range of [0, 0.2). Taking the total number of bases in this segment as the denominator and the number of bases in the base recognition result that are different or inconsistent with the reference answer (ground truth base) as the numerator, calculate the total error rate of the base category corresponding to this segment of P0 in the range of [0, 0.2), and take the logarithm of this total error rate to convert it into a Q value. In this way, the true Q value of the base category corresponding to this segment is determined. The true Q value can be further assigned to each base in this segment. In this way, the true Q value or true error rate of each base category in this group or this segment is determined.
[0078] In some other embodiments, training the specified model in the training module 144 further includes obtaining the classification probability of the base recognition result in the training set; dividing the base recognition result into a groups based on the classification probability and / or the base category in the base recognition result. For each group, by comparing the base category in this group with the corresponding ground truth base, determine the error rate of the base category in the base recognition result of this group. a is a natural number, and the value of a can achieve that the logarithmic difference of the error rates of the base categories in the adjacent groups of base recognition results is less than 2; and associating a predetermined quality score with the error rate of the base category.
[0079] In some preferred examples, the setting or adjustment of the size of a follows the following principle: after dividing P0 into a segments or a groups, the difference in the true quality values (true Q values) of the base categories corresponding to adjacent segments is not significant. Specifically, the difference in the true quality values of the base categories of adjacent segments is less than 2, which can be regarded as not significant; more specifically, for example, any value within the range of [0.5, 2) can be regarded as "not significant", such as the difference being less than or not greater than 0.5, 0.8, 1, 1.1, 1.2, 1.5, etc. More preferably, the difference in the true Q values of the base detections of adjacent segments is less than 1, which is regarded as not significant. If the preset limit is exceeded, P0 can be re-segmented or further segmented, that is, a larger value of a is taken. By following this principle for multiple tests, the value of a is usually much greater than 10. The inventor comprehensively tests and examines the quality requirements of sequencing data for various detection applications and the sequencing data obtained by sequencing multiple samples including human samples and non-human samples using sequencing platforms based on different sequencing principles or different commercial sequencing platforms, in order to establish such preferred principles or related setting or value-taking rules. In this way, it is beneficial to quickly and automatically generate a better model or a model that meets the user's input tasks. The system 100 can, for example, determine a more suitable a or an a that conforms to the principle through iterative operations.
[0080] Compared with the method of directly or first segmenting the classification probability P0 to directly determine the bases (base categories) in the base recognition results of the corresponding segments in the above related examples. In some other specific examples, the base categories in the base recognition results can be classified first. The bases can be classified according to the current (round) base recognition type, the first n (round) base recognition types, the last m (round) base recognition types, or any combination of any two or all three of them. m and n are each independently natural numbers. For example, classify according to the base categories such as A, C, G, and T recognized in the current (current round), or classify according to the base categories such as CC, AA, AT, CG, etc. recognized in the current (round) base and the base categories recognized in the previous round or the next round, or classify according to the base categories such as ATC, TAT, TTT, GCA, etc. recognized in the current (round) base and the previous and next (round) bases, or classify the base categories according to any combination of the above two or more, and so on.
[0081] In related examples where the base categories in the base recognition results are first classified, the classified base categories are often directly referred to as base categories or bases. If there may be ambiguity or confusion, sometimes the classified base categories are referred to as new base categories. The classification probability of the classified base categories can be the classification probability of the current (round) base category among them. For example, it is still P0 in the above related embodiments, or it can be the corrected P0. For example, the P0 is corrected or calibrated based on the classification probabilities or base recognition probabilities of non-current bases in the new base category, such as the previous base and / or the subsequent base. It can also correct the mapping relationship between this P0 and the Q value. Specifically, for example, the classification probabilities of non-current bases in this base classification are relatively low, such as all less than 0.8. To some extent, it can be regarded that the current base category is "dragged down" by one or some of the previous or subsequent and relatively unreliable bases, resulting in a decrease in the credibility of this base category. Its P0 can be adjusted, or the mapping relationship between P0 and the Q value can be adjusted. For example, multiply it by a coefficient less than 1, such as 0.9, to update P0, or subtract 1 or 2 from the Q value corresponding to this P0, or assign it to the adjacent lower-level segmented Q value to correct the classification probability of this base category or the mapping relationship between the classification probability of this base category and the Q value. Or, remove the bases with relatively low classification probability values (relatively unreliable bases) in this base category to update this base category, and replace this base category with the updated base category to enter the subsequent analysis process. Or, even directly remove this base category so that this base classification does not participate in the subsequent process. In this way, it helps to establish or generate a more accurate correspondence relationship.
[0082] Furthermore, the subsequent processing of the method of first classifying the base categories in the base recognition results can be similar to the methods or ways in the previous examples. For example, segment based on the classification probability P0 of the classified base categories, calculate the true error rate and / or true Q value of the bases in each segment, and count or clarify the number or distribution of the base categories of each true error rate and / or true Q value, etc., to establish the correspondence relationship between each classification probability and the predetermined quality score. The processes, methods, and technical effects of the above related examples are equally applicable to processing these new base categories and associated data, and will not be elaborated here.
[0083] The so-called predetermined quality score refers to the value-taking requirements for segmented predetermined Q values. For example, it can be a parameter that can be set and changed by the user input, such as from System 100, or a parameter that will be automatically selected or adjusted by System 100 according to the task input by the user. Currently, for the Q values calculated based on various Phred algorithms or the Q values output by various sequencing platforms, consecutive integers and non-consecutive integers are two common types. Specifically, the predetermined output Q value is a consecutive integer. For example, the value range of the Q value is all integers within [0, 40], all integers within [0, 45], or all integers within [0, 50], etc.; the predetermined output Q value is a non-consecutive integer, also known as the Q-score binning method, which divides the consecutive Q values into several intervals (bins). For example, the Q values form an arithmetic sequence or are composed of non-consecutive integers without an obvious rule. For example, they are 0, 10, 20, 30, 40, and 50, or they are composed of 0, 10, 17, 21, 28, 35, 40, and 50. Both consecutive Q values and non-consecutive Q values are widely used in currently commercially available sequencing platforms. Relatively speaking, non-consecutive Q values are more common. Non-consecutive Q values can simplify data processing, improve robustness, and facilitate visualization, significantly enhancing the efficiency and stability of data analysis. However, non-consecutive Q values may also lead to partial information loss. In specific applications, which way the Q value is suitable or expected to be presented generally requires weighing the pros and cons.
[0084] In a certain example, the predetermined quality score conforms to the true Q value in the above-related example. That is, the true Q value meets the value-taking requirements of the predetermined quality score. Thus, the established correspondence between the classification probabilities of each base category and the predetermined quality score is the establishment of the so-called quality score model.
[0085] In some other examples, the predetermined quality score is a set of score value-takings different from the true Q value. Usually, it is necessary to further update the correspondence or mapping relationship based on the established correspondence between the classification probabilities of the base or each group of bases and the true Q value, combined with the quantity ratio or distribution of the bases with each true Q value / true error rate, such as the number ratio or distribution, and the new Q value value-taking requirements, in order to establish or update the correspondence, that is, to establish or update the quality score model that meets the user input task or requirements.
[0086] Specifically, in some specific examples where the predetermined quality score is different from the true Q value and the value range of the true Q value includes the value range of the predetermined quality score, the so-called correlation between the predetermined quality score and the error rate of the base category includes, based on the error rate or accuracy rate of the base category in each grouped base recognition result, the proportion or distribution of the number of bases in each group of the base category, and the predetermined quality score requirement, merging the groups into new groups with fewer numbers of groups, adjusting the attribution of the base categories of adjacent new groups, and calculating the error rate of the base categories of each new group, so that the error rate of the base categories of the final new group corresponds to the predetermined quality score, in order to establish a mapping relationship between the classification probability of the base category and the predetermined quality score.
[0087] It can be understood that the so-called predetermined quality score, that is, the value requirement of the predetermined Q value, is equivalent to predetermining the number of new segments and the error rate of each new segment. By merging one or more adjacent segments into new segments with fewer numbers, and adjusting the base attribution of adjacent new segments (equivalent to adjusting the proportion or distribution of the number of bases in each segment) and calculating the error rate of the base categories of the new segment, finally making the error rate of the segment calculated based on P0 of the bases in each new segment be the predetermined error rate of the segment (the error rate of the new segment or the value of the new Q value). In this way, in order to establish a mapping relationship between the classification probability of each base category and the predetermined quality score, the so-called quality score model is established or generated.
[0088] Specifically, in an example, based on the statistics of the training data, including the true Q value (true error rate) of each base (base category) in each original group or segment, the distribution or distribution trend of the number of bases or the proportion of the number of bases in each group or segment, and this new quality score requirement (predetermined quality score), merging the bases of the original groups or segments and re-grouping or segmenting, and making the error rate corresponding to the bases of each segment after re-grouping or segmenting (the true error rate of the bases of the segment after re-segmenting) consistent with the predetermined quality score, in order to establish a corresponding relationship between the new groups of bases and the predetermined quality score, so as to update and generate the desired quality score model.
[0089] In system 100, the implementation of what is referred to herein as merging one or more adjacent segments into a smaller number of new segments, allocating the bases among the adjacent new segments, and calculating the error rate of the corresponding new segments can be achieved through iterative calculation. For example, if it is intended to divide segments with one integer per segment, such as consecutive integers from 0 to 50, into four new segments 10, 20, 30, and 40, the corresponding true Q for the new segment 10 may be 11 and the corresponding true Q value for the new segment 20 may be 19 during the first division. This indicates that the new segments are not reasonable enough or do not meet the predetermined requirements. Therefore, the bases in the adjacent new segments are reallocated again. For example, the lowest score (corresponding base) that was assigned to the new segment 20 in the previous step is reallocated to the new segment 10 in the previous step, so as to reduce the true Q value corresponding to the new segment 10 and increase the true Q value corresponding to the new segment 20. Iterative updates are performed in this way until the true Q value / true error rate corresponding to the final new segments is consistent with the predetermined quality score.
[0090] It can be understood that calculating or determining the correct rate or accuracy rate, as referred to in the relevant embodiments, has the same meaning as calculating or determining the error rate. The correct rate and the error rate are corresponding and can be converted into each other. In addition, the correct rate and the correct probability or accuracy rate or accurate probability in the relevant embodiments, unless otherwise specified, generally have the same meaning; the corresponding error rate and error probability and error occurrence probability, etc., are the same, and can be replaced unless otherwise specified.
[0091] In some other specific examples, it further includes merging and regrouping based on the error rate of the base categories in each group base recognition result and the quantity proportion or distribution of the base categories in each group, so that the error rate of the base categories in the new group corresponds to the first quality score, in order to establish a mapping relationship between the classification probability of each base category and the first quality score; based on the error rate of the base categories in the new group and the quantity proportion or distribution of the base categories in each new group, merging and regrouping the new group, so that the error rate of the base categories in the final group corresponds to the predetermined quality score, in order to establish a mapping relationship between the classification probability of each base category and the predetermined quality score, where the values in the first quality score include all the values of the predetermined quality score. In this way, two mapping relationships can be directly generated, and two corresponding quality score models or a cascaded quality score model can be output for the user to compare and select.
[0092] The first quality score is the specified quality score, which can be referred to as the standard quality score. In one example, the standard quality score comes from user input, such as a parameter that the user can set, adjust, or input by themselves. In another example, the standard quality score or the corresponding mapping relationship is an optional parameter built into system 100, and a mapping relationship can be established. It is a parameter or relationship that system 100 determines whether to select or learn according to the task input by the user.
[0093] In a specific example, the value of the first quality score is a continuous integer, such as a continuous integer Q value that adapts to or can cover the Q values of most current sequencing platforms, such as the continuous integers [0, 45] or [0, 50], while the predetermined quality score is a non - continuous integer, such as to meet or adapt to the tasks input by the user or the detection requirements of specified applications, etc. If the predetermined Q value is the same, and the upstream base recognition algorithm or the machine - learning base recognition model has changed or it is unknown whether the upstream base recognition algorithm or model is the same, then in most cases, only the first mapping relationship, that is, the mapping relationship between the classification probability of the previous base category and the standard quality score, needs to be updated and adjusted, without adjusting the subsequent mapping relationship. In this way, a quality score model can be quickly generated, facilitating the conversion and evaluation of the actual quality of sequencing data with the same or different Q values from different sequencing platforms, etc.
[0094] Moreover, it can also quickly and conveniently meet some special Q - value segmentation requirements. Specifically, the first mapping relationship (the previous mapping relationship) maps all bases to the standard segmented interval (standard quality score) in any scenario. Therefore, these example solutions can provide standard intermediate results during the performance verification of the Q - value algorithm and the troubleshooting of algorithm problems, which is helpful for performance verification and problem troubleshooting. For example, based on the performance of the previous mapping relationship, it can be evaluated whether the mapping relationship from P0 to the continuous integer Q value of [0, 50] is reasonable, and based on the performance of the subsequent mapping relationship, it can be evaluated whether the segmented mapping from the continuous integer Q value of [0, 50] to the non - continuous integer Q value is reasonable, etc.
[0095] In related examples, please refer to Figure 3 , the user input further includes a reference sequence, and the training data generation module 142 further includes an alignment sub - module 1422. The alignment sub - module 1422 is used to align the sequencing result to the reference sequence to generate an alignment result. Based on the alignment result from the alignment sub - module, the label associated with the feature is determined. For example, the corresponding base or base combination on the reference sequence is used as the corresponding ground - truth base to generate training data.
[0096] Furthermore, in a specific example, the alignment sub - module 1422 is connected to the sequencing system or sequencing platform that generates the sequencing result and is connected to the base recognition software provided or carried by the sequencing system or sequencing platform itself to align the sequencing result to the reference sequence to determine the corresponding base on the reference sequence.
[0097] In addition, during the process of constructing training data by the training data generation module 142, sequences with poor sequencing quality, unreliable alignment positions, or relatively low credibility can be excluded to improve the reliability of the generated training data. Training the model based on highly reliable training data is conducive to obtaining a more accurate and reliable prediction model. For example, in a certain example, for sequences that fail to successfully align with the reference sequence or whose alignment results are unreliable or have low credibility, such sequences or bases are re-assigned as "N" or marked as "N", and the bases or sequences marked as N are not used as training data or do not participate in the subsequent model training process, which can improve the reliability of the generated training data or the model trained based on such training data.
[0098] In some implementations, please refer Figure 4 , the machine learning layer 140 further includes an evaluation module 146. The evaluation module 146 is connected to the training data generation module 142, the training module 144, and the output layer 160, and is used to evaluate the performance of the quality score model using a test set before outputting the quality score model in any relevant example, so as to determine that the quality score models meet the preset requirements. Here, the test set is a part of the training data that does not overlap with the training set and the validation set, and meeting the preset requirements means that at least one of the accuracy rate, precision rate, recall rate, and speed of the classification prediction results of the quality score model for the test set meets the preset standard or can achieve the task input by the user.
[0099] In some embodiments, the specified model is selected from at least one of logistic regression, naive Bayes, support vector machine, and tree models such as lightGBM. The specified model can be a single model or a cascade model or composite model containing multiple single models.
[0100] In some embodiments, please refer Figure 5 , the machine learning layer 140 further includes a selection module 148 for determining the specified model based on the user input. It can be that the user directly selects a certain specified model through the user input, or the system 100 automatically selects the specified model according to the task input by the user.
[0101] In addition, in some related embodiments, the user input further includes sequencing data (intermediate sequencing data). The sequencing data includes images of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs; or the sequencing data includes the positions and intensity values of multiple nucleic acid molecules in one or more surface regions where they are located, and the sequencing image is an image of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs.
[0102] The sequencing data is processed using a base calling model to obtain the so-called sequencing results. The base calling model or algorithm can be the base calling software provided by the sequencing platform or a machine learning base calling model automatically generated by the system 100. In some embodiments, refer to Figure 6 A or Figure 6 B, the system 100 further includes a base calling model automatic generation system 110 connected to or placed in the input layer 120. The base calling model automatic generation system 110 is used to generate and output a base calling model.
[0103] According to some specific examples, the input layer 120 is the first input layer 120, the machine learning layer 140 is the first machine learning layer 140, and the output layer 160 is the first output layer 160. Refer to Figure 7 , the base calling model automatic generation system 110 includes: a second input layer 130 for receiving user input from the user including the so-called sequencing data and optionally one or more tasks included. The sequencing data is obtained by detecting nucleic acid molecules connected to multiple positions on the surface through surface imaging sequencing technology; a second machine learning layer 150 including a second training data generation module 152 and a second training module 154. The second training data generation module 152 is used to process the user input from the second input layer 130 to generate second training data. The so-called second training data includes multiple pairs of associated second features and second labels. The second features describe the intensities corresponding to nucleic acid molecules at one or more positions, and the second labels include one selected from the bases A, T, C, and G. The second training module 154 is used to train a second specified model based on at least a part of the second training data, including performing multiple iterative trainings on the second specified model using the second training set and evaluating the performance of the second specified model after each iterative training using the second validation set to obtain a base calling model. The second training set and the second validation set are each independently a non-overlapping part of the second training data; and a second output layer 170 for outputting the base calling model. The second output layer 170 is connected to the input layer 120.
[0104] The acquisition methods and technical characteristics of the sequencing data or intermediate sequencing data in the foregoing related embodiments are equally applicable to the base calling model automatic generation system 110 and will not be elaborated here.
[0105] In some examples, the so-called second labels further include combinations formed by any two, three, or four selected from the bases A, T, C, and G. The second labels are assumed to be the correct base calling results. The acquisition methods and technical feature characteristics of the corresponding reference true value bases in the foregoing related embodiments are equally applicable to the base calling model automatic generation system 110.
[0106] In some examples, refer to Figure 8 , the user input from the second input layer 130 further includes a reference sequence, and the second training data generation module 152 further includes a second alignment sub-module 1522. The second alignment sub-module 1522 is configured to align the sequence reads or base reads in the sequencing result to the reference sequence to generate a second alignment result; determine the second label associated with the second feature based on the second alignment result to generate second training data. The descriptions of the technical features and characteristics of the first alignment sub-module 1422 in the relevant embodiments are equally applicable to the second alignment sub-module 1522. In a specific example, the first alignment sub-module 1422 and the second alignment sub-module 1522 are the same module.
[0107] In some examples, refer to Figure 9 , the second machine learning layer 150 further includes a second evaluation module 156 connecting the second training data generation module 152, the second training module 154, and the second output layer 170, which is configured to evaluate the performance of the base recognition model using the second test set before outputting the base recognition model, so as to determine that the base recognition model meets the preset requirements. The second test set is a part of the second training data that does not overlap with the second training set and the second validation set. The so-called meeting the preset requirements means that at least one of the accuracy rate, precision rate, recall rate, and speed of the classification prediction result of the base recognition model for the second test set meets the preset standard or that the task input by the user can be achieved.
[0108] In some examples, the second specified model is selected from at least one of logistic regression, support vector machine, tree model, and neural network. The second specified model is a single model or a composite or cascaded model including multiple models. In a specific example, the second specified model includes a neural network model, such as a neural network model for semantic segmentation such as a convolutional neural network (CNN). More specifically, it is at least one of U-Net, DeepLab, SegFormer, PSPNet, FCN, GCN, and their respective variants. In a specific example, the base recognition model automatic generation system 110 trains the second specified model using the solution disclosed in CN119479831A, which is hereby incorporated by reference in its entirety.
[0109] In some examples, refer to Figure 10, the second machine learning layer 150 further includes a second selection module 158 for determining the type of the second specified model based on the corresponding user input. The user may directly select or specify a certain model through the user input, or the task system 100 or the base recognition model automatic generation system 110 may automatically select the second specified model for the user according to the user input. The descriptions of the technical features and characteristics of the selection module 148 in the relevant examples are equally applicable to the second selection module 158. In a specific example, the selection module 148 (the first selection module) and the second selection module 158 are the same module.
[0110] Some embodiments of the present application further provide a method for developing a quality score model, including using the system 100 for automatically generating a machine learning model in any of the above embodiments to access user input, and using at least one hardware processor to run the system 100 to present the best quality score model to the user. The user input includes sequencing results and optionally one or more tasks. The sequencing results are obtained by detecting nucleic acid molecules connected to multiple positions on the surface through surface imaging sequencing technology, and the sequencing results include base recognition results and base recognition probabilities.
[0111] It can be understood that the descriptions of the technical features, additional technical features, and related technical effects of the system 100 for automatically generating a machine learning model in the above related embodiments are equally applicable to the method for developing a quality score model or the method for developing a quality score model and an associated base recognition model in this embodiment, and will not be repeated here.
[0112] For example, in some embodiments, the sequencing results include at least 10 million sequence reads and / or at least 1 billion base reads.
[0113] In some embodiments, the user input is received through a model generation framework or interface; and / or, the best quality score model or a quality score model that meets the task requirements of the user input is presented to the user through the model generation framework or interface.
[0114] In some embodiments, the system is run using two or more operating systems, and at least one of file sharing, remote access, network communication, virtual machines, containers, and virtualization technologies or tools is used to implement the interaction between different operating systems. So as to quickly train and generate relevant models.
[0115] Specifically, in some examples, Windows and Linux are respectively used to run a part of the system 100, and at least one of WSL and / or SSH is used to implement the interaction between different operating systems.
[0116] SSH (Secure Shell Protocol), belonging to one of the "remote access" methods (for operating or accessing another operating system on one operating system), is applicable to the remote management of Linux and macOS and can be accessed from Windows through tools (such as PuTTY, Termius). WSL (Windows Subsystem for Linux), belonging to one of the "virtual file system and bridging tools" (bridging files or processes of different operating systems through virtualization technology), can implement running a Linux environment on Windows to achieve file and tool sharing.
[0117] In some embodiments, the operating system 100 is run using multi-threaded parallel computing. In this way, the speed of the generation model is improved. Specifically, in one example, the system 100, such as the training data generation modules 142 and / or 152 in the machine learning layer 140 and / or 150 thereof, interactively run on multiple operating systems, and a multi-threaded parallel mechanism is set on any one of the operating systems.
[0118] Some embodiments of the present application also provide a quality score model generated by using the system 100 for automatically generating a machine learning model in any of the above embodiments.
[0119] Some other embodiments of the present application also provide a computer-readable storage medium storing one or more instructions for controlling a processor to execute the method for developing a quality score model in any of the above embodiments. The computer-readable storage medium includes but is not limited to random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium well-known in the technical field.
[0120] An ordered list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. In related embodiments, the so-called computer-readable storage medium can be any device that can store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device or in combination with these instruction execution systems, apparatuses, or devices. More specific examples (non-exhaustive list) of computer-readable storage media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or other suitable media on which the aforementioned program can be printed, because the aforementioned program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then storing it in a computer memory. The various computer-readable storage media described in the related examples can represent one or more devices and / or other machine-readable storage media for storing information. The so-called "machine-readable storage medium" here can include, but is not limited to, wireless channels and various other media that can store, contain, and / or carry instructions and / or data.
[0121] Embodiments of the present application also provide a system for automatically generating a machine learning model, which is also referred to as a computer program product or a computing device. The system includes a memory storing a program and one or more processors coupled to the memory. The processors run the program to implement the method for developing a quality score model in any related embodiment.
[0122] When using a software implementation method, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more instructions or programs. When these computer programs or instructions are loaded and executed on a computer, the processes or functions in related embodiments of the present application are implemented in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0123] The so-called computer program product or computing device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The computing device may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components, the connections and associations between the components, and the functions of the components shown in the embodiments of this application are only examples and are not intended to limit the description of the embodiments of this application and / or related implementation manners.
[0124] Please refer Figure 11 , the system 100 or the electronic device 500 for automatically generating a machine learning model includes a computing unit 501, which can perform various appropriate actions and processes according to the computer program stored in the ROM (Read-Only Memory) 502 or the computer program loaded from the storage unit 508 into the RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The I / O (Input / Output) interface 505 is also connected to the bus 504.
[0125] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0126] The processor or computing unit 501 can be various general and / or special processing components with processing and computing capabilities. Some examples of computing unit 501 include, but are not limited to, CPU (Central Processing Unit), GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 executes the modules or methods or processes described in the relevant embodiments, such as implementing the functions of the machine learning layer 140 and / or 150, etc.
[0127] For example, in some examples, the method for developing a quality score model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method in the above-mentioned related embodiments may be performed. In other embodiments, the computing unit 501 may be configured in any other appropriate manner (e.g., by means of firmware) to perform the method for developing a quality score model and / or a base recognition model using an automatic training tool in any of the above-mentioned related embodiments.
[0128] It should be understood that the various parts of the system or method in the embodiments of the present application can be implemented with hardware, software, firmware or a combination thereof. In the above-mentioned related implementation manners or embodiments, multiple layers, modules, submodules or steps or methods can be implemented with software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented with hardware, it can be implemented with any one or combination of the following related prior arts in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0129] A person of ordinary skill in the art can understand that the system for automatically generating a model in any of the above-mentioned embodiments or all or part of the steps of the method for implementing any of the above-mentioned related embodiments can be completed by instructing the relevant hardware through a program, and the so-called program can be stored in a computer-readable storage medium, which, when executed, includes one of the steps or a combination of steps of implementing the example method.
[0130] In addition, each functional unit in each embodiment or example can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0131] Public solutions for establishing a Q-value model based on machine learning are, for example, CN118471340A, CN117912550A, CN117976042A, and CN118116469A, etc. Public solutions for constructing a base recognition model based on machine learning are, for example, CN118429967A, CN118429965A, and CN118969086A, etc. The system 100 for automatically generating a machine learning model of the present application is applicable to the automatic development of any publicly disclosed Q-value algorithm and / or base recognition model based on machine learning, can automate the development process of the machine learning model, and reduce the dependence on professional knowledge and experience backgrounds such as algorithm development.
[0132] The following describes the establishment, deployment, and use of a system (automatic development tool) for automatically generating a machine learning model by combining one or more publicly disclosed solutions for developing a Q-value algorithm and / or a base recognition model based on machine learning or related optimization and improvement solutions. The following example starts with automatically generating a base recognition model, and constructs training data based on the base recognition results predicted by the generated base recognition model and related parameters (such as base recognition probability) to automatically develop a corresponding quality score model. Based on the foregoing related embodiments, it can be understood that the present application embodiments do not limit the manner in which the sequencing results / base recognition results on which the automatic development of the quality score model depends are obtained, and they may not be derived from the prediction results of the base recognition model generated by this automatic training tool.
[0133] The following further exemplarily describes the specific operation modes or implementation processes that some of the above embodiments or examples may include. However, this should not be construed as limiting the scope of the subject matter of the present application to the following further exemplary descriptions or embodiments.
[0134] Establishment and evaluation of a machine learning-based base recognition model and corresponding Q-value algorithm (1) Establishment of a machine learning base recognition model
[0135] It mainly includes the following steps:
[0136] 1. Determine the usage scenario of the base recognition model to be built, and determine the corresponding platform, instrument, biochemical, fluid, etc. versions. Here, take the multi-fluorescence channel based on surface imaging, such as the so-called dual-color or four-color fluorescence high-throughput sequencing platform, as an example.
[0137] 2. Under the determined operating environment and conditions of 1, sequence the standard library sample of the known reference genome, and save the original sequencing data, such as the original fluorescence image (sequencing image).
[0138] 3. Establishment of base recognition model training data (hereinafter referred to as training data I):
[0139] (1) Use the existing basic version of base recognition software, such as the base recognition software (which may or may not involve machine learning) built into each sequencing platform, to analyze and recognize bases in the sequencing data, and output the base recognition result file (fastq file). When using the basic version of base recognition software for base recognition, retain the cluster localization information during the base recognition process, as well as the fluorescence brightness information (intensity information) of each fluorescence channel for each round of each cluster (amplicon), for subsequent extraction of model eigenvalue.
[0140] (2) Align the fastq file with the reference genome to obtain the alignment result file (sam or bam file).
[0141] (3) Based on the brightness information of each round of each cluster, optionally including the brightness information of the first N and last M rounds (N and M are independent integers respectively), and optionally other information such as the brightest base in the first N and last M rounds, the position information of the cluster, etc., make eigenvalues.
[0142] (4) Make target values based on the information after sequence alignment of the clusters. For example, the target value is the reference base corresponding to each base after alignment, sometimes also called the standard answer.
[0143] (5) One-to-one correspondence between the eigenvalues and the target values to complete the production of training data I. When making training data I, only the sequences with relatively high alignment confidence can be retained. In addition, preferably, polymorphic sites are not included in the scope of training data I. In short, when screening training data, ensure that the "target value" is accurate as much as possible, which is beneficial to training a model with more accurate prediction results. In an example, for data with a relatively high error probability of the target value, either it is not included in training data I, or the target value is marked as N, where N represents one of A, G, C, and T; in a scenario where model training is carried out using deep learning such as neural networks, the eigenvalue and target value corresponding to all clusters are written into training data I.
[0144] 4. Take a part of the training data I to train the base recognition model. For example, in this example, the training data I is divided into three parts. One part is used for the development of the base recognition model (this part of the training data I is briefly denoted as I.1), one part is used for the development of the corresponding Q-value model (this part of the training data I is briefly denoted as I.2), and one part is used for verification (this part of the training data I is briefly denoted as I.3).
[0145] 5. Establish a base recognition model: Randomly divide the training data I.1 into a training set and a test set, and establish a base recognition model based on the data set (the base recognition model is briefly denoted as model B0). The so-called "base recognition model" is the corresponding relationship between the "feature value" and the "target value" mentioned above. When establishing the model, any machine learning or deep learning algorithm can be used. In this example, the model algorithm disclosed in CN119479831A is used, and its full text is incorporated herein by reference.
[0146] In paired-end (PE) sequencing, in some preferred examples, there are often some differences in the biochemical characteristics between R1 (read1) and R2 (read2). Therefore, it is preferably to extract features and train the model for them separately. Similarly, in some scenarios involving double-sided chip sequencing, the biochemical and optical characteristics of the two surfaces may be different. Therefore, it is preferably to establish models independently based on the detection or reading of each surface, which generally helps to obtain a model with better prediction effects. Here, taking the establishment of one of the models as an example, the other model is similar and will not be repeated hereinafter. (2) Establishment of the Q-value algorithm corresponding to the base recognition model B0
[0147] 1. Establish Q-value algorithm training data based on the training data I.2 (hereinafter referred to as the Q-value algorithm training data set as training data II)
[0148] (1) Use the above-established base recognition model B0 to complete base recognition based on the feature values of the training data I.2, and obtain the base recognition probability corresponding to the model, that is, the normalized recognition probabilities of A, G, C, and T output after the model prediction for each piece of feature data. The sum of the four base recognition probabilities of each piece of data is 1; and the base recognition result predicted by the model, usually the base corresponding to the maximum base recognition probability in each piece of data.
[0149] (2) Compare the above model base recognition results with the target values of the training data I.2. Those that are consistent are marked as 1 or correct, and those that are inconsistent are marked as 0 or wrong. This correct / error information is used as the target value of the training data II.
[0150] (3) The eigenvalue of training data II can be exactly the same as that of training data I.2; it can also be the base recognition probability corresponding to each piece of data; it can also be one or several of the four base recognition probabilities. For example, the eigenvalue is the maximum base recognition probability, or the maximum and the second maximum base recognition probabilities, etc.; it can also be a derivative calculation result based on the four base recognition probabilities, such as the ratio of the maximum and the second maximum base recognition probabilities; it can also be any arbitrary combination of any part of the above types of parameters.
[0151] 2. Establish a Q-value algorithm based on training data II (hereinafter referred to as the basic scheme)
[0152] a) Perform machine learning classification algorithm training based on training data II. Any suitable machine learning model can be selected, such as logistic regression, support vector machine, tree model, etc. The obtained model is denoted as model Q0.
[0153] b) Apply the Q0 model to training data II, that is, use the Q0 model to obtain the model prediction result based on the eigenvalue of training data II. Here, the final model classification result may not be output, and only the classification probability P value given during model prediction may be output, denoted as P0.
[0154] c) Usually, P0 is a probability value with a numerical value between 0 and 1. Divide P0 evenly into a segments, and the bases in training data II are correspondingly divided into a segments respectively. The selection principle of a is that after evenly dividing P0 into a segments, the difference in the true Q values of the bases corresponding to adjacent segments is not large. In the actual operation of the example, it is generally required that the difference in the true Q values of adjacent segments is less than 1.
[0155] Here, some explanations are given for the "true Q value". In fact, the Q value is another representation of the error rate. The relationship between Q and the error rate can be written as Q = -10 * log10(err). Here, Q refers to the sequencing quality score of the base, that is, the Q value; err refers to the error probability or error rate of this base. The "true Q value" refers to the Q value calculated based on the true error rate (err) of a group of bases. For example, as mentioned above, the bases are divided into a segments according to the P0 value, then the total error rate of the bases in each segment can be calculated. For example, calculate the error rate of this segment with the number of inconsistent bases as the numerator and the total number of bases in this segment as the denominator, and calculate the corresponding Q value according to this error rate. This Q value is the "true Q value" defined here.
[0156] d) Write out the true Q values corresponding to the a segments of bases respectively.
[0157] e) Clarify the segmentation requirements of the Q value according to the algorithm development requirements. For example, in some scenarios, it is required that the value of Q takes all integers within the range of [0, 45], and in some scenarios, it is required that the value of Q takes 5 specific values within [0, 45], etc.
[0158] f) According to the above requirements for the value of Q, complete the segmentation of the Q value. In the above steps, based on the statistics of the training data II, the true Q value (true error rate) corresponding to each base in each of the a segments is known; at the same time, the distribution trend of the proportion of the number of bases in different segments is also known. Therefore, based on the error rate and the proportion of the number of bases corresponding to each of these a segments, as well as the error rate requirements in the target Q value segmentation (i.e., the value of Q mentioned above), the bases in the above a segments are merged and segmented again. After the merging and segmentation, the total error rate corresponding to the bases in each segment is consistent with the pre - required Q value.
[0159] g) Record the corresponding relationship between each final P0 value and the Q value (segmented Q value), and this relationship is the Q - value algorithm corresponding to the machine - learning algorithm.
[0160] 3. Establish a Q - value algorithm based on the training data II (hereinafter referred to as Supplementary Scheme 1)
[0161] a) The first few steps are the same as a) and b) of the basic scheme.
[0162] b) Corresponding to step c) of the basic scheme, when segmenting the bases according to the P0 value, consider the categories of base recognition to distinguish the bases. For example, in the basic scheme, all bases are grouped together and divided into a segments according to the P0 value. In Supplementary Scheme 1, the bases are first classified according to the final recognition type. For example, they can be classified according to the categories A, C, G, and T of the current base recognition; or they can be classified according to the categories of the current base and the front and back bases, such as ATC, TAT..., etc. In short, the bases can be classified according to the current base recognition type, the recognition types of the first n bases, the recognition types of the last m bases, or the pairwise combination or the combination of the three of the above, where n and m are independently integers.
[0163] c) After completing the above classification, perform operations on each type of classified base according to d) - g) of the basic scheme.
[0164] d) Record the corresponding relationship between various P0 values and Q values (segmented Q values) corresponding to each final base classification. This corresponding relationship is the so - called corresponding Q - value algorithm.
[0165] Supplementary Scheme 1, compared with the basic scheme, takes into account the potential P0 - value preference problem that may exist between different base combinations and corrects it when mapping the Q value. The preference of the P0 value for base combinations mainly comes from the preference characteristics in biochemistry, that is, the error rates corresponding to bases with the same brightness signal characteristics but different base combinations may vary greatly.
[0166] 4. Establish the Q-value algorithm based on Training Data II (hereinafter referred to as Supplementary Solution 2)
[0167] Either the basic solution or Supplementary Solution 1 can be superimposed with Supplementary Solution 2. One of the features of Supplementary Solution 2 is that the segmentation of Q-values is divided into two steps.
[0168] Specifically, in steps e)-f) of the basic solution or the corresponding steps of Supplementary Solution 1, the bases segmented according to the P0 value are directly mapped to the final Q-value segmentation. In Supplementary Solution 2, all bases are first mapped to the Q-value segmentation of all integer values in the range [0, 50].
[0169] Specifically, in the example of the basic solution, the bases segmented according to the P0 value are directly mapped to the Q-value segmentation of all integer values in the range [0, 50] for example. In Supplementary Solution 1, the bases in different groups are first mapped to the Q-value segmentation of all integer values in the range [0, 50], and then all base categories are combined according to the Q-value taking or segmentation requirements. Supplementary Solution 2, in short, is to obtain the mapping relationship between the P0 values corresponding to different categories of bases and the Q-value segmentation of all integer values in the range [0, 50] for example, which is hereinafter referred to as Mapping Relationship 2.1.
[0170] After obtaining Mapping Relationship 2.1, it is then mapped from the Q-value segmentation of all integer values in the range [0, 50] to the final required Q-value segmentation, which is hereinafter referred to as Mapping Relationship 2.2. Mapping Relationship 2.1 and Mapping Relationship 2.2 together constitute the corresponding Q-value algorithm.
[0171] One of the advantages of Supplementary Solution 2 is that Mapping Relationship 2.1 is implemented with the same logic in any scenario, which is conducive to the implementation of automatic modeling software. If there are different Q-value segmentation requirements in different scenarios, generally, only Mapping Relationship 2.2 needs to be adjusted. If the Q-value segmentation requirements are the same, and the upstream machine learning base recognition model has changed, then in most cases, only Mapping Relationship 2.1 needs to be adjusted, and usually Mapping Relationship 2.2 does not need to be adjusted.
[0172] Since Mapping Relationship 2.1 maps all bases to the standard segmentation interval in any scenario, Solution 2 can provide standard intermediate results during the performance verification of the Q-value algorithm and the troubleshooting of algorithm problems, which is helpful for performance verification and problem troubleshooting. For example, based on the performance of Mapping Relationship 2.1, it can be evaluated whether the mapping relationship from P0 to the integer Q-value in the range [0, 50] is reasonable, and based on the performance of Mapping Relationship 2.2, it can be evaluated whether the segmentation mapping from the integer Q-value in the range [0, 50] to the large-segmentation Q-value is reasonable.
[0173] The advantage of Supplementary Scheme 2 also lies in its ability to conveniently meet some special Q-value segmentation requirements. The bases corresponding to the subdivided Q-value segments (for example, the bases corresponding to integer Q-values in the range [0, 50], and also for example, the mapping relationship between the bases directly segmented according to P0 in the basic scheme and the final target Q-value segmentation is often not unique).
[0174] First of all, for example, during the product development stage, the final target Q-value is often given in the form of an interval. For example, it is required to "determine a value as the Q-value within the intervals (0, 15), (15, 20), (20, 30), (30, 40), and (40, 50) respectively", and the specific values within each interval are often not pre-defined. It can be determined according to the specific base classification situation during the development process, so there can be many kinds of Q-value segmentations that meet the requirements. Secondly, even if the specific value of the final Q-value is clear, there is more than one way to divide the bases that meet the requirements. In this case, Mapping Relationship 2.1 of Supplementary Scheme 2 maps the bases to the standard Q-value segments, which is more conducive to Mapping Relationship 2.2 to meet special Q-value segmentation requirements, such as base division requirements like "the Q35 segment cannot contain bases with a Q-value lower than 30".
[0175] (III) Model Performance Evaluation
[0176] a) Use the currently established base recognition model B0 and the corresponding Q-value algorithm to conduct tests on the validation set I.3.
[0177] b) Evaluation of the base recognition model B0
[0178] Based on the current round of brightness information in the I.3 eigenvalue, the base recognition result under the traditional algorithm (not based on machine learning model prediction) can be obtained based on the brightest base, and this result is recorded as the base recognition result 1 of I.3. Based on the eigenvalue of I.3, use the base recognition model B0 for prediction, and the obtained result is recorded as the base recognition result 2. Compare the base recognition result 1 and the base recognition result 2 with the target value of I.3 respectively, and calculate the error rate respectively (that is, the number of bases where the base recognition result is different from the target value divided by the total number of entries in the I.3 data). The reduction in the error rate of the base recognition result 2 compared to the base recognition result 1 is the improvement in the base recognition accuracy brought by the base recognition model B0.
[0179] c) Performance evaluation of the Q-value algorithm
[0180] Compare the base recognition result 2 with the target value of I.3 to obtain the correct / incorrect result of each data entry under the base recognition model B0, which is recorded as the base correct / incorrect result 2.
[0181] For the basic solution, based on the base correctness result 2, the error rate corresponding to the bases in each Q-value segment is output, and the actual Q-value is converted based on the error rate. The Q-value algorithm is judged whether it meets the expectation by comparing the difference between the Q-value assigned by the algorithm and the actual Q-value.
[0182] For supplementary solution 1, in addition to comparing the true Q-values corresponding to the bases in different Q-value segments, it is also necessary to compare the true Q-values corresponding to the bases in different Q-value segments corresponding to different base recognition types, so as to evaluate whether the total base Q-value assignment is reasonable and whether the Q-value assignment of specific types of bases is reasonable / meets the expectation.
[0183] For supplementary solution 2, in addition to comparing the true Q-values corresponding to the bases in the final Q-value segment, it is also necessary to compare the true Q-values calculated for the bases corresponding to the integer Q-values in the range of [0, 50] generated by mapping relationship 2.1.
[0184] In addition, before developing the base recognition model and the Q-value algorithm, it generally involves presetting the passing criteria and qualified thresholds for the performance of the base recognition model and the Q-value algorithm. For example, at least one of the accuracy, precision, recall, and speed of the classification prediction results of the base recognition model for the validation set I.3 meets the preset standard or can achieve the tasks input by the user, etc.
[0185] When the performance of the base recognition model or the Q-value algorithm does not meet the threshold requirements, the mechanism for redeveloping the algorithm can be triggered, or an alarm for manual intervention can be triggered. For example, the user can be reminded to increase the user input, such as increasing the sequencing data, in order to generate more training data for better model training.
[0186] Finally, the base recognition model and / or the Q-value algorithm that passes or meets the preset indicators or expected requirements after the model performance evaluation are output to complete the machine learning modeling.
[0187] Automation of the model development process
[0188] 1. For the above process of the clear data processing method and model development logic, it can be implemented in any programming language on a Linux system or a Windows system.
[0189] 2. When completing the main process of machine learning model training using the Windows system, system interaction requirements between Windows and Linux may be involved. For example, since most of the current mainstream bioinformatics analysis software is developed on the Linux system, even if there is a Windows version, its running speed is often much slower than the version under the Linux system. In the process of building a base recognition model in the above example, bioinformatics analysis steps such as alignment and error rate statistics are included. If the main process of model development is carried out under Windows, then completing this part of the bioinformatics process under Linux will help improve the efficiency of the automatic training tool.
[0190] Here, the interaction between Windows and Linux systems is achieved through the following two methods respectively.
[0191] (1) WSL2
[0192] Under the Windows system, install the Linux subsystem through the WSL mechanism, and the interaction between Windows and Linux is achieved using WSL. Advantages: The Linux subsystem starts quickly, occupies few resources, and has almost no impact on the performance of the Windows system; WSL2 has a complete Linux kernel and supports the installation of the vast majority of Linux software; WSL2 supports container technologies such as docker / singularity, and software deployment is convenient and fast. Disadvantages: The file mutual access between Windows and Linux is output through the network, and the read and write speed is slower than that of the native system.
[0193] (2) OpenSSH
[0194] Implementation method: Install OpenSSH software on both Windows and Linux, and interact through the SSH protocol. Advantages: Based on the SSH protocol, the security is greatly improved and the connection is stable; it is convenient for integration. By configuring SSH passwordless access, operations such as remote connection, task submission, and task status determination can be performed in a Python script; independent Windows and Linux nodes can make the most of the computing power of each node. Disadvantages: The installation and deployment are relatively complex, and it is necessary to deploy OpenSSH software on both Windows and Linux and configure passwordless access.
[0195] 3. In the model training stage, the running speed of the training tool can be improved through parallel computing.
[0196] For example, during the process of base recognition using the aforementioned (I) 3(1) utilization basis or traditional version of non-machine learning-based base recognition software, multiple threads can be called through code for simultaneous analysis. When developing an automatic training tool, it is generally necessary to pre-analyze the computing power and memory occupied by the basic version of the base recognition software and write the calculation formula for the maximum number of threads corresponding to different computing powers and memories. When using the automatic modeling tool (automatic training tool), the tool will automatically read the computing power and memory of the system used for modeling and automatically calculate the maximum number of threads that can run in parallel according to the maximum number of threads calculation formula.
[0197] Regarding calculating the maximum number of threads or the optimal number of threads corresponding to specific computing power and memory, those skilled in the art can reasonably calculate and determine by comprehensively considering hardware specifications (such as the number of CPU cores, hardware characteristics of the GPU such as the number of multiprocessors, thread block size, etc., memory capacity and bandwidth) and task types (such as CPU-intensive or IO-intensive). For example, calculate based on a processor such as a CPU and / or GPU, based on memory and / or based on task type, or comprehensively consider thread number calculation in order to reasonably calculate and optimize the number of threads.
[0198] (1) Whether the aforementioned (I) 3(2) is carried out on a Windows system or a Linux system, a multi-thread parallel mechanism can be set. Similarly, the calculation formula for the optimal number of parallel threads is pre-input, and the number of parallel threads is automatically set during the actual operation.
[0199] (2) The production of training data in the aforementioned (I) 3(3)-(5) can also set multi-thread parallel calculation according to the above idea.
[0200] (3) The machine learning training step of the aforementioned (I) 5 base recognition model can also set parallel operation. In addition, in the scenario where model training is performed separately on R1 and R2 data based on paired-end (PE) sequencing, the models of R1 and R2 can both be trained in parallel if the computing power and memory permit. Similarly, in some scenarios where model training is performed separately on data from multiple sides, such as double-sided chip sequencing, parallel calculation is set for each model training.
[0201] 4. The form of the feature values of the model can be determined before training to improve the data extraction speed. And when performing base recognition using the traditional or basic version algorithm in step (I) 3(1), the cached data during the base recognition process can be directly used to calculate and output the feature values simultaneously, without retaining other intermediate parameters that are not needed, thereby improving the efficiency of data calculation and reducing the time for data writing.
[0202] Using the automatic training tool
[0203] The following example shows how an automatic training tool outputs a model file based on user input such as sequencing images in a specific scenario.
[0204] Suppose a user needs to use a machine learning model in a certain special application scenario. This scenario may require special sample features, special biochemical and instrument configurations, etc., and no suitable machine learning model has been established in this scenario. Under such a premise, the user can use the system for automatically generating models (automatic training process) by themselves to build the corresponding or suitable model.
[0205] 1. Generally, the user first clarifies the sequencing conditions corresponding to the model to be established, such as biochemistry, instrument configuration, and the characteristics of the library sample, etc. Then the user performs sequencing under the above clarified sequencing conditions to obtain sequencing data for generating training data. When sequencing, for example, a library sample with a known reference sequence is used. If the user does not specifically specify a special library sample in this sequencing scenario, a library sample with a larger genome, such as a human genome sample, etc., can be used as the sequencing sample. Samples with a larger genome can make the training data contain more diverse sequence feature information, which usually helps to improve the generalization of the model. If the "special sequencing scenario" required by the user includes special library features, such as a library with extremely unbalanced bases, etc., then the sequencing library here is preferably a library with the same or similar features as the target sequencing library to help the machine learning model learn the sequencing preferences introduced by the library features.
[0206] 2. When the user performs the above sequencing, the machine learning automatic training tool is called. In current mainstream commercially available sequencing platforms that implement sequencing based on surface presentation or common sequencing processes, the sequencer often processes data including base recognition while sequencing. The software in the sequencer for implementing this data processing is the so-called traditional or basic version base recognition algorithm. The so-called traditional or basic version base recognition algorithm or tool may be a machine learning model or a software that does not involve machine learning. Here, the user can directly use the automatic training tool during the sequencing process. The automatic training tool will automatically call the traditional or basic version base recognition algorithm during the sequencing step of the sequencer, and during the process of using the traditional or basic version algorithm to complete base recognition, it can simultaneously generate the feature values required by the machine learning model based on sequencing intermediate parameters such as the brightness of the original image and the base recognition information of the context, and save them in the specified path. It can be understood that during the conventional sequencing process (without calling the automatic training tool), the so-called sequencing intermediate parameters are usually parameters that are required or usually generated during the base recognition process of the sequencer but generally do not need to be output or output to the user.
[0207] 3. The automatic training process can automatically identify the progress of sequencing. For example, after the sequencing is completed, the automatic training tool can automatically call the alignment tool, and align the obtained fastq file with the reference genome (such as the user input), and output the alignment result to the specified path.
[0208] 4. After the alignment is completed, the automatic training tool will start to automatically generate training data based on the stored feature values and the alignment result. The feature values of the training data have been all output during the sequencing process, and the target values will be extracted based on the alignment result. For details, please refer to part (1) 3 of the foregoing example. The training data will be automatically stored in the specified path.
[0209] 5. After the training dataset is generated, the automatic training tool can divide the training data, automatically complete the modeling of the machine learning model, and automatically complete the development of the Q-value algorithm according to the processes described in the foregoing (1) 4 and 5 and (2).
[0210] 6. After the machine learning model is modeled and the Q-value algorithm is developed, the automatic training tool will automatically save the model file and the algorithm file in the specified path. The model files of the machine learning and the Q-value algorithm can appear as the configuration files of the machine learning base recognition tool. As long as the configuration file is stored in the specific folder called by the base recognition tool, the next time the machine learning base recognition tool is used for sequencing, the configuration will be automatically used for base recognition and Q-value calculation. Therefore, after the modeling is completed, the user can directly copy the configuration file to the path of the corresponding configuration file of the base recognition tool on the sequencer, and the model can be directly used for the next sequencing.
[0211] In the description of this specification, the descriptions of "one embodiment", "some embodiments", "illustrative embodiments", "examples", "certain examples", "specific examples" or "embodiments" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0212] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A system for automatically generating a machine learning model, wherein the machine learning model generated by the system is used to output a quality score to evaluate the accuracy of a base recognition result, characterized in that: include: An input layer, for receiving user input including sequencing results and optionally one or more tasks from a user, wherein the sequencing results are obtained by sequencing nucleic acid molecules connected to multiple positions on a surface using a surface imaging sequencing technology, and the sequencing results include base recognition results and corresponding base recognition probabilities; The machine learning layer includes a training data generation module and a training module. The training data generation module is used to process the user input from the input layer to generate training data, wherein the training data includes a plurality of pairs of associated features and labels, wherein the features describe one or more base recognition probabilities corresponding to the base recognition results and / or the intensities corresponding to the nucleic acid molecules at one or more positions, and the labels include two categories: the base recognition results are consistent with the corresponding reference true value bases and inconsistent with each other, The training module is used to train a specified model based on a portion of the training data to obtain a quality score model, including using a training set to perform multiple iterative training on the specified model and using a validation set to evaluate the performance of the specified model after each iterative training to obtain the quality score model, wherein the training set and the validation set are each independently a non-overlapping portion of the training data; as well as The output layer is used to output a quality score model that meets the preset requirements.
2. The system according to claim 1, characterized in that The sequencing results include no less than 10 million sequence reads and / or no less than 1 billion base pairs of base reads; Optionally, the feature describes one or more base call probabilities.
3. The system according to claim 1 or 2, characterized in that: Training the specified model in the training module also includes, Obtaining classification probabilities of base recognition results in the training set; Dividing the base recognition results into a groups based on the classification probability and / or the base categories in the base recognition results, where a is a natural number not less than 5; For each group classification probability and / or corresponding base recognition result, by comparing the base category in the group base recognition result with the corresponding reference true value base, to determine the error rate or correct rate of the base category in the group base recognition result; as well as A predetermined quality score is associated with the error rate or correct rate of the base class.
4. The system according to claim 1 or 2, characterized in that: Training the specified model in the training module also includes, Obtaining classification probabilities of base recognition results in the training set; Dividing the base recognition results into a groups based on the classification probability and / or the base categories in the base recognition results, and for each group, comparing the base categories in the group with the corresponding reference true value bases to determine the error rate of the base categories in the base recognition results of the group, where a is a natural number and the value of a can achieve a logarithm difference of the error rates of the base categories in the base recognition results of adjacent groups less than 2; as well as associating a predetermined quality score with an error rate for the base class; Optionally, further comprising, Based on the error rate or correct rate of the base categories in the base recognition results of each group, the proportion or distribution of the number of base categories in each group, and the predetermined quality score, the groups are merged into new groups with fewer groups, the attribution of the base categories of adjacent new groups is adjusted, and the error rate of the base category of each new group is calculated, so that the error rate of the base category of the final new group corresponds to the predetermined quality score, so as to establish a mapping relationship between the classification probability of the base category and the predetermined quality score; Optionally, further comprising, Based on the error rates of base categories in the base recognition results of each group and the proportion or distribution of the number of base categories in each group, the groups are merged and regrouped so that the error rates of the base categories in the new groups correspond to the first quality scores, so as to establish a mapping relationship between the classification probability of each base category and the first quality score; Based on the error rates of the base categories of the new groups and the proportion or distribution of the number of base categories of each new group, the new groups are merged and regrouped so that the error rates of the base categories of the final groups correspond to the predetermined quality scores, so as to establish a mapping relationship between the classification probability of each base category and the predetermined quality score, wherein the values in the first quality score include all values of the predetermined quality score; Optionally, the value of the first quality score is a continuous integer, and the predetermined quality score is a non-continuous integer; Optionally, the user input further includes the first quality score; Optionally, it is characterized in that the user input also includes the corresponding reference true value base; Optionally, the user input further includes a reference sequence, and the training data generation module further includes an alignment submodule, the alignment submodule being used to align the sequencing result to the reference sequence to generate an alignment result; and determining the corresponding reference true value base according to the alignment result from the alignment submodule; Optionally, the specified model is selected from at least one of logistic regression, support vector machine and tree model; Optionally, the machine learning layer further includes an evaluation module, which is connected to the training module and the result output layer and is used to evaluate the performance of the quality score model using a test set before outputting the quality score model, so as to determine whether the quality score model meets preset requirements, wherein: The test set is a part of the training data that does not overlap with the training set and the validation set, and the satisfying the preset requirement is that at least one indicator of the accuracy, precision, recall and speed of the classification prediction result of the test set by the quality score model meets the preset standard or can achieve the task input by the user; Optionally, the user input further comprises sequencing data, the sequencing data comprising images of nucleic acid molecules at multiple locations on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs; Optionally, the user input further includes sequencing data, the sequencing data is a portion of data extracted and determined from a sequencing image, the sequencing data includes positions and intensity values of multiple nucleic acid molecules in one or more surface regions where they are located, and the sequencing image is an image of nucleic acid molecules at multiple positions on one or more surfaces obtained from one or more rounds of detection in one or more sequencing runs; Optionally, the sequencing data is processed using a base recognition model to obtain the sequencing results.
5. The system according to claim 4, characterized in that The system further comprises a base recognition model automatic generation system connected to the input layer, the base recognition model automatic generation system being used to generate and output the base recognition model; Optionally, the input layer is a first input layer, the machine learning layer is a first machine learning layer, the output layer is a first output layer, and the base recognition automatic generation system includes: A second input layer, for receiving user input including the sequencing data and optionally one or more tasks from a user, wherein the sequencing data is obtained by detecting nucleic acid molecules attached to multiple positions on a surface using a surface imaging sequencing technology; The second machine learning layer includes a second training data generation module and a second training module, The second training data generation module is used to process the user input from the second input layer to generate second training data, wherein the second training data includes a plurality of pairs of associated second features and second labels, wherein the second features describe the intensity corresponding to the nucleic acid molecules at one or more positions, and the second labels include one selected from the bases A, T, C and G. The second training module is used to train a second designated model based on at least a portion of the second training data, including performing multiple iterations of training on the second designated model using a second training set and evaluating the performance of the second designated model after each iteration of training using a second validation set, so as to obtain a base recognition model, wherein the second training set and the second validation set are each independently a non-overlapping portion of the second training data; and A second output layer, used for outputting the base recognition model; Optionally, the second tag further comprises a combination of any two, three or four of the bases selected from A, T, C and G; Optionally, the user input from the second input layer further includes a reference sequence, and the second training data generation module further includes a second alignment submodule, the second alignment submodule being used to align the sequence reads or base reads corresponding to the sequencing data to the reference sequence to generate a second alignment result; and determining a second label associated with the second feature based on the second alignment result to generate the second training data; Optionally, the second machine learning layer further includes a second evaluation module connected to the second training module and the second output layer, for evaluating the performance of the base recognition model using a second test set before outputting the base recognition model, so as to determine whether the base recognition model meets preset requirements, wherein: The second test set is a portion of the second training data that does not overlap with the second training set and the second validation set, and the satisfying the preset requirement is that at least one indicator of the accuracy, precision, recall and speed of the classification prediction result of the base recognition model for the second test set meets the preset standard or is capable of achieving a task input by a user; Optionally, the second specified model is selected from at least one of logistic regression, support vector machine, tree model and neural network; Optionally, the machine learning layer or the first machine learning layer or the second machine learning layer further comprises a selection module for determining the specified model or the first specified model or the second specified model based on corresponding user input.
6. A method for developing a quality score model, characterized in that The method comprises accessing user input using the system for automatically generating a machine learning model according to any one of claims 1 to 5, and running the system using at least one hardware processor to present an optimal quality score model to a user, wherein the user input comprises sequencing results and one or more tasks optionally included, wherein the sequencing results are obtained by detecting nucleic acid molecules connected to multiple positions on a surface using surface imaging sequencing technology, and wherein the sequencing results comprise base recognition results and base recognition probabilities.
7. The method according to claim 6, characterized in that The sequencing results include no less than 10 million sequence reads and / or no less than 1 billion base pairs of base reads; Optionally, receiving the user input via a model generation framework or interface; and / or presenting the optimal quality score model to a user via the model generation framework or interface; Optionally, the system is operated using two or more operating systems, and interaction between the different operating systems is achieved using at least one of file sharing, remote access, network communication, virtual machines, containers, and virtualization technologies or tools; Optionally, using Windows and Linux to respectively run a part of the system, and using at least one of WSL and / or SSH to implement interaction between different operating systems; Optionally, the system is run using multi-threaded parallel computing.
8. A quality score model, characterized in that: Produced using the system for automatically generating a machine learning model as described in any one of claims 1-5.
9. A computer-readable storage medium storing a plurality of instructions for controlling a processor to execute the method of claim 6 or 7.
10. A system for automatically generating a machine learning model, characterized in that: The system comprises a memory storing a program and one or more processors coupled to the memory, wherein the processors execute the program to implement the method of claim 6 or 7.
Citation Information
Patent Citations
Quality evaluation method and screening method of nucleic acid sequencing data
CN114420214A
Base determination method and device, computer equipment and storage medium
CN115910217A
Gene sequencing base quality evaluation method based on deep learning, product, equipment and medium
CN117726621A
Method for determining mass fraction of read segment, sequencing method and device
CN117976042A
Base identification method and device, equipment and storage medium
CN118942549A
Cited By
System and method for automatically generating machine learning model
WO2026184301A1