Method and system for predicting sequencing quality and computer readable storage medium

By using predictive models to determine sequencing quality during gene sequencing, the problem of low reliability of sequencing results has been solved, and the optimal use of resources and improved sequencing efficiency have been achieved.

CN121641197APending Publication Date: 2026-03-10GENEMIND BIOSCIENCES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Current gene sequencing technologies have low reliability of sequencing results, leading to resource waste and low sequencing efficiency, especially during long-term sequencing processes where detection errors are prone to occur.

Method used

By acquiring the sequencing characteristics of the mixed test samples in a specified fluid channel, a pre-trained regression model is used to predict the sequencing quality, determine the reliability of the sequencing results, and thus decide whether to continue with subsequent sequencing reactions.

Benefits of technology

This reduces the waste of reagents and human resources, improves sequencing efficiency, and ensures the accuracy and reliability of sequencing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121641197A_ABST
    Figure CN121641197A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for determining sequencing quality and a computer readable storage medium. The method for determining the sequencing quality comprises the steps that sequencing characteristics of at least two mixed to-be-tested samples in a specified fluid channel in the current sequencing operation are obtained, the sequencing characteristics can be obtained according to images collected by sequencing reactions from the Ath round to the Bth round in the current sequencing operation, and the at least two to-be-tested samples contain different sample identification codes respectively; a is a natural number greater than or equal to 1, and B is a natural number greater than A; the sequencing features are input into a pre-trained prediction model, the sequencing quality of current sequencing operation is determined according to a prediction result output by the prediction model, and the prediction model is a regression model. According to the method for determining the sequencing quality, the sequencing quality of the current sequencing operation is predicted according to the output of the prediction model, and the credibility of the sequencing result is estimated in advance, so that whether resources need to be continuously input to complete the subsequent sequencing reaction or not can be judged, the situation that resources such as reagents and manpower are wasted can be reduced, and the sequencing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Information

[0002] This application claims priority to and the benefit of the filing date of the patent application with the China National Intellectual Property Office filed on August 28, 2024, with the patent application number 202411204412.7, and incorporates it herein in its entirety by reference. TECHNICAL FIELD

[0003] The present application relates to the field of gene sequencing technology, in particular to a method, system and computer readable storage medium for predicting sequencing quality. BACKGROUND

[0004] In the field of gene sequencing technology, each sequencing chip has multiple fluid channels, and each fluid channel can be used to sequence one or more samples. Currently, the sequencing process of the sample takes a long time, and then comparison and bioinformatics analysis are needed to obtain the sequencing result, which will require a long waiting time. If the sequencing result is significantly different from the actual situation, i.e., the sequencing quality is poor, the credibility of the sequencing result will be low, resulting in waste of reagents and human resources, thereby greatly reducing the sequencing efficiency. SUMMARY

[0005] The present application provides a method, system and computer readable storage medium for determining sequencing quality to solve at least one of the above technical problems.

[0006] The method for determining sequencing quality of the present application comprises:

[0007] obtaining sequencing characteristics of at least two mixed samples to be tested in a specified fluid channel in a current sequencing run, the sequencing characteristics being able to be obtained according to images collected in A to B rounds of sequencing reactions in the current sequencing run, where A is a natural number greater than or equal to 1, and B is a natural number greater than A;

[0008] inputting the sequencing characteristics into a pre-trained prediction model, and determining the sequencing quality of the current sequencing run according to a prediction result output by the prediction model, the prediction model being a regression model.

[0009] The system for determining sequencing quality of the present application comprises a terminal device, which comprises an acquisition unit and a prediction unit;

[0010] The acquisition unit is configured to obtain sequencing characteristics of at least two mixed samples to be tested in a specified fluid channel in a current sequencing run, the sequencing characteristics being able to be obtained according to images collected in A to B rounds of sequencing reactions in the current sequencing run, where A is a natural number greater than or equal to 1, and B is a natural number greater than A;

[0011] The prediction unit is configured to input the sequencing feature into a pre-trained prediction model, and determine the sequencing quality of the current sequencing run according to a prediction result output by the prediction model, wherein the prediction model is a regression model.

[0012] A computer system of the present application comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method for determining sequencing quality of the present application when executing the computer program.

[0013] A computer readable storage medium of the present application stores a computer program, and the computer program implements the steps of the method for determining sequencing quality of the present application when executed by a processor.

[0014] In the method, system and computer readable storage medium for determining sequencing quality, the sequencing feature in the current sequencing run is input into a prediction model, and the sequencing quality of the current sequencing run is determined according to a prediction result output by the prediction model, so that the reliability of the sequencing result is estimated in advance, and it can be determined whether resources need to be continuously invested to complete subsequent sequencing reactions, so as to reduce the waste of reagents and human resources, and improve the sequencing efficiency.

[0015] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0016] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the accompanying drawings.

[0017] Figure 1 is a flowchart of the method for determining sequencing quality of the embodiment of the present application;

[0018] Figure 2 is a module schematic diagram of the system for determining sequencing quality of the embodiment of the present application;

[0019] Figure 3 is another flowchart of the method for determining sequencing quality of the embodiment of the present application;

[0020] Figure 4 is still another flowchart of the method for determining sequencing quality of the embodiment of the present application;

[0021] Figure 5 is still another flowchart of the method for determining sequencing quality of the embodiment of the present application;

[0022] Figure 6 is still another flowchart of the method for determining sequencing quality of the embodiment of the present application;

[0023] Figure 7 This is another flowchart illustrating the method for determining sequencing quality according to an embodiment of the present invention;

[0024] Figure 8 This is another flowchart illustrating the method for determining sequencing quality according to an embodiment of the present invention;

[0025] Figure 9 This is another flowchart illustrating the method for determining sequencing quality according to an embodiment of the present invention;

[0026] Figure 10 This is a schematic diagram illustrating the degree of matching between the training results and label values ​​in an embodiment of the present invention;

[0027] Figure 11 This is another flowchart illustrating the method for determining sequencing quality according to an embodiment of the present invention;

[0028] Figure 12 This is another flowchart illustrating the method for determining sequencing quality according to an embodiment of the present invention;

[0029] Figure 13 This is another flowchart illustrating the method for determining sequencing quality according to an embodiment of the present invention;

[0030] Figure 14 This is another flowchart illustrating the method for determining sequencing quality according to an embodiment of the present invention;

[0031] Figure 15 This is another schematic diagram illustrating the degree of matching between the training results and label values ​​in an embodiment of the present invention;

[0032] Figure 16 This is another schematic diagram illustrating the matching degree between the training results and label values ​​in an embodiment of the present invention;

[0033] Figure 17 This is a schematic diagram of the modules of the system for determining sequencing quality according to an embodiment of the present invention.

[0034] Explanation of key component symbols:

[0035] System 100 for determining sequencing quality;

[0036] Terminal device 10, acquisition unit 11, prediction unit 12, training unit 13;

[0037] Imaging equipment 20;

[0038] Memory 101, processor 102. Detailed Implementation

[0039] In the description of this invention, some of the disclosed content has been shown accordingly in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The following description with reference to the accompanying drawings is exemplary and is only used to explain the invention, and should not be construed as limiting the invention.

[0040] In the description of this invention, many different contents or examples are disclosed to implement different structures of the invention. To simplify the disclosure of this invention, the components and arrangements of specific examples are described below. Of course, these are merely examples and are not intended to limit the invention.

[0041] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0042] Furthermore, the present invention may repeat reference numbers and / or reference letters in different examples. Such repetition is for the purpose of simplification and clarity and does not in itself indicate the relationship between the various situations and / or settings discussed.

[0043] In this invention, the term "sequencing" can be referred to as "nucleic acid sequencing" or "gene sequencing," and these three terms are interchangeable in meaning. All three refer to the determination of the type and sequence of bases or nucleotides (including nucleotide analogs) in a nucleic acid sequence. Sequencing can include the process of binding nucleotides to a template and collecting the corresponding signals emitted by the nucleotides (including analogs). The so-called sequencing can include sequencing by synthesis (SBS, Sequencing by Synthesis) and / or sequencing by ligation (SBL, Sequencing by Ligation), and can include DNA sequencing and / or RNA sequencing.

[0044] Sequencing can include DNA sequencing and / or RNA sequencing. It includes long-fragment sequencing and / or short-fragment sequencing; the terms "long" and "short" are relative, such as nucleic acid molecules longer than 1Kb, 2Kb, 5Kb, or 10Kb being called long fragments, and those shorter than 1Kb or 800bp being called short fragments. It can include paired-end sequencing, single-end sequencing, and / or paired-end sequencing, etc. Paired-end sequencing or paired-end sequencing can refer to the readout of any two segments or parts of the same nucleic acid molecule that do not completely overlap. Sequencing can be performed using a sequencing platform. According to embodiments of this application, selectable sequencing platforms include, but are not limited to, Illumina's HiSeq, MiSeq, Nextseq, and Novaseq sequencing platforms; Thermo Fisher / Life Technologies' Ion Torrent platform; BGI Genomics' BGISEQ and MGISEQ / DNBSEQ platforms; and single-molecule sequencing platforms such as GenoCare1600 from GenoBio. The sequencing method can be single-end sequencing, paired-end sequencing, or any sequencing method supported by the selected automated sequencing platform.

[0045] The sequence obtained from sequencing is called the sequencing sequence, also known as a read.

[0046] In some examples, sequencing-by-synthesis (SBS) is used to perform multiple rounds of sequencing to obtain sequencing sequences or reads. For instance, the nucleic acid molecule to be tested (sometimes referred to as the test sample) is brought into contact with a polymerase and a modified nucleotide and placed under suitable conditions for polymerization. The modified nucleotide is controllably incorporated into the nucleic acid molecule to be tested, or a single-base extension reaction (or base extension reaction) is controllably performed. The corresponding reaction signal is detected, and the type of nucleotide incorporated into the nucleic acid molecule to be tested in this reaction is determined based on this signal. This process of multiple controllable single-base extension reactions and corresponding signal detection is performed to detect the type of nucleotide or base incorporated into the nucleic acid molecule to be tested in multiple or multiple rounds of reactions based on the reaction signal information, so as to read a portion of the sequence of the nucleic acid molecule to be tested.

[0047] The nucleic acid molecule to be tested, also known as the test template, nucleic acid template, or template, can be an unamplified single molecule or an amplified molecular cluster or long chain containing multiple identical polynucleotide molecules, such as the cloning clusters or DNA nanospheres (DNB) formed by bridge amplification or rolling circle amplification used by mainstream sequencing platforms. The nucleic acid molecule to be tested can be single-stranded, double-stranded, and / or hybridized complexes with probes or primers.

[0048] The corresponding reaction signals can be, for example, fluorescence signals, or they can be converted into image data formed by collecting these fluorescence signals. These image data are then processed and analyzed to detect the nucleotides incorporated into the nucleic acid molecule to be tested in each or every sequencing reaction, so as to determine a portion of the base sequence of the nucleic acid molecule to be tested.

[0049] Specifically, in some examples, sequencing is achieved based on surface fluorescence imaging detection. The nucleic acid molecule to be tested is attached to a solid surface. For example, the nucleotide can be modified to have or be able to bind a fluorescent label, as well as a removable inhibitor group that prevents other nucleotides from polymerizing and attaching to the next position of the nucleic acid molecule to be tested (this modified nucleotide is also called a reversible terminator). After each polymerization reaction or single base extension reaction, the fluorescent label is excited to emit a fluorescent signal, and these fluorescent signals are collected to obtain an image of the nucleic acid molecule to be tested that has undergone a single base extension reaction at a specified surface position. Then, the inhibitor group and fluorescent label are removed to perform the next polymerization reaction and signal acquisition (photographing). This process is repeated multiple times to obtain an image set of information related to the nucleotides attached to the nucleic acid molecule to be tested in each single base extension reaction.

[0050] Understandably, when a nucleic acid molecule to be tested undergoes a polymerization reaction at a designated location on the surface, it emits fluorescence. This fluorescence typically appears as a bright spot or bright patch with a higher intensity than the background signal in the corresponding location of the image acquired during that reaction. Therefore, based on the information in these image sets, including the bright spots corresponding to specific chemical characteristics (the nucleic acid molecule to be tested undergoing a polymerization reaction), it is possible to determine whether the nucleic acid molecule to be tested at the designated location has undergone a polymerization reaction. By combining this with a pre-defined correspondence between distinguishable fluorescence signals and nucleotide types, the type of nucleotide that has polymerized and ligated into the nucleic acid molecule to be tested can be detected. This allows for the determination of at least a portion of the sequence of the nucleic acid molecule to be tested, thus obtaining what is known as a read.

[0051] It should be noted that the term "nucleotide" as used herein includes ribonucleic acid or deoxyribonucleic acid, including natural nucleotides or their derivatives or modifications thereof (also known as modified nucleotides or altered nucleotides, etc.). In this document, the term "nucleotide" is sometimes used to refer to the bases contained within it, which will be readily understood by those skilled in the art based on conventional knowledge and / or context.

[0052] A sequencing run can consist of multiple sequencing rounds, with each round potentially including a single base extension reaction (one repeat). For example, four nucleotides (dATP, dTTP, dGTP, and dCTP) carrying different fluorescent labels can be placed in the same polymerization reaction system with multiple target nucleic acid molecules for base extension. After the base extension reaction, the fluorescent labels are excited to emit fluorescence signals. A multi-channel microscopy system collects the fluorescence signals at various wavelengths emitted by the different fluorescent labels, allowing each wavelength to be imaged in its corresponding channel. Analysis of these images determines the type of nucleotide incorporated or introduced into the target nucleic acid molecule. In other words, information obtained from a single base extension reaction can determine the type of nucleotide incorporated or introduced into the target nucleic acid molecule. A single sequencing round can also include multiple base extension reactions, i.e., multiple repeats. For example, four nucleotides carrying the same fluorescent label can be sequentially contacted with multiple nucleic acid molecules to be tested on a surface, and base extension reactions can be performed separately. One round of sequencing includes four base extension reactions. After each base extension reaction, the fluorescent label is excited by excitation light to emit a fluorescent signal. The fluorescence signal emitted by the fluorescent label is collected by a single-channel microscopic imaging system to form an image. Based on the image analysis, the type of nucleotide incorporated or introduced on the nucleic acid molecule to be tested can be determined. Another example is that four nucleotides can be arbitrarily combined to contact multiple nucleic acid molecules to be tested on a surface, such as in pairs or in a one-to-three combination. The nucleotides in each combination carry different fluorescent labels. Two combinations undergo base extension reactions separately. One round of sequencing includes two base extension reactions. After each base extension reaction, the fluorescent label is excited by excitation light to emit a fluorescent signal. The fluorescence signal emitted by different fluorescent labels is collected by a dual-channel microscopic imaging system (such as when combined in pairs), so that the fluorescence signal of each wavelength enters the corresponding channel for imaging to form an image. Based on the image analysis, the type of nucleotide incorporated or introduced on the nucleic acid molecule to be tested can be determined.

[0053] Currently, gene technology is being developed and applied with increasing speed. In order to determine the nucleic acid sequence for subsequent use, it is particularly important to accurately detect the nucleic acid sequence.

[0054] In practical applications, because the nucleic acid sequences or inserts of the test samples are often quite long, many rounds of sequencing reactions are required. For example, if a nucleic acid sequence contains 68 bases or nucleotides, or if the sequence length is 68, then 68 rounds of sequencing reactions are needed to detect the bases or nucleotides arranged within that sequence. Furthermore, when sequencing mixed test samples, in order to identify which test sample the detected sequence comes from, a sample identification code (called a barcode) composed of multiple bases or nucleotides is added to the nucleic acid sequence of the test samples to distinguish between different test samples, further increasing the number of sequencing rounds.

[0055] Based on the above, it's easy to see that as the number of sequencing rounds increases, the sequencing time also increases accordingly. In one example, 71 rounds of sequencing takes approximately 26 hours to complete. After all rounds of sequencing are completed, another 3 hours are needed to align the obtained sequences or reads and perform bioinformatics analysis to finally obtain the sequencing results. In other words, a significant investment of time, materials, and manpower is required before obtaining the final sequencing results. Even after obtaining the sequencing results, objective factors (such as detection errors during the sequencing process) may affect the final results, reducing their reliability. Large deviations indicate sequencing failure, resulting in wasted resources and impacting the sequencing efficiency of the samples to be tested.

[0056] The present invention involves extracting sequencing features based on the previous sequencing reactions when sequencing the sample to be tested, and then determining the sequencing quality based on the sequencing features. If the sequencing quality is good and meets the preset requirements, the subsequent sequencing reactions can be continued. If the sequencing quality is poor and does not meet the preset requirements, the subsequent sequencing reactions can be stopped in time and the sequencing protocol can be adjusted, thereby reducing the cost waste of subsequent processes and improving sequencing efficiency.

[0057] Specifically, according to the technical solution of the present invention, during the sequencing of the sample to be tested, relevant data detected in the previous sequencing reactions are processed to obtain sequencing features. These sequencing features are then input into a prediction model, causing the prediction model to output corresponding prediction results. The sequencing quality of the sample to be tested is then determined based on the prediction results. If the sequencing quality is good and meets the preset requirements, the sequencing results are considered reliable; if the sequencing quality is poor and does not meet the preset requirements, the sequencing results are considered unreliable.

[0058] It is understandable that sequencing reactions can visualize the types of fluorescently labeled bases or nucleotides incorporated into the sample, thus making it easier to determine the types of bases or nucleotides incorporated into the sample. However, in the process of achieving visualization, it is easily affected by various factors such as phasing and prephasing, which can lead to deviations between the sequencing results and the actual situation. Moreover, if the sequence length of the sample is long, such deviations will accumulate in multiple rounds of sequencing reactions, resulting in a significant difference between the final sequencing results and the actual situation.

[0059] In this invention, the prediction result can be correlated with the number of bases identified in the sample during sequencing. Specifically, after obtaining the prediction result based on the prediction model, the prediction result can be used to determine whether the number of bases identified in the sample during the current sequencing process is within a reasonable range. If it is within a reasonable range, it can be determined that the deviation between the sequencing result and the actual result is small, and the reliability of the sequencing result is high. Furthermore, after obtaining the prediction results from the prediction model, the results can be used to determine whether the number of bases identified in the current sequencing process is sufficient. If the number of bases identified is sufficient, it means that a considerable number of bases will participate in the sequencing reaction, thereby increasing the proportion of bases that normally participate in the sequencing reaction among all bases that should participate in the sequencing reaction, reducing the deviation between the sequencing results and the actual results, and thus improving the reliability of the sequencing results. If the number of bases identified is insufficient, it means that fewer bases participate in the sequencing reaction, and the proportion of bases that do not normally participate in the sequencing reaction among all bases that should participate in the sequencing reaction increases. This will increase the deviation between the sequencing results and the actual results, thereby reducing the reliability of the sequencing results. In this case, it can be determined that the current sequencing protocol needs to be changed and the sequencing reaction restarted to avoid wasting subsequent resources.

[0060] In this invention, the prediction model can be a pre-trained model, enabling it to output prediction results when inputting sequencing features. Before using the prediction model to determine sequencing quality, it can be trained using an initial model. Specifically, training features obtained from sequencing training samples can be input into the initial model. After generating output values, the initial model can compare these output values ​​with the label values ​​of the training samples. Based on the comparison results, it can be determined whether the output values ​​of the initial model match the actual situation, and the training of the initial model can be adjusted and optimized until the difference between the output values ​​and the label values ​​is sufficiently small. The label values ​​of the training samples can be the true values ​​obtained from pre-sequencing the training samples.

[0061] Please refer to Figure 1 One method for determining sequencing quality according to the present invention may include:

[0062] 02: Obtain the sequencing features of at least two mixed test samples in the current sequencing run within a specified fluid channel. The sequencing features can be obtained from the images collected in the sequencing reactions of rounds A to B in the current sequencing run. The at least two test samples contain different sample identification codes, where A is a natural number greater than or equal to 1 and B is a natural number greater than A.

[0063] 03: Input the sequencing features into the pre-trained prediction model, and determine the sequencing quality of the current sequencing run based on the prediction results output by the prediction model. The prediction model is a regression model.

[0064] The method for determining sequencing quality according to the present invention can be implemented using the system 100 for determining sequencing quality according to the present invention. Specifically, please refer to... Figure 2 The system 100 includes a terminal device 10, which includes an acquisition unit 11 and a prediction unit 12. The acquisition unit 11 can acquire the sequencing features of at least two mixed test samples within a specified fluid channel during the current sequencing run. These sequencing features are obtained from images acquired during sequencing reactions A through B in the current sequencing run. Each of the at least two test samples contains a different sample identification code, where A is a natural number greater than or equal to 1, and B is a natural number greater than A. The prediction unit 12 can input the sequencing features into a pre-trained prediction model and determine the sequencing quality of the current sequencing run based on the prediction results output by the prediction model. The prediction model is a regression model.

[0065] In the above-mentioned method and system 100 for determining sequencing quality, the sequencing characteristics of the current sequencing run are input into the prediction model, and the sequencing quality is determined based on the prediction results output by the prediction model. The reliability of the sequencing results is estimated in advance, thereby determining whether it is necessary to continue to invest resources to complete the subsequent sequencing reaction. This can reduce the waste of reagents and manpower resources and help improve sequencing efficiency.

[0066] Specifically, during the sequencing process, each sequencing round can generate corresponding sequencing features. By obtaining sequencing features from images acquired under multiple sequencing reactions, the sequencing quality of sequencing reactions A to B can be determined based on the sequencing features, and the sequencing quality of the remaining sequencing rounds can be inferred, thereby predicting the reliability of the sequencing results of the current sequencing run.

[0067] In addition, sequencing results can be the sequences obtained by sequencing the sample to be tested. Each sequencing run can perform a sequencing of the sample to be tested, thereby obtaining sequencing results, which can then be used to determine the sequencing accuracy and quality of the sample to be tested.

[0068] exist Figure 2In addition, system 100 may also include imaging device 20. Specifically, system 100 can use imaging device 20 to acquire signals generated by the sample during sequencing to form an image, and can use terminal device 10 to obtain sequencing features based on the acquired image.

[0069] In this invention, the prediction results include the ratio of sequencing throughput of at least two test samples in a specified fluid channel obtained in the current sequencing run.

[0070] It is understandable that the sequencing quality of a sequencing run can be determined by whether its throughput is within a reasonable range. Furthermore, the sequencing throughput of a single run is affected by the amount of relevant information reflected in the acquired images. More relevant information means more bases can be identified, making the sequencing results less susceptible to accidental factors (such as a single base pairing error during base extension), and thus more accurately determining the sequencing quality.

[0071] In this invention, the prediction model may include a first sub-model, and the sequencing features may include a first type of features.

[0072] Please refer to Figure 3 Step 03 (inputting sequencing features into a pre-trained prediction model and determining the sequencing quality for the current sequencing run based on the prediction results output by the prediction model) may include:

[0073] 031: Input the first type of features obtained by sequencing the sample identification codes of at least two test samples in the A to C rounds of sequencing reactions into the first sub-model to output the ratio of sequencing throughput of at least two test samples;

[0074] 032: When the sequencing throughput ratio is within the first expected range, determine that the sequencing quality of the current sequencing run meets the preset requirements, that is, the sequencing results are reliable;

[0075] 033: If the sequencing throughput ratio exceeds the first expected range, it is determined that the sequencing quality of the current sequencing run does not meet the preset requirements, that is, the sequencing results are unreliable.

[0076] The method for determining sequencing quality according to the present invention can be implemented using the system 100 for determining sequencing quality according to the present invention. Specifically, please refer to... Figure 2The prediction unit 12 can be used to: input the first type of features collected by sequencing the sample identification codes of at least two test samples in the A to C rounds of sequencing reactions into the first sub-model to output the sequencing throughput ratio of at least two test samples; if the sequencing throughput ratio is within the first expected range, determine that the sequencing quality of the current sequencing run meets the preset requirements, that is, the sequencing result is reliable; if the sequencing throughput ratio exceeds the first expected range, determine that the sequencing quality of the current sequencing run does not meet the preset requirements, that is, the sequencing result is unreliable.

[0077] This helps improve the accuracy of predictions of actual sequencing results.

[0078] Specifically, each type of sample to be tested will generate a certain amount of data during the sequencing process, i.e., sequencing throughput. The sequencing throughput of a sample to be tested can be the amount of data generated per unit time. In this invention, different types of samples to be tested can be sequenced simultaneously, and the ratio of sequencing throughput can be the ratio of sequencing throughput between different types of samples to be tested.

[0079] Before sequencing at least two mixed test samples within a designated fluid channel, a ratio exists between the concentrations of the test samples within the channel. Based on this concentration ratio, a first expected range for the sequencing throughput ratio between the test samples can be determined. If the first type of feature is input into the first sub-model and the output sequencing throughput ratio between the test samples is within the first expected range, then the sequencing quality meets the preset requirements, and the sequencing results are reliable. If the output sequencing throughput ratio between the test samples is not within the first expected range, then the sequencing quality does not meet the preset requirements, and the sequencing results are unreliable. For example, if the concentration ratio between the test samples within the designated fluid channel is 1.2:1, and the first test sample suffers relatively more losses during sequencing due to various factors, then the first expected range for the sequencing throughput ratio between the two test samples is 1:1.

[0080] It is understandable that when sequencing two or more mixed samples is required, the above-described scheme can be used to determine the sequencing throughput ratio among the various samples, and then the sequencing quality can be determined by the sequencing throughput ratio. In scenarios where multiple samples need to be sequenced simultaneously, this method helps to improve sequencing efficiency while ensuring that the sequencing results match the actual results. It should be noted that the aforementioned different samples can be multiple different samples of the same type, or multiple different samples of different types.

[0081] The first sub-model can be a pre-trained learning model. In some cases, the first sub-model can be loaded into the terminal device 10, so that the collected first-class features can be input into the terminal device 10, and the terminal device 10 can run the first sub-model to output the sequencing throughput ratio between at least two types of test samples.

[0082] Alternatively, step 03 (inputting sequencing features into a pre-trained prediction model and determining the sequencing quality of the current sequencing run based on the prediction results output by the model) can include:

[0083] The first type of features collected from the sequencing reactions A to C are input into the first sub-model to output a first probability value. If the first probability value is greater than or equal to the first probability threshold, it is determined that the sequencing throughput ratio between at least two types of test samples meets the standard, thereby determining that the sequencing quality meets the preset requirements, i.e., the sequencing result is reliable. If the first probability value is less than the first probability threshold, it is determined that the sequencing throughput ratio between at least two types of test samples does not meet the standard, thereby determining that the sequencing quality does not meet the preset requirements, and the sequencing result is unreliable.

[0084] For example, in one example, the first type of features collected from sequencing rounds A to C are input into the first sub-model, outputting a first probability value of 0.6. This first probability value characterizes the likelihood that the sequencing throughput ratio between at least two test samples meets the standard. The first probability threshold is 0.5. Since the first probability value is greater than the first probability threshold, it is determined that the sequencing throughput ratio between at least two test samples meets the standard, thus confirming that the sequencing quality meets the preset requirements, i.e., the sequencing results are reliable. In another example, the first type of features collected from sequencing rounds A to C are input into the first sub-model, outputting a first probability value of 0.4. This first probability value characterizes the likelihood that the sequencing throughput ratio between at least two test samples meets the standard, with a first probability threshold of 0.5. Since the first probability value is less than the first probability threshold, it is determined that the sequencing throughput ratio between at least two test samples does not meet the standard, thus confirming that the sequencing quality does not meet the preset requirements, i.e., the sequencing results are unreliable.

[0085] In this invention, C is the minimum number of sequencing rounds required to complete the sequencing of the sample identification code, and C is a natural number greater than A and less than B.

[0086] Specifically, when A is 1, it corresponds to obtaining the first type of feature from the first round of sequencing reaction, thus allowing for the rapid acquisition of the sequencing throughput ratio between different types of test samples. C can be determined by the number of bases contained in the sample identification code. For example, for a test sample with sample identification code AAA, C can be 3, meaning that three rounds of sequencing reactions are required to complete the sequencing of sample identification code AAA.

[0087] In this invention, the prediction results include the densities of at least two test samples in a specified fluid channel, the prediction model may include a second sub-model, and the sequencing features may include second-class features collected during sequencing of the insert fragments of at least two test samples in the E to B rounds of sequencing reactions, wherein E is a natural number greater than A and less than B.

[0088] Please refer to Figure 4 Step 03 (inputting sequencing features into a pre-trained prediction model and determining the sequencing quality for the current sequencing run based on the prediction results output by the prediction model) may include:

[0089] 034: Input the second type of features collected during sequencing of insert fragments of at least two test samples in the E to B rounds of sequencing reactions into the second sub-model to output the density of at least two test samples in a specified fluid channel;

[0090] 035: When the density is within the second expected range, determine that the sequencing quality of the current sequencing run meets the preset requirements, that is, the sequencing results are reliable;

[0091] 036: If the density exceeds the second expected range, it is determined that the sequencing quality of the current sequencing run does not meet the preset requirements, that is, the sequencing results are unreliable.

[0092] The method for determining sequencing quality according to the present invention can be implemented using the system 100 for determining sequencing quality according to the present invention. Specifically, please refer to... Figure 2 The prediction unit 12 can be used to: input the second type of features collected during sequencing of insert fragments of at least two test samples in the E to B rounds of sequencing reactions into the second sub-model to output the density of at least two test samples in a specified fluid channel; if the density is within the second expected range, determine that the sequencing quality of the current sequencing run meets the preset requirements, i.e., the sequencing result is reliable; if the density exceeds the second expected range, determine that the sequencing quality of the current sequencing run does not meet the preset requirements, i.e., the sequencing result is unreliable.

[0093] This helps improve the accuracy of predictions of actual sequencing results.

[0094] In related technologies, sequencing chips include multiple fluid channels. The sample to be tested is "planted" within these fluid channels, thus immobilizing the sample. In some cases, the density of the sample within the fluid channels can be defined as the number of samples per unit area of ​​the fluid channel. Higher density results in more sequencing sequences and higher sequencing throughput; lower density results in fewer sequencing sequences and lower sequencing throughput.

[0095] It is understandable that, although the probability of sequencing errors is relatively low, they are still random and difficult to detect. By determining the density of the sample to be tested in a specified fluid channel, the probability of sequencing errors occurring in the sample to be tested can be determined, and thus the degree of impact of sequencing errors on the final sequencing results can be determined.

[0096] Specifically, in this invention, before sequencing at least two mixed test samples within a designated fluid channel, the concentration of the test samples is known. Based on the concentration of the test samples, a second expected range corresponding to the density of the test samples within the designated fluid channel can be determined. Alternatively, regardless of the concentration of the introduced test samples, the second expected range corresponding to the density of the test samples within the designated fluid channel can be directly set. If the second type of feature is input into the second sub-model, and the output shows that the density of the test samples within the designated fluid channel is within the second expected range, then it is determined that the sequencing quality meets the preset requirements, i.e., the sequencing results are reliable. If the output shows that the density of the test samples within the designated fluid channel is not within the second expected range, then it is determined that the sequencing quality does not meet the preset requirements, i.e., the sequencing results are unreliable. For example, the second expected range corresponding to the density of the test samples within the designated fluid channel is 3-8, in units of particles / square micrometer. Another example is that the second expected range corresponding to the density of the test samples within the designated fluid channel is 4-6, in units of particles / square micrometer.

[0097] The second sub-model can be a pre-trained learning model. In some cases, the second sub-model can be loaded into the terminal device 10, so that the collected second type of features can be input into the terminal device 10, and the terminal device 10 can run the second sub-model to output the density of the test sample in the specified fluid channel.

[0098] Alternatively, step 03 (inputting sequencing features into a pre-trained prediction model and determining the sequencing quality of the current sequencing run based on the prediction results output by the prediction model) can include:

[0099] The second type of features collected from sequencing rounds E to B are input into the second sub-model to output a second probability value. If the second probability value is greater than or equal to the second probability threshold, it is determined that the density of the sample in the designated fluid channel meets the standard, thus confirming that the sequencing quality meets the preset requirements, i.e., the sequencing result is reliable. If the second probability value is less than the second probability threshold, it is determined that the density of the sample in the designated fluid channel does not meet the standard, thus confirming that the sequencing quality does not meet the preset requirements, i.e., the sequencing result is unreliable. For example, in one example, the second type of features collected from sequencing rounds E to B are input into the second sub-model, outputting a second probability value of 0.6. This second probability value characterizes the probability of whether the density of the sample in the designated fluid channel meets the standard. The second probability threshold is 0.5. Since the second probability value is greater than the second probability threshold, it is determined that the density of the sample in the designated fluid channel meets the standard, thus confirming that the sequencing quality meets the preset requirements, i.e., the sequencing result is reliable. In another example, the second type of features collected from the E to B rounds of base extension reactions are input into the second sub-model, and the output second probability value is 0.4. This second probability value is used to characterize the probability that the density of the sample to be tested in the specified fluid channel meets the standard. The second probability threshold is 0.5. Since the second probability value is less than the second probability threshold, it is determined that the density of the sample to be tested in the specified fluid channel does not meet the standard, thereby determining that the sequencing quality does not meet the preset requirements, that is, the sequencing results are unreliable.

[0100] In this invention, A can be 1, and E can be a natural number greater than 1 and less than B.

[0101] Specifically, E can be calibrated through pre-testing or determined based on the actual sequencing process. A larger B value results in more Type II features, leading to a better match between the sequencing results obtained from these features and the actual sequencing results. However, a larger B value also means more sequencing rounds are required, increasing the time taken. Therefore, the B value should not be too large. The goal is to obtain accurate sequencing results while minimizing the number of sequencing reactions. In some cases, B can be 6, 7, or 8.

[0102] In this invention, the prediction model may include a first sub-model and a second sub-model. Sequencing features may include a first type of feature acquired during sequencing of the sample identification codes of the at least two test samples in sequencing rounds A to C, and a second type of feature acquired during sequencing of the insert fragments of the at least two test samples in sequencing rounds E to B, wherein C is the minimum number of sequencing rounds required to complete the sequencing of the sample identification codes, C is a natural number greater than A and less than E, and E is a natural number greater than C and less than B.

[0103] Please refer to Figure 5Step 03 (inputting sequencing features into a pre-trained prediction model and determining the sequencing quality of the current sequencing run based on the prediction results output by the prediction model) may include:

[0104] 037: Input the first type of features collected by sequencing the sample identification codes of at least two test samples in the sequencing reactions of rounds A to C into the first sub-model to output the sequencing throughput ratio between at least two test samples, and input the second type of features collected by sequencing the insert fragments of at least two test samples in the sequencing reactions of rounds E to B into the second sub-model to output the density of at least two test samples in the specified flow channel;

[0105] 038: If the sequencing throughput ratio is within the first expected range and the density is within the second expected range, the sequencing quality of the current sequencing run is determined to meet the preset requirements, that is, the sequencing results are reliable.

[0106] 039: If the sequencing throughput ratio exceeds the first expected range and / or the density exceeds the second expected range, it is determined that the sequencing quality of the current sequencing run does not meet the preset requirements, that is, the sequencing results are unreliable.

[0107] The method for determining sequencing quality according to the present invention can be implemented using the system 100 for determining sequencing quality according to the present invention. Specifically, please refer to... Figure 2 The prediction unit 12 can be used to: input the first type of features collected during sequencing of sample identification codes of at least two test samples in the A to C rounds of sequencing reactions into the first sub-model to output the sequencing throughput ratio between at least two test samples; and input the second type of features collected during sequencing of insert fragments of at least two test samples in the E to B rounds of sequencing reactions into the second sub-model to output the density of the test sample in the specified flow channel; if the sequencing throughput ratio is within the first expected range and the density is within the second expected range, determine that the sequencing quality of the current sequencing run meets the preset requirements, i.e., the sequencing result is reliable; if the sequencing throughput ratio exceeds the first expected range and / or the density exceeds the second expected range, determine that the sequencing quality of the current sequencing run does not meet the preset requirements, i.e., the sequencing result is unreliable.

[0108] This helps improve the accuracy of predictions of actual sequencing results.

[0109] Specifically, in this invention, by combining the sequencing throughput ratio and density to comprehensively judge the sequencing quality and the reliability of the sequencing results, the accuracy of predicting the final sequencing results can be further improved. Specifically, the sequencing quality is considered to meet the preset requirements, i.e., the sequencing results are reliable, only when both the sequencing throughput ratio and density are within the first expected range and the density is within the second expected range. If at least one of the following conditions is met, the sequencing quality is considered to not meet the preset requirements, i.e., the sequencing results are unreliable.

[0110] In addition, in some cases, the terminal device 10 can determine the sequencing throughput ratio between different types of test samples through one process and determine whether the sequencing quality meets the preset requirements and whether the sequencing results are reliable based on the sequencing throughput ratio. It can also determine the density of the test sample in a specified fluid channel through another process and determine whether the sequencing quality meets the preset requirements and whether the sequencing results are reliable based on the density. In other words, the scheme that judges based on the sequencing throughput ratio and the scheme that judges based on the density can be executed in parallel, thereby saving data processing time and improving processing efficiency.

[0111] Alternatively, step 03 (inputting sequencing features into a pre-trained prediction model and determining the sequencing quality of the current sequencing run based on the prediction results output by the prediction model) can include:

[0112] The first type of features collected from sequencing rounds A to C are input into the first sub-model to output a first probability value, and the second type of features collected from base extension rounds E to B are input into the second sub-model to output a second probability value. If the first probability value is greater than or equal to the first probability threshold and the second probability value is greater than or equal to the second probability threshold, it is determined that the sequencing quality of the current sequencing run meets the preset requirements, i.e., the sequencing result is reliable. If the first probability value is less than the first probability threshold and / or the second probability value is less than the second probability threshold, it is determined that the sequencing quality of the current sequencing run does not meet the preset requirements, i.e., the sequencing result is unreliable.

[0113] For example, in one instance, the first type of features collected from rounds A to C of base extension reactions are input into the first sub-model, and the second type of features collected from rounds E to B of base extension reactions are input into the second sub-model. The outputs a first probability value of 0.7 and a second probability value of 0.6. The first probability value is used to characterize whether the sequencing throughput ratio between at least two test samples meets the standard, and the second probability value is used to characterize whether the density of the test sample in the specified fluid channel meets the standard. The first probability threshold and the second probability threshold are both 0.5. Since the first probability value is greater than the first probability threshold and the second probability value is greater than the second probability threshold, it is determined that the sequencing quality of the current sequencing run meets the preset requirements, that is, the sequencing results are reliable. In another example, the first type of features collected from the base extension reactions from rounds A to C are input into the first sub-model, and the second type of features collected from the base extension reactions from rounds E to B are input into the second sub-model. The first probability value is output as 0.7, the second probability value is output as 0.4, and the first probability threshold and the second probability threshold are both 0.5. Since the second probability value is less than the second probability threshold, it is determined that the sequencing quality of the current sequencing run does not meet the preset requirements, that is, the sequencing results are unreliable.

[0114] In this invention, A can be 1.

[0115] Specifically, when A is 1, it corresponds to obtaining the first type of feature from the first round of base extension reaction, thus quickly obtaining the sequencing throughput ratio between different types of samples. C and E can be calibrated through pre-testing or determined based on the actual sequencing process. In some cases, C can be 3, meaning three sequencing reactions are needed to complete the sequencing of the sample identifier. The larger the value of B, the more second-type features are obtained, and the better the sequencing results obtained through the second-type features match the actual sequencing results. However, a larger value of B means more sequencing reactions are required, and the time required is longer, so the value of B should not be too large. It is necessary to consider obtaining accurate sequencing results while minimizing the number of sequencing reactions. In some cases, the value of B can be 6, 7, or 8.

[0116] In this invention, the first expected range can have a maximum value and a minimum value. When at least two test samples include only two test samples, the minimum value can be in the range of [0.8, 1], and the maximum value can be in the range of [1, 1.25].

[0117] This makes it easy to determine whether the sequencing results are reliable.

[0118] Furthermore, the selection of the endpoints of the first expected range can be determined or adjusted based on the actual sequencing process. In some cases, when sequencing two samples, A and B, if it can be determined that sample A may not be accurately measured due to various factors during sequencing, the endpoints of the first expected range can be adjusted based on the loss of samples A. For example, in one case, the concentration ratio between samples A and B is 1.2:1. The minimum endpoint of the first expected range can be set to 1, and the maximum endpoint can be set to 1.2. Correspondingly, the sequencing throughput ratio of samples A and B can be set to A:B = 1.2:1 before sequencing preparation. However, considering that some bases of sample A cannot be accurately measured, the desired final sequencing throughput ratio of the two samples is A:B = 1:1.

[0119] In this invention, the image is obtained by acquiring the signal generated by the sequencing reaction of the current sequencing run through a multi-channel microscopic imaging system, which includes a first channel and a second channel.

[0120] The first type of feature may include at least one of the following:

[0121] The ratio of the total number of sequences belonging to different test samples in the A to C rounds of sequencing reactions;

[0122] In sequencing rounds A to C, the ratio of the number of signals generated in the first channel to the number of signals generated in the second channel; in sequencing rounds A to C, the ratio of the sum of the number of signals generated in the first channel to the sum of the number of signals generated in the second channel; in sequencing rounds D to C, the ratio of the sum of the number of signals generated in the first channel to the sum of the number of signals generated in the second channel, where D is a natural number greater than A and less than C.

[0123] In this way, it becomes clear how to obtain the first type of features.

[0124] In this invention, the ratio of the total number of sequences belonging to different test samples can be the ratio of the total number of sequences containing A and T bases belonging to different test samples, or the ratio of the total number of sequences containing C and G bases belonging to different test samples. Specifically, when at least two test samples include only two test samples, and the sample identification codes of the two test samples are "AAA" and "TTT" respectively, the ratio of the total number of sequences containing A and T bases belonging to different test samples in the A to C rounds of sequencing reactions can determine the sequencing throughput ratio between the two test samples.

[0125] Of course, it is understandable that in other cases, the sample to be tested can also use other sample identification codes. Specifically, in some cases, the sample identification code of one sample to be tested can be "CCC", and the sample identification code of another sample to be tested can be "GGG".

[0126] In this invention, the ratio of the total number of sequences containing A bases and T bases belonging to different test samples can be calculated using the following formula:

[0127] (A1T0+A2T0+A3T0+A2T1+A3T1+A3T2) / (A0T1+A0T2+A0T3+A1T2+A1T3+A2T3), where "A1T0+A2T0+A3T0+A2T1+A3T1+A3T2" represents the total number of sequences belonging to the test samples with sample identification code AAA; "A0T1+A0T2+A0T3+A1T2+A1T3+A2T3" represents the total number of sequences belonging to the test samples with sample identification code TTT;

[0128] “A1T0” indicates the number of sequences containing 1 A base and 0 T bases;

[0129] “A2T0” indicates the number of sequences containing 2 A bases and 0 T bases;

[0130] “A3T0” indicates the number of sequences containing 3 A bases and 0 T bases;

[0131] “A2T1” indicates the sequence number containing 2 A bases and 1 T base;

[0132] “A3T1” indicates the sequence number containing 3 A bases and 1 T base;

[0133] "A3T2" indicates the sequence number containing 3 A bases and 2 T bases;

[0134] “A0T1” indicates the number of sequences containing 0 A bases and 1 T base;

[0135] “A0T2” indicates the number of sequences containing 0 A bases and 2 T bases;

[0136] “A0T3” indicates the number of sequences containing 0 A bases and 3 T bases;

[0137] “A1T2” indicates the sequence number containing 1 A base and 2 T bases;

[0138] "A1T3" indicates the sequence number containing 1 A base and 3 T bases;

[0139] "A2T3" indicates a sequence number containing 2 A bases and 3 T bases.

[0140] This helps improve the accuracy of sample identification code detection.

[0141] Specifically, ideally, in rounds A to C (e.g., rounds 1-3 of sequencing), sequencing the sample identification code AAA of one type of sample should yield three A bases, while sequencing the sample identification code TTT of another sample should yield three T bases. However, in reality, under-sequencing or incorrect sequencing may occur. For example, in rounds A to C (e.g., rounds 1-3 of sequencing), sequencing the sample identification code AAA of one type of sample might only yield one A base, while the other two A bases are either not detected or are mistakenly sequenced as T bases.

[0142] By calculating the ratio of the total number of sequences belonging to the sample identification code AAA to the total number of sequences belonging to the sample identification code TTT, the situation of under-detection or incorrect base detection can be largely avoided, thereby improving the detection accuracy of the sample identification code.

[0143] Please refer to Figure 6 and Figure 7 In this invention, step 02 (obtaining the sequencing characteristics of at least two mixed test samples within a specified fluid channel during the current sequencing run) may include:

[0144] 011: Input the first training feature into the first initial regression model to obtain the first initial prediction result. The first training feature includes at least one first-class feature obtained by sequencing the training samples.

[0145] 012: The first initial prediction result is evaluated by indicators to obtain the first evaluation result, and the model parameters corresponding to the first training feature in the first initial regression model are adjusted according to the first evaluation result for retraining.

[0146] The method for determining sequencing quality according to the present invention can be implemented using the system 100 for determining sequencing quality according to the present invention. Specifically, please refer to... Figure 2 The terminal device 10 may include a training unit 13. The training unit 13 may be used to: input a first training feature into a first initial regression model to obtain a first initial prediction result, wherein the first training feature includes at least one first type of feature obtained by sequencing training samples; perform index evaluation processing on the first initial prediction result to obtain a first evaluation result, and adjust the model parameters corresponding to the first training feature in the first initial regression model according to the first evaluation result for retraining.

[0147] In this way, the sequencing features can be automatically identified and detected through the learning model.

[0148] Specifically, the first initial regression model can be a regression model, that is, a model trained by using the first training feature as the independent variable and the sequencing throughput ratio as the dependent variable. The training samples can be samples with corresponding sequences that have already been sequenced, that is, the sequences of the training samples are known. The actual sequencing throughput ratio of the training samples can be predetermined. The first training feature may include one or more first-class features, and the first initial prediction result is the sequencing throughput ratio predicted after inputting the first training feature into the first initial regression model.

[0149] In some cases, evaluating the initial prediction results can involve comparing the output of the initial regression model with the actual sequencing throughput, and then determining the initial evaluation result based on the closeness of the two results. If the two are sufficiently close, it can be determined whether the training of the initial regression model has achieved the expected training effect. If there is still a certain deviation, it can be determined that the expected training effect has not been achieved, and the initial regression model needs to be trained again.

[0150] In some cases, the initial regression model can be:

[0151] y1=β 10 +β 11 X 11 ++β 12 X 12 +...+β 1n X 1n ;

[0152] Where, β 1i (i = 1, ..., n) are the model parameters in the first initial regression model, X 1i (i = 1, ..., n) represent the first training features input to the first initial regression model. The above model can be understood as follows: in the first initial regression model, corresponding weights, β, are assigned to the input first training features. 10 It can be considered as an independent factor that is not affected by the first type of feature during the sequencing process.

[0153] Depending on the specific testing scenario, one type of first-class feature from the training samples can be selected and substituted into one of the first training features, or multiple types of first-class features from the training samples can be selected and substituted into their respective first training features. When it is necessary to adjust the model parameters in the initial regression model, the weights of one of the first training features can be adjusted, or the weights of multiple first training features can be adjusted.

[0154] Please refer to Figure 8 and Figure 9In this invention, step 012 (performing index evaluation processing on the first initial prediction result to obtain a first evaluation result, and adjusting the model parameters corresponding to the first training feature in the first initial regression model according to the first evaluation result for retraining) may include:

[0155] 013: If the goodness of fit of the first initial regression model meets the first evaluation condition based on the first evaluation result, and / or the absolute mean error of the first initial regression model meets the second evaluation condition, then the first initial regression model obtained by the current training can be used as the first sub-model.

[0156] The method for determining sequencing quality according to the present invention can be implemented using the system 100 for determining sequencing quality according to the present invention. Specifically, please refer to... Figure 2 The training unit 13 can be used to determine that the first initial regression model obtained by the current training can be used as the first sub-model when the goodness of fit of the first initial regression model meets the first evaluation condition based on the first evaluation result, and / or the absolute mean error of the first initial regression model meets the second evaluation condition.

[0157] In this way, it can be determined whether the first initial regression model has completed training.

[0158] Specifically, please combine Figure 10 ,exist Figure 10 In the diagram, the horizontal axis represents the sequencing data number obtained from sequencing the two training samples. For example, the first sequencing data obtained from the first sequencing run on the two training samples is numbered 1, and so on. Each sequencing data includes multiple reads or sequencing sequences. The vertical axis represents the ratio of sequencing throughput between the two training samples in the same channel. The dots forming the line segment represent the actual ratio of sequencing throughput between the two training samples in the same channel, while the scattered triangles represent the ratio of sequencing throughput between the two training samples in the same channel predicted by the first initial regression model. The actual ratio of sequencing throughput between the two training samples in the same channel can be used as the label value for the first initial prediction result with the same number.

[0159] In some cases, the first evaluation criterion can be a goodness-of-fit of the first initial regression model greater than 0.8, and the second evaluation criterion can be an absolute mean error of the first initial regression model less than 0.1. For example, if the goodness-of-fit is determined to be 0.85 and the absolute mean error to be 0.0564 based on the output of the first initial regression model, then the current first initial regression model can be determined to have a high predictive effect on the test sample. The closer the goodness-of-fit of the first initial regression model is to 1, the better the fit to the actual situation. The closer the absolute mean error of the first initial regression model is to 0, the better the fit to the actual situation.

[0160] In this invention, the evaluation of the first initial prediction result can be achieved using the following formula:

[0161]

[0162] Among them, R1 2 This represents the goodness of fit of the first initial regression model, where m1 represents the number of initial predictions, and y 1i This represents the label value corresponding to the i-th initial prediction result. Let y1' represent the i-th initial prediction result, and y1' represent the average label value corresponding to the initial prediction result.

[0163] It is understandable that by determining the goodness of fit of the first initial regression model, it can be determined whether the results output by the first initial regression model are sufficiently fitted to the corresponding label values, and thus can be used to judge whether the first initial prediction results can be used to determine the sequencing quality.

[0164] In this invention, the evaluation of the first initial prediction result can be achieved using the following formula:

[0165]

[0166] Where MAE1 represents the absolute mean error of the first initial regression model, m1 represents the number of first initial predictions, and y 1i This represents the label value corresponding to the i-th initial prediction result. This represents the i-th initial prediction result.

[0167] It is understandable that by determining the absolute mean error of the first initial regression model, it is possible to determine whether the output of the first initial regression model will deviate from the corresponding tag value due to different sequencing conditions under multiple tests and training, and thus it can be used to judge whether the first initial prediction result can be used to determine the sequencing quality.

[0168] In this invention, the second desired range can have a maximum value and a minimum value. The minimum value of the second desired range (unit: individuals / square micrometer) can range from [2, 5]. The maximum value of the second desired range (unit: individuals / square micrometer) can range from [5, 9].

[0169] In this way, sequencing quality can be easily determined, that is, whether the sequencing results are reliable.

[0170] Specifically, the selection of the end value of the second expected range can be adjusted according to different actual situations. In some cases, if the density requirement standard is relatively low, the second expected range can be selected as [3,8], with the unit being cells / square micrometer. If the density requirement standard is relatively high, the second expected range can be selected as [4,6], with the unit being cells / square micrometer.

[0171] In this invention, the second type of feature may include at least one of the following:

[0172] In the E to B rounds of sequencing reactions, the signals generated by each round of sequencing reactions correspond to the number of insert fragments of the at least two types of test samples in the first channel and the number of insert fragments of the at least two types of test samples in the second channel;

[0173] In sequencing rounds E to B, the number of signals generated in the first channel and the number of signals generated in the second channel in each sequencing round;

[0174] In the E to B rounds of sequencing reactions, the number of bases identified is obtained by base identification of the signals entering the first channel generated by each round of sequencing reaction and / or the number of bases identified is obtained by base identification of the signals entering the second channel generated by each round of sequencing reaction.

[0175] In this way, it becomes clear how to obtain the second type of features.

[0176] In one embodiment, E can be 4 and B can be 6. In another embodiment, E can be 4 and B can be 8. In some cases, a larger value for B indicates that more rounds of sequencing reactions are needed to predict the final sequencing results, thereby improving the accuracy of the prediction. However, a larger value for B means more rounds of sequencing reactions are required, and the time required is also longer. Therefore, the value of B should not be too large. It is necessary to consider obtaining accurate sequencing results while minimizing the number of sequencing reactions.

[0177] Please refer to Figure 11 and Figure 12 In this invention, step 02 (obtaining the sequencing characteristics of the sample to be tested in the current sequencing run) may include:

[0178] 014: Input the second training feature into the second initial regression model to obtain the second initial prediction result. The second training feature includes at least one second-class feature obtained from the training samples.

[0179] 015: The second initial prediction result is evaluated by indicators to obtain the second evaluation result, and the model parameters corresponding to the second training feature in the second initial regression model are adjusted according to the second evaluation result for retraining.

[0180] The method for determining sequencing quality according to the present invention can be implemented using the system 100 for determining sequencing quality according to the present invention. Specifically, please refer to... Figure 2 The terminal device 10 may include a training unit 13. The training unit 13 may be used to: input a second training feature into a second initial regression model to obtain a second initial prediction result, wherein the second training feature includes at least one second type of feature obtained by sequencing the training samples; perform index evaluation processing on the second initial prediction result to obtain a second evaluation result, and adjust the model parameters corresponding to the second training feature in the second initial regression model according to the second evaluation result for retraining.

[0181] In this way, the sequencing features can be automatically identified and detected through the learning model.

[0182] Specifically, the second initial regression model can be a regression model, in which the second training feature is used as the independent variable and the density of the training sample in the specified fluid channel is used as the dependent variable, thereby training the model. The training sample can be a sample with a measured corresponding sequence, that is, the sequence of the training sample is known. The actual density of the training sample in the specified fluid channel can be predetermined. The second training feature may include one or more second-class features, and the second initial prediction result is the density of the training sample in the specified fluid channel.

[0183] In some cases, evaluating the second initial prediction results can involve comparing the output of the second initial regression model with the actual density, and then determining the second evaluation result based on the degree of similarity between the two. If the two are sufficiently close, it can be determined whether the training of the second initial regression model has achieved the expected training effect. If there is still a certain deviation, it can be determined that the expected training effect has not been achieved, and the second initial regression model needs to be trained again.

[0184] In some cases, the second initial regression model can be:

[0185] y2=β 20 +β 21 X 21 ++β 22 X 22 +...+β 2n X 2n ;

[0186] Where, β 2i (i = 1, ..., n) are the model parameters in the second initial regression model, X2i (i = 1, ..., n) represent the second training features input into the second initial regression model. The above model can be understood as follows: in the second initial regression model, corresponding weights, β, are assigned to the input second training features. 20 It can be considered an independent factor that is not affected by the second type of feature during the sequencing process.

[0187] Depending on the specific testing scenario, one type of second-class feature from the training samples can be selected and substituted into one of the second-training features, or multiple types of second-class features from the training samples can be selected and substituted into their respective second-training features. When it is necessary to adjust the model parameters in the second initial regression model, the weights of one of the second-training features can be adjusted, or the weights of multiple second-training features can be adjusted.

[0188] Please refer to Figure 13 and Figure 14 In this invention, step 015 (performing index evaluation processing on the second initial prediction result to obtain a second evaluation result, and adjusting the model parameters corresponding to the second training feature in the second initial regression model based on the second evaluation result for retraining) may include:

[0189] 016: If the goodness of fit of the second initial regression model meets the third evaluation condition based on the second evaluation result, and / or the absolute mean error of the second initial regression model meets the fourth evaluation condition, then the second initial regression model obtained by the current training can be used as the second sub-model.

[0190] The method for determining sequencing quality according to the present invention can be implemented using the system 100 for determining sequencing quality according to the present invention. Specifically, please refer to... Figure 2 Training unit 13 can be used to determine that the currently trained second initial regression model can be used as the second sub-model if the goodness of fit of the second initial regression model meets the third evaluation condition based on the second evaluation result, and / or the absolute mean error of the second initial regression model meets the fourth evaluation condition.

[0191] In this way, it can be determined whether the second initial regression model has completed training.

[0192] Specifically, please combine Figure 15 and Figure 16 ,exist Figure 15 and Figure 16In the diagram, the horizontal axis represents the sequencing data number obtained from sequencing the two training samples. For example, the first sequencing data obtained from the first sequencing run on the two training samples is numbered 1, and so on. Each sequencing data includes multiple reads or sequencing sequences. The vertical axis represents the density of the two training samples within a specified fluid channel. The dots connected by a line segment represent the actual density of the two training samples within the specified fluid channel, and the triangles connected by a line segment represent the density of the training samples predicted by the second initial regression model. The actual density can be used as the label value for the second initial prediction result with the same number.

[0193] In some cases, the third evaluation criterion can be a goodness-of-fit of the second initial regression model greater than 0.8, and the fourth evaluation criterion can be an absolute mean error of the second initial regression model less than 0.1. For example, please combine this with... Figure 15 , Figure 15 This corresponds to the scenario where the second training feature is obtained through sequencing responses from the fourth to sixth rounds of the training samples. In this case, based on the output of the second initial regression model, the goodness of fit is determined to be 0.985, and the absolute mean error is 0.079. Therefore, it can be determined that the current second initial regression model has a high predictive performance in matching the test samples. The closer the goodness of fit of the second initial regression model is to 1, the better the fit to the actual situation. Similarly, the closer the absolute mean error of the second initial regression model is to 0, the better the fit to the actual situation.

[0194] Furthermore, in some cases, increasing the value of B (or using the second type of features corresponding to more rounds of sequencing responses as the second training features) can improve the fit between the output of the second initial regression model and the label values. For example, please combine this with... Figure 16 , Figure 16 This corresponds to the case where the second training feature is obtained through the sequencing responses of the training samples from the fourth to the eighth round. In this case, the goodness of fit is determined to be 0.992 and the absolute mean error is 0.048 based on the output of the second initial regression model. Compared with the case where the second training feature is obtained through the sequencing responses of the training samples from the fourth to the sixth round, the fitting effect can be further improved.

[0195] In this invention, the evaluation of the second initial prediction result can be achieved through the following formula:

[0196]

[0197] Among them, R2 2 This indicates the goodness of fit of the second initial regression model, where m² represents the number of second initial predictions, and y 2i This represents the label value corresponding to the i-th second initial prediction result. Let y1 represent the i-th second initial prediction result, and y2′ represent the average label value corresponding to the second initial prediction result.

[0198] It is understandable that by determining the goodness of fit of the second initial regression model, it can be determined whether the results output by the second initial regression model are sufficiently fitted to the corresponding label values, and thus can be used to judge whether the second initial prediction results can be used to determine the sequencing quality.

[0199] In this invention, the evaluation of the second initial prediction result can be achieved through the following formula:

[0200]

[0201] Where MAE2 represents the absolute mean error of the second initial regression model, m2 represents the number of second initial prediction results, and y 2i This represents the label value corresponding to the i-th second initial prediction result. This represents the i-th second initial prediction result.

[0202] It is understandable that by determining the absolute mean error of the second initial regression model, it is possible to determine whether the output of the second initial regression model will deviate from the corresponding tag value due to different sequencing conditions under multiple tests and training, and thus it can be used to judge whether the second initial prediction result can be used to determine the sequencing quality.

[0203] Please refer to Figure 17 A system 100 for determining sequencing quality according to the present invention may include a memory 101 and a processor 102. The memory 101 may store a computer program. When the processor 102 executes the computer program, it can implement the steps of the method for determining sequencing quality according to the present invention.

[0204] For example, when a computer program is executed by processor 102, methods for determining sequencing quality can be implemented, including:

[0205] 02: Obtain the sequencing features of at least two mixed test samples in the specified fluid channel during the current sequencing run. The sequencing features can be obtained from the images collected in the sequencing reactions of rounds A to B during the current sequencing run, where A is a natural number greater than or equal to 1 and B is a natural number greater than A.

[0206] 03: Input the sequencing features into the pre-trained prediction model, and determine the sequencing quality of the current sequencing run based on the prediction results output by the prediction model. The prediction model is a regression model.

[0207] In the aforementioned system 100 for determining sequencing quality, the sequencing characteristics of the current sequencing run are input into the prediction model. Based on the output of the prediction model, the sequencing quality of the current sequencing run is predicted, and the reliability of the sequencing results is estimated in advance. This allows for a determination of whether further resources need to be invested to complete subsequent sequencing reactions, reducing the waste of reagents and manpower, and improving sequencing efficiency.

[0208] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor 102, implements the method for determining sequencing quality according to any of the above embodiments.

[0209] For example, when a computer program is executed by processor 102, methods for determining sequencing quality can be implemented, including:

[0210] 02: Obtain the sequencing features of at least two mixed test samples in the specified fluid channel during the current sequencing run. The sequencing features can be obtained from the images collected in the sequencing reactions of rounds A to B during the current sequencing run, where A is a natural number greater than or equal to 1 and B is a natural number greater than A.

[0211] 03: Input the sequencing features into the pre-trained prediction model, and determine the sequencing quality of the current sequencing run based on the prediction results output by the prediction model. The prediction model is a regression model.

[0212] In the aforementioned computer-readable storage medium, the sequencing characteristics of the current sequencing run are input into the prediction model. Based on the output of the prediction model, the sequencing quality of the current sequencing run is predicted, and the reliability of the sequencing results is estimated in advance. This allows for a determination of whether further resources need to be invested to complete subsequent sequencing reactions, reducing the waste of reagents and manpower, and improving sequencing efficiency.

[0213] It is understood that computer-readable storage media can include: any entity or device capable of carrying a computer program, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media, etc. Computer programs can include computer program code. Computer program code can be in the form of source code, object code, executable files, or certain intermediate forms, etc.

[0214] In some embodiments of the present invention, the prediction unit 12 may be a microcontroller chip integrating a processor, memory, communication module, etc. The processor may be a central processing unit (CPU), a graphics processing unit (GPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0215] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0216] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a system including a processing module or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0217] Although embodiments of the present invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to the embodiments of the present invention without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method of determining sequencing quality, the method comprising: The method comprises: obtaining sequencing characteristics of at least two mixed test samples in a specified fluid channel in a current sequencing run, the sequencing characteristics being capable of being obtained according to images collected in A to B rounds of sequencing reactions in the current sequencing run, wherein the at least two test samples respectively contain different sample identification codes, A is a natural number greater than or equal to 1, and B is a natural number greater than A; inputting the sequencing characteristics into a pre-trained prediction model, and determining the sequencing quality of the current sequencing run according to a prediction result output by the prediction model, the prediction model being a regression model.

2. The method of claim 1, wherein the prediction result comprises a ratio of sequencing fluxes of the at least two test samples in the specified fluid channel obtained in the current sequencing run; optionally, the prediction model comprises a first sub-model, and the sequencing characteristics comprise first-type characteristics collected by sequencing the sample identification codes of the at least two test samples in A to C rounds of sequencing reactions; wherein C is the minimum number of sequencing rounds required to complete sequencing of the sample identification codes, and C is a natural number greater than A and less than B; the inputting the sequencing characteristics into the pre-trained prediction model and determining the sequencing quality of the current sequencing run according to the prediction result output by the prediction model comprises: inputting the first-type characteristics into the first sub-model to output the ratio of the sequencing fluxes; in a case where the ratio of the sequencing fluxes is within a first expected range, determining that the sequencing quality of the current sequencing run meets a preset requirement; in a case where the ratio of the sequencing fluxes exceeds the first expected range, determining that the sequencing quality of the current sequencing run does not meet the preset requirement.

3. The method of claim 1, wherein the prediction result comprises a density of the at least two test samples in the specified fluid channel; optionally, the prediction model comprises a second sub-model, and the sequencing characteristics comprise second-type characteristics collected by sequencing inserts of the at least two test samples in E to B rounds of sequencing reactions; wherein E is a natural number greater than A and less than B; the inputting the sequencing characteristics into the pre-trained prediction model and determining the sequencing quality of the current sequencing run according to the prediction result output by the prediction model comprises: inputting the second-type characteristics into the second sub-model to output the density of the at least two test samples in the specified fluid channel; in a case where the density is within a second expected range, determining that the sequencing quality of the current sequencing run meets a preset requirement; in a case where the density exceeds the second expected range, determining that the sequencing quality of the current sequencing run does not meet the preset requirement; optionally, A is 1, and E is a natural number greater than 1 and less than B.

4. The method of claim 1, wherein The prediction model comprises a first sub-model and a second sub-model, and the sequencing features comprise first-type features collected by sequencing sample identification codes of the at least two to-be-tested samples in A to C rounds of sequencing reactions and second-type features collected by sequencing insert fragments of the at least two to-be-tested samples in E to B rounds of sequencing reactions; wherein C is the minimum number of sequencing rounds required to complete sequencing of the sample identification codes, C is a natural number greater than A and less than E, and E is a natural number greater than C and less than B; The prediction result comprises a ratio of sequencing fluxes of the at least two to-be-tested samples in the specified fluid channel obtained in the current sequencing run and a density of the at least two to-be-tested samples in the specified fluid channel; The step of inputting the sequencing features into a pre-trained prediction model and determining sequencing quality of the current sequencing run according to a prediction result output by the prediction model comprises: inputting the first-type features into the first sub-model to output a ratio of sequencing fluxes between the at least two to-be-tested samples and inputting the second-type features into the second sub-model to output a density of the at least two to-be-tested samples in the specified fluid channel; in a case where the ratio of sequencing fluxes is within a first expected range and the density is within a second expected range, determining that the sequencing quality of the current sequencing run meets a preset requirement; in a case where the ratio of sequencing fluxes exceeds the first expected range and / or the density exceeds the second expected range, determining that the sequencing quality of the current sequencing run does not meet the preset requirement; Optionally, A is 1. Optionally, the first expected range has a maximum end value and a minimum end value. In a case where the at least two to-be-tested samples only comprise two to-be-tested samples, the minimum end value has a value range of [0.8, 1] and / or the maximum end value has a value range of [1, 1.25]. Optionally, the image is obtained by collecting signals generated by a sequencing reaction of the current sequencing run through a multi-channel microscopic imaging system, and the multi-channel microscopic imaging system comprises a first channel and a second channel. The first-type features comprise at least one of: a quantity ratio of total numbers of sequences belonging to different to-be-tested samples in the A to C rounds of sequencing reactions; a ratio of a quantity of signals generated by each round of sequencing reaction in the first channel to a quantity of signals generated by each round of sequencing reaction in the second channel in the A to C rounds of sequencing reactions; a ratio of a sum of quantities of signals generated by all rounds of sequencing reaction in the first channel to a sum of quantities of signals generated by all rounds of sequencing reaction in the second channel in the A to C rounds of sequencing reactions; a ratio of a sum of quantities of signals generated by all rounds of sequencing reaction in the first channel to a sum of quantities of signals generated by all rounds of sequencing reaction in the second channel in D to C rounds of sequencing reactions, wherein D is a natural number greater than A and less than C; Optionally, the quantity ratio of total numbers of sequences belonging to different to-be-tested samples is a quantity ratio of total numbers of sequences containing A bases and T bases belonging to different to-be-tested samples or a quantity ratio of total numbers of sequences containing C bases and G bases belonging to different to-be-tested samples. Optionally, the quantity ratio of the total number of sequences containing A bases and T bases belonging to different to-be-tested samples can be calculated by the following formula: (A1T0+A2T0+A3T0+A2T1+A3T1+A3T2) / (A0T1+A0T2+A0T3+A1T2+A1T3+A2T3), wherein, "A1T0+A2T0+A3T0+A2T1+A3T1+A3T2" represents the total number of sequences belonging to the to-be-tested sample with sample identification code AAA; "A0T1+A0T2+A0T3+A1T2+A1T3+A2T3" represents the total number of sequences belonging to the to-be-tested sample with sample identification code TTT; "A1T0" represents the number of sequences containing 1 A base and 0 T base; "A2T0" represents the number of sequences containing 2 A bases and 0 T bases; "A3T0" represents the number of sequences containing 3 A bases and 0 T bases; "A2T1" represents the number of sequences containing 2 A bases and 1 T base; "A3T1" represents the number of sequences containing 3 A bases and 1 T base; "A3T2" represents the number of sequences containing 3 A bases and 2 T bases; "A0T1" represents the number of sequences containing 0 A bases and 1 T base; "A0T2" represents the number of sequences containing 0 A bases and 2 T bases; "A0T3" represents the number of sequences containing 0 A bases and 3 T bases; "A1T2" represents the number of sequences containing 1 A base and 2 T bases; "A1T3" represents the number of sequences containing 1 A base and 3 T bases; "A2T3" represents the number of sequences containing 2 A bases and 3 T bases; Optionally, the step of obtaining the sequencing characteristics of at least two mixed to-be-tested samples in the current sequencing run in the specified fluid channel comprises: inputting the first training characteristics into the first initial regression model to obtain the first initial prediction result, wherein the first training characteristics comprise at least one first type of characteristics obtained by sequencing the training samples; performing index evaluation processing on the first initial prediction result to obtain a first evaluation result, and adjusting the model parameters corresponding to the first training characteristics in the first initial regression model according to the first evaluation result to retrain; Optionally, the step of performing index evaluation processing on the first initial prediction result to obtain a first evaluation result, and adjusting the model parameters corresponding to the first training characteristics in the first initial regression model according to the first evaluation result to retrain comprises: determining that the first initial regression model obtained by the current training can be used as the first sub-model when it is determined that the goodness of fit of the first initial regression model meets the first evaluation condition according to the first evaluation result, and / or the absolute mean error of the first initial regression model meets the second evaluation condition; Optionally, the index evaluation processing on the first initial prediction result can be realized by the following formula: wherein R1 2 represents a goodness of fit of the first initial regression model, m1 represents a number of the first initial prediction results, represents a label value corresponding to the i-th first initial prediction result, represents the i-th first initial prediction result, y1' represents an average value of label values corresponding to the first initial prediction results; Optionally, the index evaluation processing on the first initial prediction result can be realized by the following formula: Wherein, MAE1 represents the absolute mean error of the first initial regression model, m1 represents the number of the first initial prediction results, represents the label value corresponding to the i-th first initial prediction result, represents the i-th first initial prediction result; Optionally, the second expected range has a maximum end value and a minimum end value; The minimum end value ranges from 2 to 5, and the maximum end value ranges from 5 to 9, both in units of pieces per square micrometer. Optionally, the image is acquired by a multi-channel microscopic imaging system, and the multi-channel microscopic imaging system includes a first channel and a second channel. The second type of feature includes at least one of the following: In the E to B rounds of sequencing reactions, the number of inserts corresponding to the at least two samples in the first channel and the number of inserts corresponding to the at least two samples in the second channel; In the E to B rounds of sequencing reactions, the number of signals in the first channel and the number of signals in the second channel; In the E to B rounds of sequencing reactions, the number of recognized bases obtained by base recognition of signals entering the first channel and the number of recognized bases obtained by base recognition of signals entering the second channel; Optionally, the step of obtaining the sequencing characteristics of at least two mixed samples in the current sequencing run in the specified fluid channel includes: Inputting the second training feature into the second initial regression model to obtain a second initial prediction result, wherein the second training feature includes at least one second type of feature obtained by sequencing the training sample; Performing index evaluation processing on the second initial prediction result to obtain a second evaluation result, and adjusting the model parameters corresponding to the second training feature in the second initial regression model according to the second evaluation result to retrain; Optionally, the step of performing index evaluation processing on the second initial prediction result to obtain a second evaluation result, and adjusting the model parameters corresponding to the second training feature in the second initial regression model according to the second evaluation result to retrain includes: If the goodness of fit of the second initial regression model meets a third evaluation condition and / or the absolute mean error of the second initial regression model meets a fourth evaluation condition according to the second evaluation result, it is determined that the second initial regression model obtained by the current training can be used as the second sub-model; Optionally, the index evaluation processing of the second initial prediction result can be realized by the following formula: wherein R2 2 represents a goodness of fit of the second initial regression model, m2 represents a number of the second initial prediction results, y 2i represents a label value corresponding to the i-th second initial prediction result, represents the i-th second initial prediction result, y2' represents an average value of label values corresponding to the second initial prediction results; Optionally, the index evaluation processing of the second initial prediction result can be realized by the following formula: wherein MAE2 represents an absolute mean error of the second initial regression model, m2 represents a number of the second initial prediction results, represents a label value corresponding to the i-th second initial prediction result, represents the i-th second initial prediction result.

5. A system for determining sequencing quality, comprising: The system includes a terminal device, and the terminal device includes an acquisition unit and a prediction unit; The acquisition unit is configured to obtain sequencing characteristics of at least two mixed samples in a current sequencing run in a specified fluid channel, wherein the sequencing characteristics can be obtained from images acquired in the first A to B rounds of sequencing reactions in the current sequencing run, where A is a natural number greater than or equal to 1, and B is a natural number greater than A. The prediction unit is configured to input the sequencing features into a pre-trained prediction model, and determine the sequencing quality of the current sequencing run according to a prediction result output by the prediction model, wherein the prediction model is a regression model.

6. The system of claim 5, wherein, the prediction result comprises a ratio of sequencing fluxes of at least two samples to be tested in the specified fluid channel obtained in the current sequencing run; optionally, the prediction model comprises a first sub-model, and the sequencing features comprise first type features of sample identification codes of the at least two samples to be tested collected by sequencing in A to C sequencing reactions, wherein C is a natural number greater than A and less than B; the prediction unit is configured to: input the first type features into the first sub-model to output the ratio of sequencing fluxes; determine that the sequencing quality of the current sequencing run meets a preset requirement when the ratio of sequencing fluxes is within a first expected range; determine that the sequencing quality of the current sequencing run does not meet the preset requirement when the ratio of sequencing fluxes is beyond the first expected range.

7. The system of claim 5, wherein, the prediction result comprises a density of the at least two samples to be tested in the specified fluid channel; optionally, the prediction model comprises a second sub-model, and the sequencing features comprise second type features of insert fragments of the at least two samples to be tested collected by sequencing in E to B sequencing reactions, wherein E is a natural number greater than A and less than B; the prediction unit is configured to: input the second type features into the second sub-model to output the density of the at least two samples to be tested in the specified fluid channel; determine that the sequencing quality of the current sequencing run meets a preset requirement when the density is within a second expected range; determine that the sequencing quality of the current sequencing run does not meet the preset requirement when the density is beyond the second expected range; and optionally, A is 1, and E is a natural number greater than 1 and less than B.

8. The system of claim 5, wherein, the prediction model comprises a first sub-model and a second sub-model, and the sequencing features comprise first type features of sample identification codes of the at least two samples to be tested collected by sequencing in A to C sequencing reactions and second type features of insert fragments of the at least two samples to be tested collected by sequencing in E to B sequencing reactions; the prediction result comprises a ratio of sequencing fluxes of at least two samples to be tested in the specified fluid channel obtained in the current sequencing run and a density of the at least two samples to be tested in the specified fluid channel; wherein C is a minimum number of sequencing rounds required to complete sequencing of the sample identification codes, C is a natural number greater than A and less than E, and E is a natural number greater than C and less than B; the prediction unit is configured to: inputting the first type of features collected from the sequencing reactions in the A to C rounds into the first sub-model to output a sequencing flux ratio between the at least two samples to be tested, and inputting the second type of features collected from the sequencing reactions in the E to B rounds into the second sub-model to output a density of the at least two samples to be tested in the specified fluid channel; determining that the sequencing quality of the current sequencing run meets a preset requirement when the sequencing flux ratio is within a first expected range and the density is within a second expected range; determining that the sequencing quality of the current sequencing run does not meet the preset requirement when the sequencing flux ratio is beyond the first expected range and / or the density is beyond the second expected range; Optionally, A is 1. Optionally, the first expected range has a maximum end value and a minimum end value. When the at least two samples to be tested include only two samples to be tested, the minimum end value has a value range of [0.8, 1] and / or the maximum end value has a value range of [1, 1.25]. Optionally, the image is acquired by a multi-channel microscopic imaging system from signals generated by the sequencing reactions of the current sequencing run, the multi-channel microscopic imaging system including a first channel and a second channel; and the first type of features include at least one of the following: a quantity ratio of total numbers of sequences belonging to different samples to be tested determined from the image in the sequencing reactions in the A to C rounds; a ratio of a number of signals generated by each of the sequencing reactions in the A to C rounds in the first channel to a number of signals generated by each of the sequencing reactions in the A to C rounds in the second channel; a ratio of a sum of numbers of signals generated by all of the sequencing reactions in the A to C rounds in the first channel to a sum of numbers of signals generated by all of the sequencing reactions in the A to C rounds in the second channel; a ratio of a sum of numbers of signals generated by all of the sequencing reactions in the D to C rounds in the first channel to a sum of numbers of signals generated by all of the sequencing reactions in the D to C rounds in the second channel, wherein D is a natural number greater than A and smaller than C; Optionally, the quantity ratio of total numbers of sequences belonging to different samples to be tested determined from the image is a quantity ratio of total numbers of sequences containing A bases and T bases belonging to different samples to be tested determined from the image, or is a quantity ratio of total numbers of sequences containing C bases and G bases belonging to different samples to be tested determined from the image; and optionally, the quantity ratio of total numbers of sequences containing A bases and T bases belonging to different samples to be tested determined from the image can be calculated by the following formula: (A1T0+A2T0+A3T0+A2T1+A3T1+A3T2) / (A0T1+A0T2+A0T3+A1T2+A1T3+A2T3), wherein "A1T0+A2T0+A3T0+A2T1+A3T1+A3T2" represents total numbers of sequences belonging to a sample to be tested with a sample identification code of AAA; and "A0T1+A0T2+A0T3+A1T2+A1T3+A2T3" represents total numbers of sequences belonging to a sample to be tested with a sample identification code of TTT. "A1T0" represents a sequence number containing 1 A base and 0 T base; "A2T0" represents a sequence number containing 2 A bases and 0 T bases; "A3T0" represents a sequence number containing 3 A bases and 0 T bases; "A2T1" represents a sequence number containing 2 A bases and 1 T base; "A3T1" represents a sequence number containing 3 A bases and 1 T base; "A3T2" represents a sequence number containing 3 A bases and 2 T bases; "A0T1" represents a sequence number containing 0 A bases and 1 T base; "A0T2" represents a sequence number containing 0 A bases and 2 T bases; "A0T3" represents a sequence number containing 0 A bases and 3 T bases; "A1T2" represents a sequence number containing 1 A base and 2 T bases; "A1T3" represents a sequence number containing 1 A base and 3 T bases; "A2T3" represents a sequence number containing 2 A bases and 3 T bases; Optionally, the terminal device comprises a training unit, which is configured to: input a first training feature into a first initial regression model to obtain a first initial prediction result, the first training feature comprising at least one first type of feature obtained by sequencing a training sample; perform index evaluation processing on the first initial prediction result to obtain a first evaluation result, and adjust the model parameters corresponding to the first training feature in the first initial regression model according to the first evaluation result to retrain; Optionally, the training unit is further configured to: determine that the first initial regression model obtained by the current training can be used as the first sub-model when it is determined that the goodness of fit of the first initial regression model meets a first evaluation condition according to the first evaluation result, and / or the absolute mean error of the first initial regression model meets a second evaluation condition; Optionally, the index evaluation processing on the first initial prediction result can be realized by the following formula: wherein R1 2 represents a goodness of fit of the first initial regression model, m1 represents a number of the first initial prediction results, represents a label value corresponding to the i-th first initial prediction result, represents the i-th first initial prediction result, y1' represents an average value of label values corresponding to the first initial prediction results; Optionally, the index evaluation processing on the first initial prediction result can be realized by the following formula: Wherein, MAE1 represents the absolute mean error of the first initial regression model, m1 represents the number of the first initial prediction results, represents the label value corresponding to the i-th first initial prediction result, represents the i-th first initial prediction result; Optionally, the second expected range has a maximum end value and a minimum end value; the value range of the minimum end value is [2, 5], and the unit is: pieces per square micrometer, and / or the value range of the maximum end value is [5, 9], and the unit is: pieces per square micrometer; Optionally, the image is obtained by a multi-channel microscopic imaging system collecting signals generated by the sequencing reaction of the current sequencing operation, and the multi-channel microscopic imaging system comprises a first channel and a second channel; The second type of feature comprises at least one of the following: In the E to B rounds of sequencing reactions, the number of insert fragments corresponding to the at least two samples in the first channel and the number of insert fragments corresponding to the at least two samples in the second channel generated by each round of sequencing reaction; In the E to B rounds of sequencing reactions, the number of signals in the first channel and the number of signals in the second channel generated by each round of sequencing reaction; In the E to Bth sequencing reaction, the number of recognized bases obtained by base recognition of the signal into the first channel generated by each round of sequencing reaction and / or the number of recognized bases obtained by base recognition of the signal into the second channel generated by each round of sequencing reaction; Optionally, the terminal device comprises a training unit, which is configured to: input the second training features into the second initial regression model to obtain a second initial prediction result, the second training features comprising at least one second type of features obtained by sequencing the training samples; perform index evaluation processing on the second initial prediction result to obtain a second evaluation result, and adjust the model parameters corresponding to the second training features in the second initial regression model according to the second evaluation result to retrain; Optionally, the training unit is further configured to: determine that the second initial regression model obtained by the current training can be used as the second sub-model when it is determined according to the second evaluation result that the goodness of fit of the second initial regression model meets a third evaluation condition and / or the absolute mean error of the second initial regression model meets a fourth evaluation condition; Optionally, the index evaluation processing on the second initial prediction result can be implemented by the following formula: wherein R2 2 represents a goodness of fit of the second initial regression model, m2 represents a number of the second initial prediction results, y 2i represents a label value corresponding to the i-th second initial prediction result, represents the i-th second initial prediction result, y2' represents an average value of label values corresponding to the second initial prediction results; Optionally, the index evaluation processing on the second initial prediction result can be implemented by the following formula: wherein MAE2 represents an absolute mean error of the second initial regression model, m2 represents a number of the second initial prediction results, represents a label value corresponding to the i-th second initial prediction result, represents the i-th second initial prediction result.

9. A computer system, characterized by The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 4.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 4.