Quality parameter determination method and device for base identification

By calculating the first parameter in the recognition result data of the base recognition model and filtering it with the preset Q0 value, the target quality parameters are obtained, which solves the problem of difficulty in evaluating the quality of machine learning base recognition results in the prior art, and improves the accuracy of the evaluation.

CN120199333APending Publication Date: 2025-06-24GENEMIND BIOSCIENCES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311787490.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately evaluate the quality parameters of base recognition results based on machine learning, making it difficult to effectively evaluate the credibility and error rate of base recognition results.

Method used

By obtaining the recognition result data of the base recognition model for the nucleic acid sample to be identified, the first parameter is calculated based on the base probability distribution, and the first parameter is filtered based on the preset Q0 value to obtain the target quality parameter.

Benefits of technology

A more accurate quality parameter evaluation of machine learning base recognition results is achieved, and the credibility and error rate evaluation accuracy of base recognition results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199333A_ABST
    Figure CN120199333A_ABST
Patent Text Reader

Abstract

The invention discloses a quality parameter determination method and device for base recognition, and the method comprises the steps: obtaining recognition result data of a base recognition model for a to-be-recognized nucleic acid sample, the recognition result data comprising base probability distribution for the current base extension reaction of the to-be-recognized nucleic acid sample; based on the basic group probability distribution, calculating to obtain a first parameter; and filtering the first parameter based on a preset Q0 value to obtain a target quality parameter. According to the method and the device, the target quality parameter is determined by using the information output by the base recognition model, so that the target quality parameter can more accurately judge the base recognition result of machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and more specifically, to a method and device for determining quality parameters for base recognition. Background Art

[0002] There are various types of data in a gene sequencing result file, among which there are at least two important types of data: basecall result data and quality score data. The quality score is usually represented by a Q value. The significance of the Q value lies in scoring each recognized or output base to determine the credibility of the base. It can be seen that the application accuracy of the quality score Q value will affect the final base recognition result. Summary of the Invention

[0003] In view of this, this application provides the following technical solutions:

[0004] A method for determining quality parameters for base recognition, comprising:

[0005] Obtaining recognition result data of a base recognition model for a nucleic acid sample to be recognized, where the recognition result data includes the base probability distribution of the current base extension reaction for the nucleic acid sample to be recognized;

[0006] Calculating a first parameter based on the base probability distribution;

[0007] Filtering the first parameter based on a preset Q0 value to obtain a target quality parameter, where Q0 = -10×log 10 e, where e is the base recognition error rate.

[0008] Optionally, the calculating a first parameter based on the base probability distribution includes:

[0009] Obtaining the maximum probability parameter in the current base extension reaction based on the base probability distribution;

[0010] Calculating a first parameter based on the maximum probability parameter.

[0011] Optionally, the calculating a first parameter based on the maximum probability parameter includes:

[0012] Determining a second parameter Q1 based on the formula Q1 = -10×log 10 (1 - P max ), where P max represents the maximum probability parameter;

[0013] Comparing the second parameter with a preset value, and if the second parameter is greater than the preset value, rounding the second parameter to zero to obtain a first parameter;

[0014] If the second parameter is not greater than the preset value, round the second parameter to an integer to obtain a first parameter; wherein, the value range of the first parameter is a positive integer from 0 to the preset value.

[0015] Optionally, it further includes:

[0016] Determine the base recognition error rate corresponding to each of the first parameters;

[0017] Based on the error rate of base recognition corresponding to each of the first parameters, statistically obtain the first base recognition error rate corresponding to the first parameters greater than a preset first threshold, and statistically obtain the second base recognition error rate corresponding to the first parameters less than a preset second threshold;

[0018] If the first base recognition error rate and / or the second base recognition error rate is greater than the error rate threshold, optimize the first parameter based on the error rate of base recognition corresponding to each of the first parameters to obtain a target first parameter;

[0019] Among them, based on a preset Q0 value, filtering the first parameter to obtain a target quality parameter, including:

[0020] Based on the preset Q0 value, process the target first parameter to obtain a target quality parameter.

[0021] Optionally, the optimizing the first parameter based on the error rate of base recognition corresponding to each of the first parameters to obtain a target first parameter includes:

[0022] Based on the formula Q2 = -10×log 10 (2×P ID ) to update the first parameter, where Q2 represents the target first parameter corresponding to the current first parameter, and P ID represents the error rate corresponding to the set formed by the current first parameter.

[0023] Optionally, the filtering the first parameter based on the preset Q0 value to obtain a target quality parameter includes:

[0024] Based on the preset Q0 value, determine the segmented data corresponding to the preset Q0 value;

[0025] Based on the segmented data, filter the first parameter to obtain a target quality parameter so that the base filtering range of the target quality parameter matches the base filtering range of the Q0 value.

[0026] Optionally, it further includes:

[0027] Filter the base recognition results of the base recognition model based on the target quality parameter to obtain the target base recognition results.

[0028] Optionally, it further includes:

[0029] Obtain training sample data, where the training sample data includes optical data of the original base channel;

[0030] Use the true base recognition result corresponding to the training sample data as the training target, and train the training sample data to obtain a base recognition model.

[0031] Optionally, the calculating the first parameter based on the base probability distribution includes:

[0032] Obtain the test sample data of the base recognition model and the predicted base recognition result obtained by recognizing the test sample data based on the base recognition model, where the test sample data includes the true base recognition result;

[0033] Determine the base recognition probability screening condition based on the correspondence between the predicted base recognition result and the true base recognition result;

[0034] Based on the base recognition probability screening condition, perform probability parameter screening in the probability distribution to obtain the first parameter.

[0035] Optionally, the determining the base recognition probability screening condition based on the correspondence between the predicted base recognition result and the true base recognition result includes:

[0036] Obtain the data where the predicted base recognition result of the current base extension reaction in each test sample data is consistent with the true base recognition result, and statistically obtain the base prediction probability of the corresponding current base reaction in the data;

[0037] Determine the screening range of the base recognition probability based on the predicted probability.

[0038] A quality parameter determination device for base recognition includes:

[0039] An acquisition unit for obtaining the recognition result data of the base recognition model for the nucleic acid sample to be recognized, where the recognition result data includes the base probability distribution of the current base extension reaction for the nucleic acid sample to be recognized;

[0040] A calculation unit for calculating a first parameter based on the base probability distribution;

[0041] A processing unit for filtering the first parameter based on a preset Q0 value to obtain a target quality parameter, where Q0 = -10×log10 e, where e is the base recognition error rate.

[0042] Optionally, the calculation unit includes:

[0043] A first obtaining subunit, configured to obtain a maximum probability parameter in the current base extension reaction based on the base probability distribution;

[0044] A first calculating subunit, configured to perform calculations based on the maximum probability parameter to obtain a first parameter.

[0045] Optionally, the first calculating subunit is specifically configured to:

[0046] Based on the formula Q2' = -10×log 10 (1 - P max ), determine a second parameter Q1, where P max represents the maximum probability parameter;

[0047] Compare the second parameter with a preset value. If the second parameter is greater than the preset value, round the second parameter down to an integer to obtain the first parameter;

[0048] If the second parameter is not greater than the preset value, round the second parameter to an integer to obtain the first parameter; wherein, the value range of the first parameter is a positive integer from 0 to the preset value.

[0049] Optionally, it further includes:

[0050] A first determining unit, configured to determine the base recognition error rate corresponding to each of the first parameters;

[0051] A statistics unit, configured to statistically obtain a first base recognition error rate corresponding to the first parameter greater than a preset first threshold and a second base recognition error rate corresponding to the first parameter less than a preset second threshold based on the error rate of base recognition corresponding to each of the first parameters;

[0052] An optimization unit, configured to, if the first base recognition error rate and / or the second base recognition error rate is greater than an error rate threshold, optimize the first parameter based on the base recognition error rate corresponding to each of the first parameters to obtain a target first parameter;

[0053] Wherein, the processing unit is specifically configured to:

[0054] Perform a filtering process on the target first parameter based on the preset Q0 value to obtain a target quality parameter.

[0055] Optionally, the optimization unit is specifically configured to:

[0056] Based on the formula Q2 = -10×log 10 (2×P ID ), update the first parameter, where Q2 represents the target first parameter corresponding to the current first parameter, and P ID represents the error rate corresponding to the set formed by the current first parameter.

[0057] Optionally, the processing unit includes:

[0058] A determination subunit, configured to determine segmented data corresponding to the preset Q0 value based on the preset Q0 value;

[0059] A processing subunit, configured to perform filtering processing on the first parameter based on the segmented data to obtain a target quality parameter, so that the base filtering range of the target quality parameter matches the base filtering range of the Q0 value.

[0060] Optionally, it further includes:

[0061] A filtering unit, configured to filter the base recognition result of the base recognition model based on the target quality parameter to obtain a target base recognition result.

[0062] Optionally, it further includes: a model training unit, and the model training unit is configured to:

[0063] Obtain training sample data, where the training sample data includes optical data of the original base channel;

[0064] Use the true base recognition result corresponding to the training sample data as a training target, and train the training sample data to obtain a base recognition model.

[0065] Optionally, the calculation unit includes:

[0066] An identification subunit, configured to obtain test sample data of the base recognition model and a predicted base recognition result obtained by recognizing the test sample data based on the base recognition model, where the test sample data includes a true base recognition result;

[0067] A second determination subunit, configured to determine a base recognition probability screening condition based on the correspondence between the predicted base recognition result and the true base recognition result;

[0068] A screening subunit, configured to perform probability parameter screening in the probability distribution based on the base recognition probability screening condition to obtain a first parameter.

[0069] Optionally, the second determination subunit is specifically configured to:

[0070] Obtain data where the predicted base recognition result of the current base extension reaction in each test sample data is consistent with the true base recognition result, and statistically obtain the base prediction probability of the corresponding current base reaction in the data.

[0071] Based on the predicted probability, determine the screening range of the base recognition probability.

[0072] As can be seen from the above technical solutions, the present application discloses a method and device for determining quality parameters for base recognition, including: obtaining recognition result data of a base recognition model for a nucleic acid sample to be recognized, where the recognition result data includes the base probability distribution of the current base extension reaction for the nucleic acid sample to be recognized; calculating a first parameter based on the base probability distribution; and filtering the first parameter based on a preset Q0 value to obtain a target quality parameter. In the present application, the information output by the base recognition model is used to determine the target quality parameter, so that the target quality parameter can more accurately evaluate the base recognition result of machine learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0074] Figure 1 It is a schematic flowchart of a method for determining quality parameters for base recognition provided by an embodiment of the present application;

[0075] Figure 2 It is a schematic diagram of a curve corresponding to the error rate and Q1 provided by an embodiment of the present application;

[0076] Figure 3 It is a schematic diagram of another curve corresponding to the error rate and target Q1 provided by an embodiment of the present application;

[0077] Figure 4 It is a schematic diagram of yet another curve corresponding to the error rate and target Q1 provided by an embodiment of the present application;

[0078] Figure 5 It is a schematic diagram of a comparison curve of the alignment error rate in the filtered sequence provided by an embodiment of the present application for Q0 and Q T ;

[0079] Figure 6 It is a schematic diagram of a comparison curve of the loss in the filtered sequence provided by an embodiment of the present application for Q0 and Q T ;

[0080] Figure 7A Q0 and Q provided by an embodiment of the present application T Schematic diagram of the comparison curve of the total sequence loss of the filtered sequences of each

[0081] Figure 8 A target quality parameter Q provided by an embodiment of the present application T Schematic diagram of the curve of the relevant information with the error rate

[0082] Figure 9 Schematic diagram of a decreasing error rate curve provided by an embodiment of the present application

[0083] Figure 10 Schematic diagram of another decreasing error rate curve provided by an embodiment of the present application

[0084] Figure 11 Schematic diagram of the error rate curve of the data retained after filtering provided by an embodiment of the present application

[0085] Figure 12 Schematic diagram of another error rate curve of the retained after filtering provided by an embodiment of the present application

[0086] Figure 13 Schematic diagram of the structure of a quality parameter determination device for base recognition provided by an embodiment of the present application Detailed implementation manners

[0087] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0088] The terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may include unlisted steps or units.

[0089] In an embodiment of the present application, the term "sequencing" can also be referred to as "nucleic acid sequencing" or "gene sequencing", and the three can be used interchangeably in expression, all referring to the determination of the types and arrangement orders of bases or nucleotides (including nucleotide analogs) in a nucleic acid molecule. The so-called sequencing includes the process of binding nucleotides to a template and collecting the corresponding signals emitted by the nucleotides (including analogs). The so-called sequencing includes sequencing by synthesis (sequencing while synthesizing, SBS) and / or sequencing by ligation (sequencing while ligating, SBL), including DNA sequencing and / or RNA sequencing, including long-fragment sequencing and / or short-fragment sequencing. The so-called long fragments and short fragments are relative. For example, a nucleic acid molecule longer than 1 Kb, 2 Kb, 5 Kb or 10 Kb can be called a long fragment, and a nucleic acid molecule shorter than 1 Kb or 800 bp can be called a short fragment.

[0090] Sequencing generally includes multiple rounds to achieve the determination of the types and arrangement orders of multiple bases or nucleotides on a nucleic acid template. In the embodiments of the present application, each round of "the process of achieving the determination of the types and arrangement orders of multiple bases or nucleotides on a nucleic acid template" is called "one round of sequencing". "One round of sequencing" (cycle) is also called "sequencing round", and can be defined as a single base extension of four nucleotides / bases. In other words, "one round of sequencing" can be defined as completing the determination of the type of base or nucleotide at any specified position on the template. For a sequencing platform that realizes sequencing based on polymerization or ligation reactions, one round of sequencing includes the process of enabling four nucleotides (including nucleotide analogs) to bind to the so-called nucleic acid template through base complementarity and collecting the corresponding signals emitted. Among them, for a platform that realizes sequencing based on polymerization reactions, the reaction system includes reaction substrate nucleotides, polymerase and nucleic acid template. A sequence (sequencing primer) is bound to the nucleic acid template. Based on the principle of base pairing and the principle of polymerization reaction, the added reaction substrate nucleotides are connected to the sequencing primer under the catalysis of the polymerase to achieve the binding of the nucleotide to a specific position on the nucleic acid template. Usually, one round of sequencing can include one or multiple base extensions (repeat). For example, four nucleotides are added to the reaction system in sequence, and base extensions and the collection of corresponding reaction signals are carried out respectively. One round of sequencing includes four base extensions; for another example, four nucleotides are added to the reaction system in any combination, such as a combination of two or a combination of one and three, and the two combinations carry out base extensions and the collection of corresponding reaction signals respectively. One round of sequencing includes two base extensions; for yet another example, four nucleotides are added to the reaction system simultaneously for base extension and the collection of reaction signals. One round of sequencing includes one base extension.

[0091] In the embodiments of the present application, the term "channel" refers to four channels formed in different ways during the sequencing process, which can screen and distinguish the channels derived from the four bases A, C, G, T or U. For example, the so-called channel may refer to the fluorescence signal optical channels formed by using different excitation lights, different fluorescence filter color plates, etc. during the sequencing photographing process, which can screen and distinguish the fluorescence signals of the four fluorescence bases derived from A, C, G, T or U. During actual sequencing, photos are taken in four different fluorescence channels to form fluorescence images. Ideally, only the signals of the fluorescence base categories corresponding to the channel exist in each fluorescence channel. However, in actual situations, due to the influence of fluorescence crosstalk, in addition to the fluorescence signals of the corresponding fluorescence bases, the fluorescence signals of other bases will also appear in each channel.

[0092] There are various types of data in the gene sequencing result file, including at least two important types of data: basecall result data and quality score data. Among them, the quality score is usually represented by the Q value. The significance of the Q value lies in scoring each identified or output base to determine the credibility of the base.

[0093] Currently, when performing base identification based on fluorescence signals, the parameter Q value is characterized by the parameter obtained by statistically analyzing the optical data corresponding to the four bases in the current base extension reaction for evaluating the credibility of the base.

[0094] Taking the second-generation sequencing technology as an example, the second-generation sequencing technology utilizes the characteristics that different fluorescent molecules have different fluorescence emission wavelengths, and uses different fluorescent molecules to label the substrates of the base extension reaction. After the base extension reaction occurs, a laser is used to irradiate the fluorescent molecules to excite the fluorescent molecules to generate fluorescence signals, and an optical sensor is used to obtain the fluorescence signals of specific wavelengths. Finally, based on the fluorescence signals, the type of the base bound to the fluorescent molecule is identified. Taking dual-color sequencing as an example, four different fluorescent labels are used to label four different bases (A, T or U, G, C), and four bases are added simultaneously to complete a base extension reaction. A laser is irradiated to excite the fluorescent molecules to generate fluorescence signals, and an optical imaging system is used to collect the images of the fluorescence signals generated by different excitation wavelengths. In one method, image processing of the fluorescence image and positioning of the fluorescence point positions are performed for base cluster detection. According to the base cluster detection results of multiple fluorescence images corresponding to the sequencing signal responses of different base types, a template is constructed to construct the positions of all base cluster template points (cluster). Then, according to the template, the optical data of the filtered image is extracted (that is, mainly the extraction of fluorescence intensity), and then the fluorescence intensity is corrected. Finally, the score is calculated by identifying the base based on the maximum intensity of the positions of each base cluster template point, and a base sequence file is obtained.

[0095] In each round of sequencing, each identified base is assigned a quality value score for evaluating the accuracy of the identification result, known as the quality score (usually denoted as Q value or Qphred). The quality score can reflect the confidence and error rate (e) of base identification during the sequencing process. The calculation method of the Q value is as follows: Qphred = -10 × log 10 e. From this formula, it can be seen that the larger the Q value, the smaller the probability of identification error and the higher the confidence. The Q values commonly used in statistics correspond to different error rates. In high-throughput sequencing, Q20 can be used as the threshold for base filtering, and Q30 is also often used to evaluate the sequencing quality. Among them, Q20 means that in the base identification results of the sequencing output, the probability of correct base identification is 99%, and the probability of incorrect base identification is 1%; Q30 corresponds to the situation where in the base identification results of the sequencing output, the probability of correct base identification is 99.9%, and the probability of incorrect base identification is 0.1%.

[0096] In a traditional sequencing method based on fluorescence signal for base identification, the quality score (hereinafter referred to as Q0 value for the sake of distinguishing from the quality parameter in the embodiments of the present application) algorithm is based on the numerical values of the fluorescence intensities of the four bases in the fluorescence image to statistically calculate the error rate, and then generate the Q0 value. This algorithm has undergone strict error rate distribution statistical tests, so it is a relatively standard Q value for the bases identified by traditional basecall. Currently, with the development of base identification technology, there are more and more algorithms for base identification based on machine learning. The error rate of the bases output by the base identification algorithm based on machine learning is lower than that of the traditional base identification, which can effectively reduce the error rate and improve the base quality. However, the way this algorithm outputs bases does not completely rely on the fluorescence intensity information of the four bases in the fluorescence image. Therefore, using the original quality parameter (Q0 value) cannot well reflect the error rate information of the bases output by the base identification based on machine learning.

[0097] In view of this, the embodiments of the present application propose a method for determining the quality parameter of a base identification algorithm that can be applied to machine learning to better evaluate the accuracy of base identification based on machine learning.

[0098] See Figure 1 , which is a schematic flowchart of a method for determining the quality parameter for base identification provided by the embodiments of the present application. The method may include:

[0099] S101. Obtain the identification result data of the base identification model for the nucleic acid sample to be identified.

[0100] In the embodiments of the present application, the base recognition model outputs image data formed by one or more fluorescence images generated based on the sequencing signals of the bases participating in the base extension reaction. Among them, the fluorescence image to be measured may refer to the original fluorescence image taken on the surface of the sequencing chip where the nucleic acid sample to be recognized is located during each round of sequencing, or may also be the corrected fluorescence image after the original fluorescence image taken on the surface of the sequencing chip where the nucleic acid sample to be recognized is located during each round of sequencing is corrected. In the embodiments of the present application, the network result of the base recognition model is used to extract features from the multi-channel input image, and the base recognition result corresponding to the input image data of each channel is determined based on the extracted features, that is, the recognition result data is obtained. The base recognition result can have different presentation forms. For example, the base type corresponding to the current base extension reaction position or site can be directly output, or the base recognition result can be characterized in the form of probability data of each base being recognized. In one implementation manner, the base recognition result may include the base type or the probability data of each base being recognized output based on the signal at the position where the base extension reaction occurs in the fluorescence image. Specifically, the recognition data includes the base probability distribution for the current base extension reaction of the nucleic acid sample to be recognized. In the embodiments of the present application, the base probability distribution refers to the respective probabilities that the base types participating in the base extension reaction are recognized as the four bases A, T (or U), G, and C for a specified round of base extension reaction. It should be understood that for each current base extension reaction of the nucleic acid sample to be recognized, the sum of the respective probabilities of the four bases A, T (or U), G, and C is 1. For example, the probability data of the position where the base signal acquisition center indicating whether the base extension reaction position or site is a certain base type is located. The probability value of the position where the base signal acquisition center is located indicates the probability that the base signal belongs to the base type A, C, G, T, or U, and the sum of the probabilities of the four base types is 1.

[0101] S102. Calculate a first parameter based on the base probability distribution.

[0102] In the embodiments of the present application, the base probability distribution data is used to determine the target quality parameter. However, the base probability distribution is generally a percentage parameter. In order to more accurately generate the target quality parameter for base evaluation, the obtained base probability distribution data needs to be transformed to obtain the first parameter, that is, the first parameter is obtained by calculating the base probability distribution data in a corresponding data format. Then, based on the preset Q0 value, the first parameter is processed to obtain the target quality parameter, so that the target quality parameter can also have a base filtering function similar to the Q0 value.

[0103] Specifically, when calculating the first parameter based on the base probability distribution, it can be based on the maximum probability parameter in the current base extension reaction, or probability parameters that meet specific conditions can be selected for processing.

[0104] In one embodiment, the process of calculating the first parameter based on the base probability distribution includes: obtaining the maximum probability parameter in the current base extension reaction based on the base probability distribution; and calculating the first parameter based on the maximum probability parameter. Further, calculating the first parameter based on the maximum probability parameter includes:

[0105] Based on the formula Q1 = -10×log 10 (1 - P max ), determining the second parameter Q1, where P max represents the maximum probability parameter; comparing the second parameter with a preset value, if the second parameter is greater than the preset value, rounding down the second parameter to an integer to obtain the first parameter; if the second parameter is not greater than the preset value, rounding the second parameter to an integer to obtain the first parameter; wherein, the value range of the first parameter is a positive integer from 0 to the preset value.

[0106] In this method, the maximum value of the predicted probability corresponding to each base in the output result of the base recognition model is first separately counted and corresponding bases are paired. Thus, while obtaining a value representing the base, the value can represent its corresponding maximum predicted probability value, and the higher the probability, the higher the credibility. Since probabilities are usually expressed as percentage values, such as 75%, 60%, etc., in order to more accurately generate the target quality parameter for base evaluation, the probability value needs to be transformed, that is, transformed into a value within a preset range, namely: the probability value is transformed into a value less than or equal to the preset value. In some embodiments, the preset value is 100, that is, this probability needs to be converted into a value between 0 - 100 to obtain the target quality parameter. Specifically, first convert the maximum probability parameter into a value used for calculation when determining the target quality parameter, that is, the second parameter Q1, where Q1 = -10×log 10 (1 - P max ). Since the final target quality is an integer between 0 - 100, the preset value is set to 100. If the calculated second parameter is greater than 100, then round down the second parameter to an integer to obtain the first parameter; if the second parameter is not greater than 100, then round the second parameter to an integer to obtain the first parameter.

[0107] S103. Filter the first parameter based on a preset Q0 value to obtain the target quality parameter.

[0108] After obtaining the first parameter, which can be a numerical value ranging from 0 to 100, and using the preset Q0 value for comparison, the error rate of the values from 0 to 100 can be statistically analyzed according to the definition and segmentation criteria of the Q0 value to assign a new target quality parameter, so as to obtain the target quality parameter that can be applied to evaluate the base recognition result in the base recognition model in the embodiments of the present application. As described above, in the embodiments of the present application, Q0 = -10×log 10 e, where e is the base recognition error rate.

[0109] In one implementation, based on the preset Q0 value, the first parameter is filtered to obtain the target quality parameter, including: based on the preset Q0 value, determining the segmented data corresponding to the Q0 value; filtering the first parameter based on the segmented data to obtain the target quality parameter, so that the base filtering range of the target quality parameter matches the base filtering range of the preset Q0 value. Exemplarily, according to the preset Q0 value, the base recognition error rate e corresponding to the preset Q0 value can be obtained, and based on this base recognition error rate e, the first parameter that meets the preset Q0 value can be obtained, that is, the first parameter can also be processed according to this segmentation method to obtain the target quality parameter.

[0110] Through research, it is found that there is a certain relationship between the base misrecognition rate and the magnitude of the quality parameter. When the quality parameter is too large or too small, the error rate will be relatively high. Therefore, in order to obtain an accurate quality parameter for machine learning, the first parameter can be optimized, and then the target quality parameter can be obtained based on the optimized first parameter. In the embodiments of the present application, it further includes:

[0111] Determining the base recognition error rate corresponding to each first parameter;

[0112] Based on the base recognition error rate corresponding to each first parameter, statistically obtaining the first base recognition error rate corresponding to the first parameter greater than the preset first threshold, and statistically obtaining the second base recognition error rate corresponding to the first parameter less than the preset second threshold;

[0113] If the first base recognition error rate and / or the second base recognition error rate is greater than the error rate threshold, optimizing the first parameter based on the base recognition error rate corresponding to each first parameter to obtain the target first parameter;

[0114] Among them, based on the preset Q0 value, filtering the first parameter to obtain the target quality parameter, including:

[0115] Processing the target first parameter based on the base filtering condition of the original quality parameter to obtain the target quality parameter.

[0116] Further, optimize the first parameter based on the base recognition error rate corresponding to each first parameter to obtain the target first parameter, including: based on the formula Q2 = -10×log 10 (2×P ID ) to update the first parameter, where Q2 represents the target first parameter corresponding to the current first parameter, and P ID represents the error rate corresponding to the set formed by the current first parameter. Thus, the accuracy of the quality parameter corresponding to the first parameter can be fed back according to the error rate corresponding to the set formed by the current first parameter, and the current first parameter exceeding the error rate preset value can be optimized to improve the accuracy of the quality parameter, and further improve the accuracy of base recognition.

[0117] For example, the data used for statistics is from a set of sequencing pictures (fov) in an experiment on the human genes of a complete standard human sample (HG001) using the latest fixed version of the sequencer and reagents at that time with a machine learning model, and basecall (base recognition) is performed on these sequencing pictures, and the base prediction probabilities and the parameters used for base prediction generated during this process are printed. Complete the comparison between the basecall result file and the human gene Hg19 reference, and obtain the correct base corresponding to each predicted base from the obtained sam file. Then use the intermediate parameters printed by the machine learning basecall to assign which base each probability corresponds to and the base on the gene index (reference). We determine whether the probability is correct or not according to whether the type of the base is consistent with the reference. And write the correctly measured base as 1 and assign it to the corresponding maximum probability, and write the wrongly measured base as 0 and assign it to the corresponding maximum probability. So far, two very important pieces of information will be obtained, the maximum probability of the base and the correct and wrong marks corresponding to the base probability. For example, the value 1 indicates that the base recognition result obtained based on the base recognition model is consistent with the marked true value result, and vice versa, the value 0 is used to indicate.

[0118] In the embodiments of the present application, in order to facilitate the distinction between the original quality parameter and the target quality score, the original quality parameter can be characterized as the Q0 value, and the target quality score can be characterized as the Q T value. Correspondingly, the algorithm for determining the original quality parameter is simply referred to as the Q0 algorithm, and the algorithm for determining the target quality parameter is simply referred to as the Q T algorithm.

[0119] According to the obtained base probability distribution of the current base extension reaction of the nucleic acid sample to be recognized, use Q1 = -10×log 10 (1 - P max) Numerically convert the probability to obtain a Q1 value. Round the obtained Q1 value to get an integer Q1 value. Then, statistically analyze the error rate for all Q1 values and perform error rate simulation calculations for each Q1 value to obtain a relationship graph between the error rate and the Q1 value as shown in Figure 2 . Among them, the closer the curve is to 0, the lower the error rate, indicating better classification. According to Figure 2 , it can be found that there are still large fluctuations in the high-Q1 part and the error rate is relatively high. Therefore, it is difficult to give a relatively accurate quality score based solely on the Q1 value. Therefore, it is necessary to partition it and calculate the Q value corresponding to each Q1 in its corresponding interval.

[0120] First, compare and statistically analyze the Q1 value data. Specifically, according to the calculation method of "the number of correct bases in the data greater than the specified Q1 \ the total number of data greater than the specified Q1", statistically analyze the error rate of the set above a specific value (i.e., the preset first threshold). The curve formed by the statistical results is as shown in Figure 3 . In Figure 3 , the closer the curve is to 0, the lower the error rate, indicating better filtering effect. Figure 3 In it, the horizontal axis is the target first parameter, which can be expressed as the target Q1 value, and the vertical axis is the error rate of calculating the bases above the filtering threshold (i.e., greater than the preset first threshold), which can be expressed as the error rate corresponding to the target Q1 value.

[0121] According to the calculation method of "the number of correct bases in the data less than the specified Q1 \ the total number of data less than the specified Q1", statistically analyze the error rate of the set of Q1 value data below a specific Q1 value (such as the preset second threshold). The curve formed by the statistical results is as shown in Figure 4 . Figure 4 In it, the closer the curve is to 0, the lower the error rate, indicating better filtering effect. Figure 4 In it, the horizontal axis is the target first parameter optimized through the preset second threshold, which can be expressed as the target Q1 value, and the vertical axis is the error rate of calculating the bases above the filtering threshold (i.e., less than the preset second threshold), which can be expressed as the error rate corresponding to the target Q1 value.

[0122] It can be seen that there is a certain relationship between the error rate and the magnitude of Q1. When Q1 is too large or too small, the error rate will be relatively high. Therefore, based on this phenomenon, a relatively low Q1 value can be given to this segmented Q1, and then a higher Q T value can be assigned according to the Q1 interval with a low error rate.

[0123] Further, perform Q value conversion and round to an integer for the error rate (P ID ) corresponding to each set formed by Q1. The specific assignment method is as follows:

[0124] Q2=-10×log 10 (2×P ID )

[0125] Among them, Q2 represents the target first parameter corresponding to the current first parameter, P ID Indicates the error rate corresponding to the set formed by the current first parameter.

[0126] Design a more accurate Q based on the classification error mode and error rate reduction mode of Q0 T The value is segmented and the error rate and the loss of filtered bases are statistically analyzed. According to the error rate, the corresponding values ​​of the Q values ​​corresponding to Q2 and the traditional Q0 value can be obtained. And in this way, we can know which parts need to be adjusted compared with the Q0 value. The best segment Q2 can be found by searching. Here, the Q2 value is combined with the machine learning algorithm and the preset Q0 value combined with the traditional algorithm to compare the data after Q_X_70 filtering ("Q_X_70" filtering means filtering out: when the Q of the base in the sequence is a number X or above Q value and the high Q bases in the sequence account for less than 70% of the total number of bases in the sequence, retain the sequence if and only if the base Q in the sequence is X or above Q value and the high Q bases in the sequence account for 70% or more of the total number of bases in the sequence), and the corresponding curve graph is as follows Figure 5 As shown, in Figure 5 The closer the curve is to 0, the lower the error rate is, indicating that the filtering effect is better. Figure 5 The algorithm for determining the original quality parameters is referred to as the Q0 algorithm, and the algorithm for determining the target quality parameters is referred to as the Q T algorithm, Figure 5 The horizontal axis represents the Q value. In the curve corresponding to the Q0 algorithm, the Q value represents the Q0 value. T The Q value in the curve corresponding to the algorithm represents Q T value; Figure 5 The vertical axis in represents the error rate. See also Figure 6 , Figure 6 The closer the curve is to 0, the fewer correct sequences are filtered incorrectly, indicating a better filtering effect. Figure 6 The algorithm for determining the original quality parameters is referred to as the Q0 algorithm, and the algorithm for determining the target quality parameters is referred to as the Q T algorithm, Figure 6 The horizontal axis represents the Q value. In the curve corresponding to the Q0 algorithm, the Q value represents the Q0 value. T The Q value in the curve corresponding to the algorithm represents Q T value; Figure 5 The vertical axis in represents the percentage of sequence loss. Figure 7 , Figure 7 shows the Q0 value and Q T The total sequence loss comparison curve of the sequences after filtering is shown in Figure 2. Figure 7Among them, the closer the curve is to 0, the fewer sequences are filtered, indicating less sequence loss and better filtering effect. From Figure 7 it can be seen that after segmentation, the sequence loss is at the corresponding Q T value. Without much flux loss, the error rate is better than or close to the traditional Q0 value, which is in line with the optimized result and its error decline trend.

[0127] To ensure the distribution of the error rate of the Q T value conforms to that of the traditional Q0 value, in the embodiment of the present application, the same data is used to draw Figure 8 . From Figure 8 in (a), (b), (c), and (d), it can be seen that at the corresponding Q T value, its error rate meets the error rate consistent with the traditional Q0 value, but the Q T value has less base loss, and this phenomenon is consistent with the trend and result reflected in the above pictures. Thus, it can be determined that the Q T value segmentation is the real quality score segmentation, and it is used as the target quality parameter of the basecall algorithm of the machine learning model. In Figure 8 the new Q new algorithm is the Q T algorithm, and the traditional Q new algorithm is the Q0 algorithm.

[0128] After obtaining the target quality parameter, the base recognition result of the base recognition model can be filtered based on the target quality parameter to obtain the target base recognition result.

[0129] In the embodiment of the present application, based on the recognition result data of the base recognition model, a target quality parameter is generated, and the target quality parameter is applied to the filtering of the base recognition result of the corresponding base recognition model. In one implementation manner, the process of generating the base recognition model includes: obtaining training sample data, where the training sample data includes optical data of the original base channel; using the real base recognition result corresponding to the training sample data as the training target, training the training sample data to obtain the base recognition model. The model structure of the base recognition model can be a neural network model or a logistic regression model. For example, an isolation forest model can be used. The real base recognition result is marked in the training sample, and during the training process, the prediction result gets closer and closer to the real base recognition result until the corresponding threshold is met, and the training ends.

[0130] In the embodiment of the present application, in addition to using the maximum probability parameter in the current base extension reaction to determine the target quality parameter, eligible probability parameters can also be used to determine the target quality parameter. In one implementation manner of the embodiment of the present application, based on the base probability distribution, that is, calculating to obtain the first parameter, including:

[0131] Obtain the test sample data for the base recognition model, as well as the predicted base recognition results obtained by recognizing the test sample data based on the base recognition model. The test sample data includes the true base recognition results; based on the correspondence between the predicted base recognition results and the true base recognition results, determine the base recognition probability screening conditions; based on the base recognition probability screening conditions, perform probability parameter screening in the probability distribution to obtain the first parameter.

[0132] Further, based on the correspondence between the predicted base recognition results and the true base recognition results, determine the base recognition probability screening conditions, including:

[0133] Obtain the data where the predicted base recognition results of the current base extension reaction in each test sample data are consistent with the true base recognition results, and statistically obtain the base prediction probability of the corresponding current base reaction in the data; based on the predicted probability, determine the screening range of the base recognition probability.

[0134] For example, the base probability distribution of the site corresponding to the current base extension reaction is [0.1, 0.2, 0.4, 0.3]. The maximum probability of this site is 0.4, and the base category corresponding to the maximum probability is base C. And the true base recognition result of this site is also base C. Statistically obtain the range of the probability range when the prediction result is consistent with the true base result. For example, it is usually concentrated between 0.4 and 0.6. If it is too large or too small, there may be problems with the sample prediction. Then use this type of screening range as the probability parameter for subsequent applications, so as to calculate the target quality parameter. Specifically, in this embodiment, only the probability parameter applied is different, and the subsequent processing process is similar to the processing process of applying the maximum probability parameter in the foregoing embodiment, which will not be elaborated here.

[0135] Next, a specific application scenario is used to illustrate the method for determining the quality parameter of base recognition in the embodiments of the present application and the effect of filtering the base recognition results.

[0136] The data source adopted in this application scenario is:

[0137] The data used for statistics is sourced from several sequencing images (fov) in a set of experiments on human genes of a complete standard human sample (HG001) using a machine learning model basecall with the latest fixed version of the sequencer and reagents at that time. Basecall is performed on these sequencing images, and the base prediction probabilities and parameters used for base prediction generated during this process are printed. The basecall result file is aligned with the human gene Hg19 reference, and the correct base corresponding to each predicted base is obtained from the aligned sam file. Then, the intermediate parameters printed by the machine learning basecall are used to assign which base each probability corresponds to and the base on the gene index (reference). Whether the type of base is consistent with the reference determines the correctness of this probability. The correctly measured base is written as 1 and assigned to the corresponding maximum probability, and the wrongly measured base is written as 0 and assigned to the corresponding maximum probability. Thus, two very important pieces of information will be obtained: the maximum probability of the base and the correct and wrong marks (numerical values 1 and 0) corresponding to the base probability.

[0138] According to the obtained data, numerical conversion is performed on the probability to obtain the second parameter Q1 = -10×log 10 (1 - P max ). Compare the second parameter with the preset value. If the second parameter is greater than the preset value, round the second parameter down to the nearest integer to obtain the first parameter; if the second parameter is not greater than the preset value, round the second parameter to the nearest integer to obtain the first parameter; where the value range of the first parameter is a positive integer from 0 to the preset value. Specifically, round Q1 to an integer value and limit the maximum value of this value to 100 (if this value exceeds 100, then this value will be assigned 100). Then, using the probability and this first parameter as data feature values (i.e., these data are the input data for the input model), and using 0 and 1 to represent the probability values of correct and wrong, fit them into the Isolation Forest model to distinguish the correct and wrong sets and obtain the base probability.

[0139] According to the probability for division, its main purpose is to segment the region where the base is located to better form a corresponding relationship with the traditional Q0 value (i.e., the original quality parameter in the embodiments of the present application). Then, use homologous data completely unrelated to the above training data for prediction and statistical separation. By dividing the probability values, Figure 9 、 Figure 10 、 Figure 11 and Figure 12 can be obtained. By Figures 9 - 12 , it can be concluded that compared with the traditional Q0 value, the segmentation of the target quality parameter is better and the error rate is lower. In actual applications, it is a relatively optimal feasible selection scheme. It should be noted that in Figure 9 、 Figure 10 、Figure 11 and Figure 12 The new Q_new algorithm is the Q T algorithm. The traditional Q_new algorithm is the Q0 algorithm. Correspondingly, the filtered Q value threshold refers to the data obtained by filtering the error rate through a preset first threshold or a preset second threshold, and the target first parameter obtained after optimizing the first parameter. For specific implementation details, reference can be made to the descriptions of the foregoing embodiments.

[0140] In an embodiment of the present application, a quality parameter determination device for base recognition is further provided, and this device can be used for the above-mentioned quality parameter determination method for base recognition.

[0141] See Figure 13 , the quality parameter determination device for base recognition includes:

[0142] An acquisition unit 201, configured to obtain recognition result data of a base recognition model for a nucleic acid sample to be recognized, where the recognition result data includes a base probability distribution for the current base extension reaction of the nucleic acid sample to be recognized;

[0143] A calculation unit 202, configured to calculate a first parameter based on the base probability distribution;

[0144] A processing unit 203, configured to perform filtering processing on the first parameter based on a preset Q0 value to obtain a target quality parameter, where Q0 = -10×log 10 e, where e is the base recognition error rate.

[0145] Optionally, the calculation unit includes:

[0146] A first acquisition subunit, configured to obtain a maximum probability parameter in the current base extension reaction based on the base probability distribution;

[0147] A first calculation subunit, configured to perform calculation based on the maximum probability parameter to obtain a first parameter.

[0148] Optionally, the first calculation subunit is specifically configured to:

[0149] Based on the formula Q1 = -10×log 10 (1 - P max ), determine a second parameter Q1, where P max represents the maximum probability parameter;

[0150] Compare the second parameter with a preset value. If the second parameter is greater than the preset value, perform rounding down to zero on the second parameter to obtain a first parameter;

[0151] If the second parameter is not greater than the preset value, round the second parameter to an integer to obtain a first parameter; wherein, the value range of the first parameter is a positive integer from 0 to the preset value.

[0152] Optionally, it further includes:

[0153] A first determination unit, configured to determine a base recognition error rate corresponding to each of the first parameters;

[0154] A statistical unit, configured to statistically obtain a first base recognition error rate corresponding to the first parameter greater than a preset first threshold and a second base recognition error rate corresponding to the first parameter less than a preset second threshold based on the error rate of base recognition corresponding to each of the first parameters;

[0155] An optimization unit, configured to, if the first base recognition error rate and / or the second base recognition error rate is greater than an error rate threshold, optimize the first parameter based on the base recognition error rate corresponding to each of the first parameters to obtain a target first parameter;

[0156] Wherein, the processing unit is specifically configured to:

[0157] Based on the preset Q0 value, perform filtering processing on the target first parameter to obtain a target quality parameter.

[0158] Optionally, the optimization unit is specifically configured to:

[0159] Based on the formula Q2 = -10×log 10 (2×P ID ) to update the first parameter, where Q2 represents the target first parameter corresponding to the current first parameter, and P ID represents the error rate corresponding to the set formed by the current first parameter.

[0160] Optionally, the processing unit includes:

[0161] A determination subunit, configured to determine segmented data corresponding to the preset Q0 value based on the preset Q0 value;

[0162] A processing subunit, configured to perform filtering processing on the first parameter based on the segmented data to obtain a target quality parameter, so that the base filtering range of the target quality parameter matches the base filtering range of the Q0 value.

[0163] Optionally, it further includes:

[0164] A filtering unit, configured to filter the base recognition result of the base recognition model based on the target quality parameter to obtain a target base recognition result.

[0165] Optionally, it further includes: a model training unit, which is used for:

[0166] Obtaining training sample data, where the training sample data includes optical data of the original base channel;

[0167] Using the true base recognition result corresponding to the training sample data as the training target, training the training sample data to obtain a base recognition model.

[0168] Optionally, the calculation unit includes:

[0169] An identification subunit, configured to obtain test sample data of the base recognition model, and a predicted base recognition result obtained by recognizing the test sample data based on the base recognition model, where the test sample data includes a true base recognition result;

[0170] A second determination subunit, configured to determine a base recognition probability screening condition based on the correspondence between the predicted base recognition result and the true base recognition result;

[0171] A screening subunit, configured to perform probability parameter screening in the probability distribution based on the base recognition probability screening condition to obtain a first parameter.

[0172] Optionally, the second determination subunit is specifically configured to:

[0173] Obtain data in which the predicted base recognition result of the current base extension reaction in each test sample data is consistent with the true base recognition result, and statistically obtain the base prediction probability of the corresponding current base reaction in the data;

[0174] Based on the predicted probability, determine the screening range of the base recognition probability.

[0175] It should be noted that the specific implementation of each unit and subunit in this embodiment can refer to the corresponding content in the previous text, and will not be elaborated here.

[0176] In another embodiment of the present application, a readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for determining quality parameters for base recognition as described in any one of the above.

[0177] In another embodiment of the present application, an electronic device is further provided. The electronic device may include:

[0178] A memory, configured to store an application program and data generated by running the application program;

[0179] A processor for executing the application program to implement the method for determining quality parameters for base recognition as described in any one of the above.

[0180] It should be noted that the specific implementation of the processor in this embodiment can refer to the corresponding content in the foregoing, and will not be elaborated here.

[0181] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0182] Those skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0183] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0184] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for determining quality parameters for base recognition, characterized in that, Including: Obtaining recognition result data of a base recognition model for a nucleic acid sample to be recognized, where the recognition result data includes a base probability distribution for the current base extension reaction of the nucleic acid sample to be recognized; Calculating a first parameter based on the base probability distribution; Based on a preset Q0 value, filter the first parameter to obtain a target quality parameter, where Q0 = -10×log 10 e, where e is the base recognition error rate.

2. The method according to claim 1, wherein The calculating the first parameter based on the base probability distribution includes: Obtaining a maximum probability parameter in the current base extension reaction based on the base probability distribution; Calculating a first parameter based on the maximum probability parameter; Optionally, the calculating the first parameter based on the maximum probability parameter includes: Based on the formula Q1 = -10×log 10 (1 - P max ), determine the second parameter Q1, where P max represents the maximum probability parameter; Comparing the second parameter with a preset value, and if the second parameter is greater than the preset value, rounding down the second parameter to an integer to obtain the first parameter; If the second parameter is not greater than the preset value, rounding the second parameter to an integer to obtain the first parameter; where the value range of the first parameter is a positive integer from 0 to the preset value.

3. The method according to any one of claims 1-2, characterized in that, Also including: Determining a base recognition error rate corresponding to each of the first parameters; Based on the base recognition error rate corresponding to each of the first parameters, statistically obtaining a first base recognition error rate corresponding to the first parameters greater than a preset first threshold, and statistically obtaining a second base recognition error rate corresponding to the first parameters less than a preset second threshold; If the first base recognition error rate and / or the second base recognition error rate is greater than an error rate threshold, optimizing the first parameter based on the base recognition error rate corresponding to each of the first parameters to obtain a target first parameter; Wherein, filtering the first parameter based on a preset Q0 value to obtain a target quality parameter includes: Filtering the target first parameter based on the preset Q0 value to obtain a target quality parameter; Optionally, the optimizing the first parameter based on the base recognition error rate corresponding to each of the first parameters to obtain a target first parameter includes: Update the first parameter based on the formula Q2 = -10×log 10 (2×P ID ), where Q2 represents the target first parameter corresponding to the current first parameter, and P ID represents the error rate corresponding to the set formed by the current first parameter.

4. The method according to any one of claims 1 to 3, characterized in that, The filtering the first parameter based on a preset Q0 value to obtain a target quality parameter includes: Determining segmented data corresponding to the preset Q0 value based on the preset Q0 value; Filtering the first parameter based on the segmented data to obtain a target quality parameter, so that the base filtering range of the target quality parameter matches the base filtering range of the Q0 value.

5. The method according to any one of claims 1 to 4, characterized in that Also including: Filtering the base recognition result of the base recognition model based on the target quality parameter to obtain a target base recognition result.

6. The method according to any one of claims 1-5, characterized in that, Also including: Obtaining training sample data, where the training sample data includes optical data of an original base channel; Using the true base recognition result corresponding to the training sample data as a training target, training the training sample data to obtain a base recognition model.

7. The method according to any one of claims 1-6, characterized in that, The calculating the first parameter based on the base probability distribution includes: Obtaining test sample data of the base recognition model, and a predicted base recognition result obtained by recognizing the test sample data based on the base recognition model, where the test sample data includes a true base recognition result; Determine a base recognition probability screening condition based on the correspondence between the predicted base recognition result and the true base recognition result; Based on the base recognition probability screening condition, perform probability parameter screening in the probability distribution to obtain a first parameter; Optionally, the determining a base recognition probability screening condition based on the correspondence between the predicted base recognition result and the true base recognition result includes: Obtain the data in each test sample data where the predicted base recognition result of the current base extension reaction is consistent with the true base recognition result, and statistically obtain the base prediction probability of the corresponding current base reaction in the data; Based on the prediction probability, determine the screening range of the base recognition probability.

8. A quality parameter determination device for base recognition, characterized in that, It includes: An acquisition unit for obtaining the recognition result data of the base recognition model for the nucleic acid sample to be recognized, where the recognition result data includes the base probability distribution of the current base extension reaction for the nucleic acid sample to be recognized; A calculation unit for calculating a first parameter based on the base probability distribution; A processing unit, configured to perform filtering processing on the first parameter based on a preset Q0 value to obtain a target quality parameter, where Q0 = -10×log 10 e, where e is the base recognition error rate.

9. The device according to claim 8, characterized in that, The calculation unit includes: A first acquisition subunit for obtaining the maximum probability parameter in the current base extension reaction based on the base probability distribution; A first calculation subunit for calculating based on the maximum probability parameter to obtain a first parameter; Optionally, the first calculation subunit is specifically used for: Based on the formula Q1 = -10×log 10 (1 - P max ), determine the second parameter Q1, where P max represents the maximum probability parameter; Compare the second parameter with a preset value. If the second parameter is greater than the preset value, round down the second parameter to an integer to obtain a first parameter; If the second parameter is not greater than the preset value, round the second parameter to an integer to obtain a first parameter; wherein, the value range of the first parameter is a positive integer from 0 to the preset value; Optionally, the device further includes: A first determination unit for determining the base recognition error rate corresponding to each first parameter; A statistics unit for statistically obtaining the first base recognition error rate corresponding to the first parameter greater than a preset first threshold and the second base recognition error rate corresponding to the first parameter less than a preset second threshold based on the base recognition error rate corresponding to each first parameter; An optimization unit for, if the first base recognition error rate and / or the second base recognition error rate is greater than an error rate threshold, optimizing the first parameter based on the base recognition error rate corresponding to each first parameter to obtain a target first parameter; Wherein, the processing unit is specifically used for: Based on the preset Q0 value, perform filtering processing on the target first parameter to obtain a target quality parameter; Optionally, the optimization unit is specifically used for: Based on the formula Q2 = -10×log 10 (2×P ID ), update the first parameter, where Q2 represents the target first parameter corresponding to the current first parameter, and P ID represents the error rate corresponding to the set formed by the current first parameter; Optionally, the processing unit includes: A determination subunit for determining the segmented data corresponding to the preset Q0 value based on the preset Q0 value; A processing subunit for performing filtering processing on the first parameter based on the segmented data to obtain a target quality parameter, so that the base filtering range of the target quality parameter matches the base filtering range of the Q0 value.

10. The device according to any one of claims 8-9, characterized in that, It further includes: A filtering unit for filtering the base recognition result of the base recognition model based on the target quality parameter to obtain a target base recognition result; Optionally, the device further includes: a model training unit, and the model training unit is configured to: Obtain training sample data, where the training sample data includes optical data of an original base channel; Use the true base recognition result corresponding to the training sample data as a training target, and train the training sample data to obtain a base recognition model; Optionally, the calculation unit includes: An identification subunit, configured to obtain test sample data of the base recognition model, and a predicted base recognition result obtained by recognizing the test sample data based on the base recognition model, where the test sample data includes a true base recognition result; A second determination subunit, configured to determine a base recognition probability screening condition based on the correspondence between the predicted base recognition result and the true base recognition result; A screening subunit, configured to perform probability parameter screening in the probability distribution based on the base recognition probability screening condition to obtain a first parameter; Optionally, the second determination subunit is specifically configured to: Obtain data in each test sample data where the predicted base recognition result of the current base extension reaction is consistent with the true base recognition result, and statistically obtain the base prediction probability of the corresponding current base reaction in the data; Determine a screening range of the base recognition probability based on the predicted probability.