Base recognition method, table generation method, table, related device, and sequencing system
By generating and using tables to store key-value pairs of sequencing information and base recognition results, the problem of high computational complexity of machine learning models is solved, and the rapid acquisition of base recognition results and optimization of computing resources is achieved.
Patent Information
- Application Number
- CN202510682030.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-02
AI Technical Summary
The existing machine learning models have high computational complexity in the base recognition process, resulting in slow prediction speed and occupies a lot of computing resources of the equipment, and it is impossible to quickly obtain base recognition results.
By generating and using tables, multiple sets of key-value pairs of sequencing information and their corresponding base recognition results are stored, and table query is used instead of running the base recognition model every time to obtain the base recognition results of the specified sequencing reaction.
It improves the efficiency of obtaining base recognition results, reduces the occupation of computing resources, and improves the speed and accuracy of base recognition.
Smart Images

Figure CN120581065A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of biological information processing technology, and more specifically to a base recognition method, a table generation method, a table, related devices and a sequencing system. Background Art
[0002] Next-generation gene sequencing (NGS) is a high-throughput sequencing technology that can quickly and cost-effectively generate large amounts of sequencing data. Genetic sequencing result files contain a variety of data, including at least two key components: sequence base data (typically represented by bases, consisting of A, C, G, and T) and quality score data (denoted as the Quanlity Score, or Q-value). In current applications, machine learning models can significantly reduce the error rate of base calls through complex calculations and corrections. Machine learning models can use features such as the signals generated during the sequencing process as input features and then predict the base calls based on feature learning.
[0003] The prediction process of machine learning models is computationally complex, especially when dealing with large-scale sequencing data. The prediction speed is slow. If base recognition results need to be obtained through machine learning models, it may not be possible to obtain predictions and output results in a short period of time. At the same time, it will also increase the computing power resources occupied by the device. Summary of the Invention
[0004] In view of this, this application provides the following technical solutions:
[0005] In a first aspect of the embodiments of the present application, a base recognition method is provided, comprising:
[0006] Obtaining a table comprising multiple sets of sequencing information and base recognition results forming key-value pairs with each set of sequencing information, wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding sequencing reaction or multiple consecutive sequencing reactions including the corresponding sequencing reaction;
[0007] At least one set of sequencing information of a specified round of sequencing reaction is obtained, and information query is performed in the table to determine the base recognition result of the specified round of sequencing reaction.
[0008] In a second aspect of the embodiments of the present application, a table generation method is provided, comprising:
[0009] Inputting each set of sequencing information from the plurality of sets of sequencing information into a base recognition model to obtain a base recognition result corresponding to each set of sequencing information; wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding round of sequencing reaction or a plurality of consecutive rounds of sequencing reactions including the corresponding round of sequencing reaction;
[0010] Each set of sequencing information and the corresponding base recognition results are stored in a preset table in the form of key-value pairs to obtain a table applied to base recognition.
[0011] In a third aspect of the embodiments of the present application, a table is provided, which is generated based on the table generation method described in the second aspect of the embodiments of the present application.
[0012] In a fourth aspect of the embodiments of the present application, a base recognition device is provided, comprising:
[0013] a first acquisition unit, configured to acquire a table comprising multiple sets of sequencing information and base recognition results forming key-value pairs with each set of sequencing information, wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding sequencing round or a plurality of consecutive sequencing rounds including the corresponding sequencing round;
[0014] The query unit is configured to obtain at least one set of sequencing information of a specified round of sequencing reaction, perform information query in the table, and determine a base recognition result of the specified round of sequencing reaction.
[0015] In a fifth aspect of the embodiments of the present application, a table generating device is provided, comprising:
[0016] a second acquisition unit, configured to input each set of sequencing information from the plurality of sets of sequencing information into a base recognition model to obtain a base recognition result corresponding to each set of sequencing information; wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding round of sequencing reaction or a plurality of consecutive rounds of sequencing reactions including the corresponding round of sequencing reaction;
[0017] The storage unit is used to store each set of sequencing information and the corresponding base recognition results in a preset table in the form of key-value pairs to obtain a table applied to base recognition.
[0018] In a sixth aspect of the embodiments of the present application, a sequencing system is provided, comprising:
[0019] a memory for storing an application and a table, wherein the table includes multiple sets of sequencing information and base recognition results forming key-value pairs with each set of sequencing information, wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding sequencing reaction or a plurality of consecutive sequencing reactions including the corresponding sequencing reaction;
[0020] A processor, configured to execute the application program to implement:
[0021] Get the form;
[0022] At least one set of sequencing information of a specified round of sequencing reaction is obtained, and information query is performed in the table to determine the base recognition result of the specified round of sequencing reaction.
[0023] Through the above technical solutions, it can be seen that the present application discloses a base recognition method, a table generation method, a table, a related device and a sequencing system. In the base recognition method, by obtaining a table, at least one set of sequencing information of a specified sequencing round reaction is obtained, and information query is performed in the table to determine the base recognition result of the specified sequencing reaction. Among them, the table includes multiple sets of sequencing information and base recognition results that form key-value pairs with each set of sequencing information. The base recognition result is determined based on each set of sequencing information and a base recognition model. By generating a table for each set of sequencing information and the corresponding base recognition result, the problem of low computational efficiency and complex processing of obtaining the base recognition result by running the base recognition model in each base recognition process is solved, and the efficiency of obtaining the base recognition result is improved by looking up the table. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0025] Figure 1 A schematic diagram of a base recognition method provided in an embodiment of the present application;
[0026] Figure 2 A partial schematic diagram of a table provided in an embodiment of the present application;
[0027] Figure 3 A time rise curve comparison diagram provided in an embodiment of the present application;
[0028] Figure 4 Another time rise curve comparison diagram provided in an embodiment of the present application;
[0029] Figure 5 A schematic diagram of a table of bases and quality scores provided in an embodiment of the present application;
[0030] Figure 6 A flowchart of a table generation method provided in an embodiment of the present application;
[0031] Figure 7 A schematic structural diagram of a base recognition device provided in an embodiment of the present application;
[0032] Figure 8 A schematic diagram of the structure of a table generating device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] In this application, the terms "first" and "second" are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements and may include steps or elements that are not listed.
[0035] In the embodiments of the present application, the term "sequencing" may also be referred to as "nucleic acid sequencing" or "gene sequencing". The three terms are interchangeable and all refer to the determination of the type and arrangement order of bases or nucleotides (including nucleotide analogs) in nucleic acid molecules. The so-called sequencing includes the process of binding nucleotides to nucleic acid templates and collecting the corresponding signals emitted by the nucleotides (including analogs). The so-called sequencing includes synthesis sequencing (sequencing by synthesis, SBS) and / or ligation sequencing (sequencing by ligation, SBL), including DNA sequencing and / or RNA sequencing.
[0036] Sequencing generally includes multiple rounds to realize the process of determining the type and arrangement order of multiple bases or nucleotides on the nucleic acid template. In the embodiment of the present application, each round of "process to realize the determination of the type and arrangement order of multiple bases or nucleotides on the nucleic acid template" is referred to as "one round of sequencing". "One round of sequencing" (cycle) is also called "sequencing round", which can be defined as a base extension of four nucleotides / bases. In other words, "one round of sequencing" can be defined as the determination of the base or nucleotide type at any specified position on the nucleic acid template. For sequencing platforms that realize sequencing based on polymerization or ligation reaction, one round of sequencing includes realizing that four nucleotides (including nucleotide analogs) are bound to the so-called nucleic acid template in a base complementary manner, and collecting the corresponding signal process emitted. Among them, for platforms that realize sequencing based on polymerization reaction, the reaction system includes reaction substrate nucleotides, polymerase and nucleic acid template, and a sequence (sequencing primer) is combined with the nucleic acid template. Based on the base pairing principle and the polymerization reaction principle, the added reaction substrate nucleotides are connected to the sequencing primer under the catalysis of the polymerase to realize the binding of the nucleotide to the specific position of the nucleic acid template. Generally, one round of sequencing may include one or more base extensions (repeat). For example, four nucleotides are added to the reaction system in sequence, and base extension and corresponding reaction signal collection are performed respectively, and one round of sequencing includes four base extensions. For another example, any combination of the four nucleotides is added to the reaction system, such as two-by-two combinations or one-by-three combinations, and base extension and corresponding reaction signal collection are performed for two combinations respectively, and one round of sequencing includes two base extensions. For another example, the four nucleotides are added to the reaction system at the same time for base extension and reaction signal collection, and one round of sequencing includes one base extension.
[0037] Taking the second-generation sequencing technology as an example, the second-generation sequencing technology uses the characteristics of different fluorescent molecules with different fluorescence emission wavelengths. Different fluorescent molecules are used to mark the substrates of the base extension reaction. After the base extension reaction occurs, the fluorescent molecules are irradiated with laser to excite the fluorescent molecules to generate fluorescent signals, and the fluorescent signals of specific wavelengths are obtained by optical sensors. Finally, the type of base bound to the fluorescent molecules is identified based on the fluorescent signals. Taking four-color sequencing as an example, four different fluorescent molecules are used to mark four different bases (A, T or U, G, C) respectively. Four bases are added simultaneously during a round of sequencing to complete a base extension reaction. Laser irradiation is used to excite the fluorescent molecules to generate fluorescent signals, and an optical imaging system is used to collect the fluorescent signals generated by different fluorescent molecules and form an image. In one method, the image is processed and the fluorescent signal position is located to perform base cluster detection. Template construction is performed based on the base cluster detection results of multiple images corresponding to the sequencing signal responses of different base types to construct the positions of all base cluster template points. Based on the template, optical data is extracted from the filtered image (primarily the fluorescence signal intensity), which is then corrected. Finally, the base type is identified based on the maximum intensity at each base cluster template point. After multiple rounds of sequencing, the base sequence of the nucleic acid template can be obtained.
[0038] A variety of data will appear in the file containing the second-generation sequencing results, among which the main data include sequence base data (expressed as base: composed of A, C, G, T) and quality score data (Quanlity Score, referred to as Q value). The base data error rate after complex calculation and correction by machine learning is lower, and the quality score is relatively more accurate. In the process of obtaining base recognition results through machine learning models, the calculation time of the machine learning model to predict the base recognition results based on sequencing information is relatively long, and the computing resources occupied are also large, which reduces the efficiency of base recognition result prediction and output. Therefore, a base recognition method is provided in an embodiment of the present application, which can improve the efficiency of obtaining base recognition results and reduce the occupation of computing resources.
[0039] See also Figure 1 , is a flow chart of a base recognition method provided in an embodiment of the present application, which may include the following steps:
[0040] S101. Obtain a table.
[0041] S102: Obtain at least one set of sequencing information for a specified round of sequencing reaction, perform information query in a table, and determine the base recognition result of the specified round of sequencing reaction.
[0042] In the embodiment of the present application, the table includes multiple sets of sequencing information and base recognition results that form key-value pairs with each set of sequencing information. In other words, the table in the embodiment of the present application includes multiple key-value pairs, the "key" in each key-value pair is a set of sequencing information, and the "value" is the base recognition result corresponding to the set of sequencing information. Among them, each set of sequencing information is determined based on the intensity characteristics of the signal generated by a corresponding round of sequencing reaction or multiple consecutive rounds of sequencing reactions including the same round of sequencing reaction. The base recognition result is determined based on each set of sequencing information and a base recognition model. The base recognition model is a machine learning model trained based on sequencing information samples. The base recognition results output by the base recognition model may include the base recognition type and may also include quality score data corresponding to the recognized base. Each set of sequencing information may be a signal feature that can be used to identify the corresponding base recognition result. The sequencing system can identify the base by detecting fluorescent signals, electrical signals, or other physical signals. In the embodiment of the present application, each set of sequencing information is mainly determined based on the fluorescence signal intensity characteristics. A single round of sequencing reaction refers to the sequencing information obtained from a single round of sequencing reaction. The multiple consecutive rounds of sequencing reactions encompassing that round of sequencing reaction can, for example, include the current round of sequencing reaction and its corresponding previous and subsequent rounds of sequencing reactions. For example, if the current round is the fifth round of sequencing reaction, the corresponding multiple consecutive rounds of sequencing reactions can be the third, fourth, fifth, and sixth rounds of sequencing reactions. The signal intensity characteristics obtained through multiple rounds of sequencing reactions can enhance the accuracy of base recognition results during subsequent base recognition.
[0043] In the table generated for querying base call results, each set of sequencing information forms a key-value pair with its corresponding base call result. That is, each set of sequencing information is stored in a corresponding, matched pair. For example, the first set of sequencing information corresponds to the first base call result, the second set of sequencing information corresponds to the second base call result, and so on. The nth set of sequencing information corresponds to the nth base call result. This allows subsequent queries of the table to retrieve the corresponding base call result based on the current set of sequencing information without having to rerun the base call model, improving the efficiency of obtaining base call results.
[0044] Accordingly, when querying the base call results corresponding to a specified sequencing round, the table can be queried for a set of sequencing information corresponding to the specified sequencing round to obtain the corresponding base call results. Alternatively, the table can be queried for a set of sequencing information determined based on the intensity characteristics of the signals generated by multiple consecutive sequencing rounds including the specified sequencing round to obtain the corresponding base call results. For example, the table can be queried based on the fluorescence brightness information of the current specified sequencing round to obtain the corresponding base type and the quality score corresponding to the base type.
[0045] In the embodiments of the present application, the input information corresponding to different base recognition models may be different. The inventors will think of all possible input features of base recognition models (i.e., multiple sets of input sequencing information) and their corresponding output results (i.e., base recognition results, which may include base recognition types and corresponding quality parameters), and store them in a table, so that the corresponding base recognition results can be obtained by querying the table based on at least one set of sequencing information of a specified round of sequencing reaction, without having to run the base recognition model every time the base recognition results of a specified round of sequencing reaction are needed, thereby improving the efficiency of obtaining base recognition results and reducing the occupation of computing resources.
[0046] The base recognition method of the embodiment of the present application is described below in conjunction with corresponding application scenarios.
[0047] In the embodiment of the present application, each set of sequencing information stored in the table is determined based on the signal intensity characteristics generated by the corresponding round of sequencing reaction or the continuous multiple rounds of sequencing reactions containing the round of sequencing reaction. In order to accurately represent the signal characteristics in the sequencing reaction, it is also possible to reduce the storage space occupied by the data storage in the table, so that subsequent table queries can be quickly implemented. In one embodiment, the intensity characteristic is the intensity characteristic of the intensity data of the signal generated by the round of sequencing reaction corresponding to each set of sequencing information or the continuous multiple rounds of sequencing reactions containing the round of sequencing reaction after segmentation processing according to a preset number of segments. Among them, the intensity data characterizes the original intensity or corrected intensity of the signal; the corrected intensity is the intensity after the original intensity is corrected; correspondingly, the correction includes at least one of crosstalk correction, phase loss correction, and overflow correction, and the correction can also be set according to the actual application scenario requirements.
[0048] In this embodiment, the raw intensity can be the intensity data of the fluorescent signal directly detected by the optical imaging system, such as the actual brightness value of the fluorescent signal detected by the optical imaging system on the image. The raw intensity of the fluorescent signal has not been processed in any way and reflects the initial state of the fluorescent signal, but may contain noise, interference, or other errors. In order to make the intensity data more accurate, the raw intensity of the fluorescent signal can be corrected to obtain a corrected intensity to eliminate noise, interference, and errors in the fluorescent signal, so that the fluorescent signal is closer to the true value, thereby improving the accuracy of base calling.
[0049] Crosstalk is a phenomenon in which the signal from one channel interferes with the signals in other channels in a multichannel optical imaging system. During the imaging process, especially in multicolor fluorescence imaging, different fluorescence channels may interfere with each other. Crosstalk correction can be used to identify and quantify the degree of interference between different channels. The signal intensity of each channel is then adjusted to eliminate interference from signals in other channels, thereby improving signal purity and accuracy. This ensures that the signal detected in each channel is derived solely from fluorescence at its corresponding wavelength, without interference from signals in other channels.
[0050] Phase dephasing manifests as phase lag (phasing or phase) or phase advance (prephasing or prephase). Phase lag refers to the fact that a nucleotide analogue that should have reacted and incorporated into the nucleic acid template in round P, but instead participated in the reaction in round P+1. Phase advance refers to the fact that a nucleotide analogue that should have reacted and incorporated into the nucleic acid template in round P, but instead participated in the reaction in round P-1, that is, crosstalk occurs between adjacent sequencing rounds in the same channel. In this case, the correct sequencing signal for that round of sequencing reaction cannot be identified, that is, the sequence information of the nucleic acid template molecule cannot be accurately obtained. Therefore, this phase error needs to be corrected to improve the accuracy of base recognition.
[0051] Overflow refers to the "overflow" of signal brightness between adjacent base clusters, resulting in overlapping brightness between adjacent base clusters. To eliminate this overlapping brightness, the sum of the surrounding brightness is subtracted from the base cluster signal according to a certain ratio. The specific overflow correction expression is shown in formula (1).
[0052]
[0053] Among them, Intsbleed(i) is the signal after overflow correction; IntsR(i) is the original signal of the central base cluster; bld is the empirical parameter; N i is the set of all spatially neighboring base cluster points of the i-th base cluster; IntsR(j) is the signal strength of the adjacent base cluster. Through overflow correction, the accuracy and reliability of the signal can be improved.
[0054] The above-mentioned correction processes are illustrated based on the corresponding application scenarios. In addition, the corresponding correction processing mode can be determined based on the actual application scenario. For example, there can be base deviation correction, chemical correction, time drift correction, etc. The embodiments of this application do not limit the various correction methods as long as they can improve the signal accuracy.
[0055] Whether the intensity data is original intensity or corrected intensity, it is usually a continuous floating-point value. If these floating-point values are directly stored in a table, more memory will be taken up, affecting subsequent query efficiency. At the same time, if these floating-point values are used for base recognition model training, the complexity of the eigenvalues mentioned below will be increased. Therefore, in an embodiment of the present application, the intensity data is the data after segmentation processing according to the preset number of segments. Wherein, the preset segmentation can be determined according to the data range corresponding to the usual intensity data. For example, the intensity data characterizes the sequencing signal brightness value, and the preset segmentation is the total interval number of the discretized segmentation of the brightness value, for converting continuous floating-point signals into integer features, thereby reducing the occupancy of storage space. Exemplarily, if the brightness range of the signal is 1 to 1000 and the preset number of segments is 10, each segment corresponds to 100 units (1-100 is segment 1, 101-200 is segment 2, and the rest may be deduced by analogy). For example, if the signal brightness representing ACGT is (100, 200, 300, 400), after being segmented according to the preset number of segments, the converted integer brightness can be (0, 1, 2, 3). In the embodiment of the present application, the specific number of segments is set based on the signal strength distribution and accuracy requirements, mainly to balance the feature discrimination and the memory efficiency of the lookup table.
[0056] In one embodiment of the present application, each set of sequencing information includes a combination of at least one or more of the following parameters:
[0057] (1) A base type combination for the M+N+1 round of sequencing reaction determined based on the signal intensity characteristics of the current round of sequencing reaction, the signal intensity characteristics of the N rounds of sequencing reactions preceding the current round of sequencing reaction, and the signal intensity characteristics of the M rounds of sequencing reactions following the current round of sequencing reaction. Wherein, M and N are both integers greater than or equal to 0. Correspondingly, the values of M and N can be determined according to actual application requirements, such as by using a preset memory space in a table, or by a preset base recognition accuracy.
[0058] In one embodiment, both M and N can take the value of 1. In this case, each set of sequencing information includes the base type combination of three rounds of sequencing reactions determined based on the signal intensity characteristics of the current round of sequencing reaction, the signal intensity characteristics of the previous round of sequencing reaction, and the signal intensity characteristics of the subsequent round of sequencing reaction.
[0059] For example, for a certain base cluster template point, the signal brightness of the current round (nth round) sequencing reaction in the four channels A / C / G / T of the optical imaging system (the four channels A / C / G / T are used to detect the signal on the A base, the signal on the C base, the signal on the G base, and the signal on the T base, respectively, and are sometimes referred to as the four base channels A, T / U, C, and G, or simply the four base channels in this article) are: A=80, C=150, G=200, T=50; the signal brightness of the previous round (n-1th round) sequencing reaction in the four channels A / C / G / T are: A=30, C=180, G=90, T=40; the signal brightness of the next round (n+1th round) sequencing reaction in the four channels A / C / G / T are: A=70, C=60, G=210, T=20. If the base type is determined based on the highest brightness, then the base type incorporated into the base cluster template site in the current sequencing reaction is G, the base type incorporated into the base cluster template site in the previous sequencing reaction is predicted to be C, and the base type incorporated into the base cluster template site in the next sequencing reaction is predicted to be G. The corresponding base type combination for these three rounds is CGG. This information can be used as sequencing information for table lookup.
[0060] In one embodiment, M can be 0 and N can be 2. In this case, each set of sequencing information includes the base type combination of three rounds of sequencing reactions determined based on the signal intensity characteristics of the current round of sequencing reaction and the signal intensity characteristics of the two rounds of sequencing reactions before the current round of sequencing reaction.
[0061] It should be noted that in the embodiment of the present application, the base combination of the corresponding sequencing round is used to obtain the corresponding base recognition result. The base combination here is determined based on the base brightness information of the corresponding sequencing round. For example, the base combination of the corresponding sequencing round is a three-round base combination corresponding to the current sequencing round, the previous sequencing round, and the next sequencing round. The base combination here is a preliminary recognition result of the current base and the previous and next bases, and the preliminary recognition result is the base type determined according to the brightness information in the sequencing, such as directly comparing the original brightness information of the four base channels of AGCT, or it can be the preliminary base recognition result determined according to the corrected brightness (the final brightness after background removal, crosstalk, phasing / prephasing correction, brightness normalization correction, etc.). That is, when obtaining the preliminary base recognition result, the base type of the current sequencing round is determined according to the original brightness of the four base channels or the channel with the brightest brightness after correction.
[0062] When using a base recognition model to obtain a base recognition result, the base recognition model can determine the final base recognition result output by the base recognition model based on the preliminary base recognition result. The base recognition model is a machine learning model that learns sequencing information to output the base type of the current sequencing round. During the base recognition model training process, in addition to obtaining preliminary base recognition results based on the brightness information described above, the model can also be trained based on other sequencing information, such as characteristics of the sequencer and characteristics of sequencing signal changes over time. This allows the use of multi-dimensional information to improve the accuracy and reliability of base recognition by the base recognition model.
[0063] (2) A base type combination for K rounds of sequencing reactions determined based on the signal intensity characteristics of the K rounds of sequencing reactions preceding the current round of sequencing reactions; wherein K is a positive number greater than or equal to 1. The specific value of K can be determined based on actual scenario requirements and the information configuration format preset in the table.
[0064] In one embodiment, the value of K can be 3. In this case, each set of sequencing information can be a base type combination of three sequencing rounds determined based on the signal intensity characteristics of the three sequencing rounds preceding the current sequencing round. For example, if the signal intensity indicates that the base type in the round preceding the current round is A, the base type in the two rounds preceding the current round is C, and the base type in the three rounds preceding the current round is G, then the sequencing information can be represented by GCA.
[0065] In the embodiment of the present application, the signal strength characteristics of the front and back rounds and the base types can be combined to determine the base recognition results of a specified sequencing round, which can enhance the robustness of the prediction.
[0066] In another embodiment of the present application, each group of sequencing information can also be the intensity feature ranking of the signal generated by the current round sequencing reaction in the four base channels of A, T / U, C, and G. If the actual value of the signal intensity generated by each round sequencing reaction is directly stored in the table, the storage space of the table will be larger, and the query efficiency will also be reduced in the subsequent table lookup process. Therefore, in the embodiment of the present application, the actual intensity feature can be converted into an intensity feature ranking (or expressed as a brightness ranking), which can both characterize the signal intensity feature and improve the table lookup efficiency. For example, the brightness of ACGT corresponding to the current round sequencing reaction is (100, 200, 300, 600), and the intensity feature ranking after conversion can be represented by (0, 1, 2, 3).
[0067] Using intensity feature ranking to represent sequencing information can save storage space in the table, but it can also lose some accuracy in some application scenarios. For example, when the intensity features representing two base channels are relatively close, the brightness ranking will lose a certain amount of accuracy. To address this problem, in one embodiment of the present application, each set of sequencing information can determine a first target parameter based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and a preset number of segments, where the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction.
[0068] In this embodiment, the maximum and second maximum values of the signal intensity of the current round of sequencing reaction can be determined based on the intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G. Specifically, when calculating the first parameter, the maximum value of the signal intensity of the current round of sequencing reaction and the second maximum value of the signal intensity of the current round of sequencing reaction can be calculated according to the actual signal intensity value of each base cluster template point position. Exemplarily, during the current round of sequencing reaction, if the intensity of the signal generated at a certain base cluster template point position in the four base channels of A, T / U, C, and G is (70, 15, 10, 5) respectively, then the maximum value of the signal intensity generated at the base cluster template point position is 70, the second maximum value is 15, and the sum of the signal intensities of the four base channels is 100. The preset number of segments is to convert the discretized data into a finite integer to facilitate fast matching in the table lookup method. The preset number of segments (W) is the total number of intervals for segmenting continuous numerical values, such as W=10. For example, if the first target parameter is represented by GN, GN = [maximum signal intensity of the current sequencing reaction / (maximum signal intensity of the current sequencing reaction + second-largest signal intensity of the current sequencing reaction)] * W. This allows the relationship between the maximum and second-largest signal intensities to be determined using the first target parameter, compensating for the loss of precision associated with using intensity feature ranking to represent signal intensity characteristics.
[0069] In another embodiment of the present application, each set of sequencing information can also be based on the second target parameter determined by the maximum signal intensity of the current round of sequencing reaction, the second parameter value and the preset number of segments. Wherein, the second parameter characterizes the sum of the signal intensities of the four base channels of the current round of sequencing reaction. The second target parameter can characterize the magnitude relationship between the maximum signal intensity and the total signal intensity of the current round of sequencing reaction, and can further compensate for the loss of precision caused by the intensity feature ranking representation of the signal intensity characteristic. For example, the preset number of segments is W, and MQ is used to represent the second target parameter, then MQ=(the maximum signal intensity of the current round of sequencing reaction / the sum of the signal intensities of the four base channels of the current round of sequencing reaction)*W.
[0070] It should be noted that the above-mentioned sequencing information is illustrated by the corresponding embodiments, and the corresponding sequencing information can also be the corresponding combination of data determined by the above-mentioned various embodiments. For example, each set of sequencing information can include the intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G; the above-mentioned first target parameter and second target parameter. For another example, each set of sequencing information can also include the intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G; the above-mentioned first target parameter and second target parameter; and the base type combination of the three rounds of sequencing reactions determined based on the signal intensity characteristics of the current round of sequencing reaction and the signal intensity characteristics of the two rounds of sequencing reactions before the current round of sequencing reaction.
[0071] Furthermore, the ranking of the intensity characteristics of the signals generated by the current round of sequencing reactions in the four base channels of A, T / U, C, and G described in the above-mentioned various embodiments is determined after sorting the original intensity or corrected intensity of the signals generated by the current round of sequencing reactions in the four base channels of A, T / U, C, and G according to the intensity size. In the embodiment of the present application, the intensity characteristics of the four base channels can be ranked directly based on the original intensity, so that the ranking information of the four base channels can be obtained quickly; or the intensity characteristics of the four base channels can be ranked based on the corrected intensity, so that the ranking information of the four base channels can be obtained more accurately. Correspondingly, the processing method of the ranking information can be determined according to the actual application scenario requirements, and the present application does not impose any restrictions on this.
[0072] In one embodiment of the present application, multiple sets of sequencing information are determined based on the intensity characteristics of the signals generated by multiple rounds of sequencing reactions. The multiple rounds of sequencing reactions are divided into multiple segments, each segment of the sequencing reaction includes a preset number of sequencing rounds, and the preset number of sequencing rounds in the same segment correspond to the same set of sequencing information.
[0073] In the process of querying base recognition results based on the table, multiple rounds of sequencing reactions are divided into several segments, each segment includes a preset number of sequencing reactions, and all rounds in the same segment share the same set of sequencing information, which can reduce the complexity of table lookup while maintaining the accuracy of base recognition. For example, the sequencer continuously outputs the signal intensity of multiple rounds of sequencing reactions, such as the signal intensity of 100 rounds of sequencing reactions. If each segment contains 25 rounds, the 100 rounds can be divided into 4 segments, segment 1 corresponding to rounds 1 to 25; segment 2 corresponding to rounds 26 to 50; segment 3 corresponding to rounds 51 to 75; segment 4 corresponding to rounds 76 to 100. Then, the intensity features within each segment can be extracted, such as the brightness ranking, the first target parameter GN that characterizes the relationship between the maximum signal intensity and the second largest signal intensity, the second target parameter MQ that characterizes the relationship between the maximum signal intensity and the total signal intensity, and the base type of the previous sequencing round (such as the base type combination of the three sequencing reactions of the current sequencing round, the previous sequencing round, and the next sequencing round). The above features are combined into a unique identifier (key) for the segment. For example, the sequencing information for the segment can be represented as: ranking (0, 1, 2, 3)_GN_MQ_base type C of the previous sequencing round. This sequencing information can then be queried in the table to obtain the base call results, such as the corresponding base call type and Q value.
[0074] Furthermore, in the process of generating the table, each set of sequencing information and base recognition results stored can be processed to reduce the amount of data storage occupied, and at the same time, the query efficiency during the query process can also be improved. Therefore, in an embodiment of the present application, data formatting processing can also be performed on information such as sequencing information and base recognition results to reduce the amount of data storage occupied. Correspondingly, each set of sequencing information, base recognition results and preset number of rounds in the table are represented in the form of integer unit data, and the data form includes a data form of a target number of bits, a type data form occupying a target byte, or at least one of a binary digital form. In this embodiment, in order to maximize memory efficiency and table lookup speed, sequencing information, base recognition results and preset number of rounds are all converted into compact integer form for storage.
[0075] The signal brightness of the current round of sequencing reaction can be truncated as needed. For example, all data can be converted to 8-bit data format, or it can be converted to 4-bit data format. For example, the signal brightness data can be converted to A=00, C=01, G=10, T=11 (2 bits), so that 2 bits are used instead of 1 byte. You can also use the smallest byte type such as int8 or unit8 to store values. For example, if the GN value is 0.75, it can be stored as 7 to avoid floating point number occupation. You can also splice multiple features into a binary string as a data index. For example, if the sequencing information includes brightness ranking and GN value, brightness ranking (01) + GN (0110) = 010110, which can be directly used as a memory address offset to save storage space.
[0076] Specifically, taking brightness ranking as an example, the original brightness values can be floating-point intensity values, such as A=80, T=50, C=150, and G=200. These values are converted to integers between 0 and 3 by intensity sorting. In this case, the data is stored as uint8, which reduces memory usage. For GN or MQ values, such as GN=0.57, the conversion can be multiplied by the preset number of segments N and rounded to the nearest integer (e.g., N=10, in which case GN=0.57×10≈6). In this case, the data is stored as int8, with a range of -128 to 127, sufficient to cover segments with N ≤ 100. For base call results, a 2-bit identifier can be used to identify the base type, such as A:00, C:01, G:10, T:11. If the predicted result is G, the value can be 10. The floating-point data for Q values can be converted to integers. The preset round number is also represented as an integer. For example, if the sequencing round number ranges from 1 to 100, the segment storage is segmented into segments of 25 rounds, with the segment number represented as uint8 (e.g., 0-4).
[0077] The method for constructing the base recognition model in the embodiment of the present application includes:
[0078] Obtain a training sample set; perform machine learning modeling on each sample in the training sample set based on a specific model structure to obtain a base call model. Each training sample in the training sample set is annotated with a feature value and a target value, where the feature value is each set of sequencing information in the table, and the target value is the base call result corresponding to each set of sequencing information. The base call result may include a base call type and / or a base sequencing quality score.
[0079] The data source of the training sample set can be raw sequencing data generated using a fixed version of sequencer, reagents, and base recognition software, such as fluorescence signal intensity data. In order to improve the accuracy of the base recognition model, the training sample set can cover a variety of sequencing scenarios, such as high / low complexity sequences, different sequencing depths, etc. The feature value is the input value of the base recognition model, which can include the brightness ranking of each round of sequencing reaction, such as the order of the intensity of the four channels A / T / C / G (such as [T(0), A(1), C(2), G(3)]); GN / MQ value: the discretized ratio (such as GN=6, MQ=5); context feature: the combination of base types in the previous and next rounds (such as previous base C+current base G+next base G). The target value is the label of the training sample, which can include base recognition type and Q value. Base recognition type: the true base (A / C / G / T) annotated by high-precision calibration (such as Sanger sequencing or consensus sequence); Q value can be a quality score calculated based on the error rate.
[0080] Among them, the specific model structure can include the lightGBM model structure, which is suitable for structured numerical features, and the output is 4 probability values (corresponding to A / C / G / T), and the highest probability is the predicted base. The model structure can also be a convolutional neural network (CNN), which is suitable for the need to capture local signal patterns, such as the spatial correlation of brightness rankings. A hybrid model, such as CNN+LightGBM, can also be used to process the original signal intensity image with CNN, splice the CNN output with structured features such as GN / MQ values, and input it into LightGBM. Then, according to the loss function, the gap between the predicted value and the true value of the base recognition model can be calculated, thereby adjusting the model parameters of the base recognition model to obtain the final base recognition model.
[0081] To facilitate data query, in an embodiment of the present application, multiple subtables can be generated based on the number of sequencing rounds, so that queries can be performed in the corresponding subtables based on the current sequencing round number, thereby improving query efficiency. Accordingly, obtaining at least one set of sequencing information for a specified round of sequencing reactions and performing information query in the table to determine the base recognition results of the specified round of sequencing reactions includes:
[0082] When the table includes at least one sub-table, a round number range corresponding to a specified round of sequencing reaction is determined; based on the round number range, a target sub-table is determined in the table; and information query is performed in the target sub-table based on at least one set of sequencing information of the specified round of sequencing reaction to determine the base recognition result of the specified round of sequencing reaction.
[0083] Each subtable in the table contains records for a corresponding round range. This record information represents multiple sets of sequencing information within that round range, along with the base call results that form key-value pairs with each set of sequencing information. The round range includes a predetermined number of sequencing reactions. For example, data from sequencing reactions from rounds 1-100 is recorded in subtable 1, data from reactions from 101-200 is recorded in subtable 2, and data from reactions from 201-300 is recorded in subtable 3. If the currently designated round is 150, data queries are performed from subtable 2, which can improve data query efficiency.
[0084] The following is an example of an actual application scenario to illustrate the base recognition method in the embodiment of the present application.
[0085] Since the signal intensity features used in training the base recognition model do not have a fixed pattern, if the sequencing information related to all intensity features is to be summarized, each set of sequencing information is combined into a unique identification ID (equivalent to an identification code, or key). The types of sequencing information covered are many (such as signal intensity information, base type combinations corresponding to the sequencing round, and related target parameters, etc.), and the electronic device memory will not be able to store them. Therefore, it is necessary to segment the continuously changing floating-point values representing the signal intensity, thereby saving the size of the table and reducing the complexity of the intensity features. The inventors sorted the signal brightness at any base cluster template point position on the image and gave each brightness a new brightness ranking based on the sorting ranking to represent the signal intensity. For example: the brightness of the signal generated at a base cluster template point position obtained by a certain round of sequencing reaction in the four base channels A / C / G / T is (100, 200, 300, 600); then the brightness ranking after conversion is (0, 1, 2, 3). At this point, the brightness feature of the brightness arrangement order can be obtained. For example, the relationship between the maximum brightness and the second-largest brightness of the signal at a certain base cluster template point position in the four base channels can be represented by the first target parameter GN, GN = [maximum brightness of the signal at that position in the current round of sequencing / (maximum brightness of the signal at that position in the current round of sequencing + second-largest brightness of the signal at that position in the current round of sequencing)] * W, where W is the preset number of segments. The relationship between the maximum brightness of the signal at that position and the total brightness of the signal at that position in the four base channels is represented by MQ, MQ = (maximum brightness of the signal at that position in the current round of sequencing / sum of the brightness of the signal at that position in the four base channels in the current round of sequencing) * W, where W is the preset number of segments. Then, the above features are rounded off to the nearest integer to obtain a GN value with a minimum value of n1 and a maximum value of n2; then, an MQ value with a minimum value of n3 and a maximum value of n4 is obtained; the brightness features of the brightness order of the four bases at the position reflected by the previous round of sequencing reaction, the brightness features of the brightness order of the four bases at the position reflected by the current round of sequencing, and the brightness features of the brightness order of the four bases at the position reflected by the next round of sequencing, the sequence features of the continuous base combination formed by the first five base types reflected by the current round of sequencing and the base types reflected by the next round of sequencing, can be used as a set of sequencing information.
[0086] Use the segmented signal intensity data to train a machine learning model (i.e., a base recognition model, hereinafter referred to as "model1"), and then use model1 to predict the results of each segment combination for each possible set of sequencing information composed of all possible segmented signal intensity data (such as the original intensity data can be segmented according to the intensity ranking), and integrate these segment combinations and prediction results into a table. Then use model1 to predict the probabilities of the corresponding four bases, and use the method of calculating the Q value using the probability values of the four bases obtained by machine learning to calculate the Q value corresponding to the base. The features of each possible set of sequencing information are merged as the ID of the hash table, and then the predicted bases of the corresponding parameters of machine learning and the Q values corresponding to the bases calculated by the machine learning predicted probability are placed in the corresponding hash table ID to form a key-value pair, so that the information data of all tables are obtained.
[0087] The large amount of data in the machine learning prediction table and the large amount of input feature data will lead to memory explosion problems during table lookup. Therefore, the data structure of the table lookup is transformed. First, the predicted bases (segmented string features) are converted into numbers to reduce the memory usage of the string. At the same time, the data structure of the data is transformed. All data formats are converted to int8 type (int8 represents a data form, which is the smallest integer unit data form in computer storage) and char (char is also a counting unit in computer storage and is the smallest data unit form in computer storage like int8) is used for counting. This method can save the size of the memory occupied by the data in the running memory during prediction, reduce the size of the memory occupied by the table, and also reduce the complexity of the key (corresponding hash table ID) required for table lookup, thereby speeding up the table lookup. The time complexity required for table lookup has also been optimized.
[0088] Due to the limited reading time and memory of the CPU, the table size read into the table has also been segmented and optimized. The segmentation method is to segment the number of sequencing rounds according to the time of the sequencing reaction. For example, 100 rounds of sequencing reactions are divided into four segments, such as segment 1, segment 2, segment 3 and segment 4, and each segment has 25 rounds as a node. For example, if the sequencing reaction performs the 5th round of sequencing reaction, then this round of sequencing reaction belongs to segment 1. If the sequencing reaction performs the 50th round of sequencing reaction, then this round of sequencing reaction belongs to segment 2. Therefore, a complex table with a length of 100 rounds can be divided into 4 smaller tables, thereby greatly reducing the complexity of the table and speeding up the table lookup speed. At the same time, it can also reduce the occurrence of errors caused by the repeated use of IDs. See Figure 2 , which shows a portion of the table, in Figure 2In the table shown, com can represent base type combinations corresponding to sequencing rounds. For example, com1, com2, and com3 can represent base type combinations for multiple sequencing rounds of the first form, multiple sequencing rounds of the second form, and multiple sequencing rounds of the third form. It should be noted that the number of sequencing rounds corresponding to each com can be the same or different, and base type combinations corresponding to the same number of sequencing rounds but different sequencing rounds can be selected. Specifically, for example, com1 represents the base type combination of the current sequencing round reaction, the base type combination of the sequencing round before the current sequencing round reaction, and the base type combination of the sequencing round after the current sequencing round reaction; com2 represents the base type combination of the current sequencing round reaction and the base type combination of the two sequencing rounds before the current sequencing round reaction; and com3 represents the base type combination of the three sequencing rounds before the current sequencing round reaction. The data corresponding to A, C, G, and T are represented by brightness ranking, GN and MQ are represented by integers, cyc represents the number of sequencing rounds, pre represents the base type identifier in the base call result, such as 0 can represent base A; and Q represents the quality score of the corresponding base type.
[0089] To further reduce the data storage space occupied by the table and improve the efficiency of table query, the features of all the chars in the table can be calculated and converted into a unique key-value binary number X1, and the corresponding base identification type and Q value are also converted into the corresponding unique value Y1. The synthesis method of the binary number X1 is to convert the number of bytes occupied by each char-based feature used in the value range of that feature into a binary number X1 represented in sequence. Then this X1 can be used as an array subscript and a mark. The unique value Y1 corresponding to the base and Q value is then used as the content of the array. This can greatly reduce memory and speed up the table lookup. Regarding the choice of array, after testing, the vector array in the C++ standard library is faster than the hash container (unordered_map, map) in the standard library. Because although the hash table lookup also has a time complexity of O(1), each table lookup requires calculating the hash address corresponding to the feature, and using the feature as the subscript of the array does not require this step. See Figure 3 and Figure 4 ,exist Figure 3 and Figure 4 The new array table lookup method represents the table query method in the embodiment of the present application. It can be concluded that the table query method adopted in the embodiment of the present application is highly efficient.
[0090] Through research, the table's data format allows for ranking the four brightness levels using the numbers 0, 1, 2, and 3. This simplifies the data ranking compared to actual base brightness without compromising the data's meaning. To account for the loss of precision between base brightness data, the table uses the ratio of the maximum brightness to the sum of the maximum and second-highest brightnesses to determine the slight difference between the maximum and second-highest brightnesses, compensating for the loss of precision caused by ranking base brightness. Brightness prominence is calculated by dividing the maximum brightness of the current test base by the sum of all base brightnesses for that base, further compensating for the loss of precision. Prediction results are recorded at the end of the table, using the numbers 0, 1, 2, and 3 to represent the four bases ACGT. The corresponding Q value is then calculated based on the predicted features for each ID. After obtaining the Q value, a new table is used to obtain the result data and brightness Q value. The re-comparison results and the actual error rate are then used to fine-tune the machine-learned Q value, resulting in a more accurate Q value. The fine-tuning method is to represent all Q values as continuous numbers, then calculate the actual error rate of this Q value under this ID, and then calculate the actual Q value based on this error rate. Then, compare this latest Q value with the Q value calculated by machine learning to get the average difference, and then perform a difference completion operation to obtain a more accurate Q value.
[0091] Through the above steps, a very small two-dimensional array table can be obtained. The array operation results are very fast and the table of corresponding bases and quality scores can be output at the same time. Figure 5 Using the table lookup method is equivalent to greatly accelerating the speed of machine learning prediction and reducing the difficulty of machine learning engineering.
[0092] In the embodiment of the present application, a table generation method is also provided. The relevant contents of the table generation method have been described in the aforementioned base recognition method and will not be described in detail here. Figure 6 The table generation method may include the following steps:
[0093] S201: Input each set of sequencing information from the multiple sets of sequencing information into a base recognition model to obtain a base recognition result corresponding to each set of sequencing information.
[0094] Each set of sequencing information is determined based on the intensity characteristics of the signal generated by a corresponding round of sequencing reaction or multiple consecutive rounds of sequencing reactions including the corresponding round of sequencing reaction;
[0095] S202: Store each set of sequencing information and the corresponding base recognition result in a preset table in the form of key-value pairs to obtain a table used for base recognition.
[0096] Optionally, the intensity feature is the intensity feature after the intensity data of the signal generated by a round of sequencing reaction corresponding to each set of sequencing information or multiple consecutive rounds of sequencing reactions including this round of sequencing reaction are segmented according to a preset number of segments; wherein the intensity data represents the original intensity or corrected intensity of the signal; the corrected intensity is the intensity after the original intensity is corrected; optionally, the correction includes at least one of crosstalk correction, phase dephasing correction, and overflow correction.
[0097] Optionally, each set of sequencing information includes a combination of at least one or more of the following parameters:
[0098] A base type combination for an M+N+1 round of sequencing reaction determined based on the signal intensity characteristics of the current round of sequencing reaction, the signal intensity characteristics of N rounds of sequencing reactions preceding the current round of sequencing reaction, and the signal intensity characteristics of M rounds of sequencing reactions following the current round of sequencing reaction;
[0099] Wherein, M and N are integers greater than or equal to 0;
[0100] A base type combination for a K-round sequencing reaction is determined based on signal intensity characteristics of K-round sequencing reactions preceding the current round of sequencing reaction;
[0101] Wherein, K is an integer greater than or equal to 1;
[0102] The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G;
[0103] A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction;
[0104] The second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value and the preset number of segments, and the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction.
[0105] Optionally, each set of sequencing information includes at least one or more combinations of the following parameters:
[0106] The base type combination of the three rounds of sequencing reactions is determined based on the signal intensity characteristics of the current round of sequencing reaction, the signal intensity characteristics of the previous round of sequencing reaction, and the signal intensity characteristics of the subsequent round of sequencing reaction;
[0107] The base type combination of the three rounds of sequencing reactions is determined based on the signal intensity characteristics of the current round of sequencing reaction and the signal intensity characteristics of the two rounds of sequencing reactions before the current round of sequencing reaction;
[0108] A base type combination of three rounds of sequencing reactions determined based on signal intensity characteristics of the three rounds of sequencing reactions preceding the current round of sequencing reaction;
[0109] The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G;
[0110] A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction;
[0111] The second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value and the preset number of segments, and the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction.
[0112] Optionally, the ranking of the intensity characteristics of the signals generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G is determined based on sorting the original intensities or corrected intensities of the signals generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G according to the intensity.
[0113] Optionally, multiple sets of sequencing information are determined based on the intensity characteristics of signals generated by multiple rounds of sequencing reactions, and the multiple rounds of sequencing reactions are divided into multiple segments, each segment of the sequencing reaction includes a preset number of sequencing rounds, and the preset number of sequencing rounds in the same segment correspond to the same set of sequencing information.
[0114] Optionally, each set of sequencing information, base recognition results and preset rounds in the table is represented in integer unit data form, and the data form includes at least one of a data form of a target number of bits, a type data form occupying a target number of bytes, or a binary digital form.
[0115] Optionally, the base calling result includes a base calling type and / or a base sequencing quality score.
[0116] Optionally, the method for constructing a base recognition model includes:
[0117] A training sample set is obtained, where each training sample in the training sample set is marked with a characteristic value and a target value. The characteristic value is each set of sequencing information in the table, and the target value is the base recognition result corresponding to each set of sequencing information.
[0118] Based on a specific model structure, machine learning modeling is performed on each sample in the training sample set to obtain a base recognition model.
[0119] Optionally, the specific model structure includes a lightGBM model structure.
[0120] Correspondingly, a table is also provided in an embodiment of the present application, which is generated based on the above-mentioned table generation method.
[0121] In the embodiment of the present application, a base recognition device is also provided. Figure 7 , the device comprises:
[0122] A first acquiring unit 301 is configured to acquire a table, the table comprising multiple sets of sequencing information and base recognition results forming key-value pairs with each set of sequencing information, wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding sequencing round or a plurality of consecutive sequencing rounds including the corresponding sequencing round;
[0123] The query unit 302 is configured to obtain at least one set of sequencing information of a specified round of sequencing reaction, perform information query in the table, and determine a base recognition result of the specified round of sequencing reaction.
[0124] Optionally, the intensity feature is the intensity feature of the intensity data of the signal generated by a round of sequencing reaction corresponding to each set of sequencing information or multiple consecutive rounds of sequencing reactions including this round of sequencing reaction after segmentation processing according to a preset number of segments; wherein, the intensity data represents the original intensity or corrected intensity of the signal; the corrected intensity is the intensity after correction of the original intensity; optionally, the correction includes at least one of crosstalk correction, phase dephasing correction, and overflow correction.
[0125] Optionally, each set of sequencing information includes a combination of at least one or more of the following parameters:
[0126] A base type combination for an M+N+1 round of sequencing reaction determined based on the signal intensity characteristics of a current round of sequencing reaction, the signal intensity characteristics of N rounds of sequencing reactions before the current round of sequencing reaction, and the signal intensity characteristics of M rounds of sequencing reactions after the current round of sequencing reaction;
[0127] Wherein, M and N are integers greater than or equal to 0;
[0128] A base type combination for a K-round sequencing reaction is determined based on signal intensity characteristics of the K-round sequencing reactions preceding the current-round sequencing reaction;
[0129] Wherein, K is an integer greater than or equal to 1;
[0130] The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G;
[0131] A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction;
[0132] The second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value and the preset number of segments, and the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction.
[0133] Optionally, each set of sequencing information includes a combination of at least one or more of the following parameters:
[0134] A base type combination for three rounds of sequencing reactions determined based on a signal intensity characteristic of a current round of sequencing reaction, a signal intensity characteristic of a previous round of sequencing reaction, and a signal intensity characteristic of a subsequent round of sequencing reaction;
[0135] A base type combination for three rounds of sequencing reactions determined based on the signal intensity characteristics of the current round of sequencing reaction and the signal intensity characteristics of the two rounds of sequencing reactions before the current round of sequencing reaction;
[0136] A base type combination of three rounds of sequencing reactions determined based on signal intensity characteristics of the three rounds of sequencing reactions preceding the current round of sequencing reaction;
[0137] The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G;
[0138] A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction;
[0139] The second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value and the preset number of segments, and the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction.
[0140] Optionally, the ranking of the intensity characteristics of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G is determined based on sorting the original intensity or corrected intensity of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G according to the intensity size.
[0141] Optionally, the multiple sets of sequencing information are determined based on the intensity characteristics of the signals generated by multiple rounds of sequencing reactions, and the multiple rounds of sequencing reactions are divided into multiple segments, each segment of the sequencing reaction includes a preset number of sequencing rounds, and the preset number of sequencing rounds in the same segment correspond to the same set of sequencing information.
[0142] Optionally, each set of sequencing information, the base recognition result and the preset number of rounds in the table is represented in an integer unit data form, and the data form includes at least one of a data form of a target number of bits, a type data form occupying a target byte, or a binary digital form.
[0143] Optionally, the base recognition result includes a base recognition type and / or a base sequencing quality score.
[0144] Optionally, the base recognition result is determined based on each set of sequencing information and a base recognition model.
[0145] Optionally, the base recognition apparatus comprises a model building unit, wherein the model building unit is configured to:
[0146] A training sample set is obtained, wherein each training sample in the training sample set is marked with a characteristic value and a target value, the characteristic value is each set of sequencing information in the table, and the target value is the base recognition result corresponding to each set of sequencing information.
[0147] Based on a specific model structure, machine learning modeling is performed on each sample in the training sample set to obtain a base recognition model.
[0148] Optionally, the specific model structure includes a lightGBM model structure.
[0149] Optionally, the query unit includes:
[0150] a first determining subunit, configured to determine, when the table includes at least one subtable, a round number range corresponding to a specified round of sequencing reaction; wherein each subtable in the table has record information corresponding to a round number range, the record information representing multiple sets of sequencing information within the round number range and base recognition results forming key-value pairs with each set of sequencing information, the round number range including a preset number of sequencing rounds;
[0151] a second determining subunit, configured to determine a target subtable in the table based on the round number range;
[0152] The query subunit is configured to perform information query in the target subtable based on at least one set of sequencing information of the designated round of sequencing reaction, and determine the base recognition result of the designated round of sequencing reaction.
[0153] Correspondingly, a table generating device is also provided in the embodiment of the present application, see Figure 8 , the device comprises:
[0154] A second acquisition unit 401 is configured to input each set of sequencing information from the plurality of sets of sequencing information into a base recognition model to obtain a base recognition result corresponding to each set of sequencing information; wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding sequencing reaction or a plurality of consecutive sequencing reactions including the corresponding sequencing reaction;
[0155] The storage unit 402 is configured to store each set of sequencing information and the corresponding base recognition result in a preset table in the form of a key-value pair, to obtain a table used for base recognition.
[0156] Optionally, the intensity feature is the intensity feature of the intensity data of the signal generated by a round of sequencing reaction corresponding to each set of sequencing information or multiple consecutive rounds of sequencing reactions including this round of sequencing reaction after segmentation processing according to a preset number of segments; wherein, the intensity data represents the original intensity or corrected intensity of the signal; the corrected intensity is the intensity after correction of the original intensity; optionally, the correction includes at least one of crosstalk correction, phase dephasing correction, and overflow correction.
[0157] Optionally, each set of sequencing information includes a combination of at least one or more of the following parameters:
[0158] A base type combination for an M+N+1 round of sequencing reaction determined based on the signal intensity characteristics of a current round of sequencing reaction, the signal intensity characteristics of N rounds of sequencing reactions before the current round of sequencing reaction, and the signal intensity characteristics of M rounds of sequencing reactions after the current round of sequencing reaction;
[0159] Wherein, M and N are integers greater than or equal to 0;
[0160] A base type combination for a K-round sequencing reaction is determined based on signal intensity characteristics of the K-round sequencing reactions preceding the current-round sequencing reaction;
[0161] Wherein, K is an integer greater than or equal to 1;
[0162] The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G;
[0163] A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction;
[0164] The second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value and the preset number of segments, and the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction.
[0165] Optionally, each set of sequencing information includes at least one or more combinations of the following parameters:
[0166] A base type combination for three rounds of sequencing reactions determined based on a signal intensity characteristic of a current round of sequencing reaction, a signal intensity characteristic of a previous round of sequencing reaction, and a signal intensity characteristic of a subsequent round of sequencing reaction;
[0167] A base type combination for three rounds of sequencing reactions determined based on the signal intensity characteristics of the current round of sequencing reaction and the signal intensity characteristics of the two rounds of sequencing reactions before the current round of sequencing reaction;
[0168] A base type combination of three rounds of sequencing reactions determined based on signal intensity characteristics of the three rounds of sequencing reactions preceding the current round of sequencing reaction;
[0169] The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G;
[0170] A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction;
[0171] The second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value and the preset number of segments, and the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction.
[0172] Optionally, the ranking of the intensity characteristics of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G is determined based on sorting the original intensity or corrected intensity of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G according to the intensity size.
[0173] Optionally, the multiple sets of sequencing information are determined based on the intensity characteristics of the signals generated by multiple rounds of sequencing reactions, and the multiple rounds of sequencing reactions are divided into multiple segments, each segment of the sequencing reaction includes a preset number of sequencing rounds, and the preset number of sequencing rounds in the same segment correspond to the same set of sequencing information.
[0174] Optionally, each set of sequencing information, the base recognition result and the preset number of rounds in the table is represented in an integer unit data form, and the data form includes at least one of a data form of a target number of bits, a type data form occupying a target byte, or a binary digital form.
[0175] Optionally, the base recognition result includes a base recognition type and / or a base sequencing quality score.
[0176] Optionally, the base recognition result is determined based on each set of sequencing information and a base recognition model.
[0177] Optionally, the device further comprises a construction unit, wherein the construction unit is configured to:
[0178] A training sample set is obtained, wherein each training sample in the training sample set is marked with a characteristic value and a target value, the characteristic value is each set of sequencing information in the table, and the target value is the base recognition result corresponding to each set of sequencing information.
[0179] Based on a specific model structure, machine learning modeling is performed on each sample in the training sample set to obtain a base recognition model.
[0180] Optionally, the specific model structure includes a lightGBM model structure.
[0181] It should be noted that the specific implementation of each unit and sub-unit in this embodiment can refer to the corresponding content in the previous text and will not be described in detail here.
[0182] In another embodiment of the present application, a sequencing system is provided, comprising: a memory for storing an application and a table, the table comprising multiple sets of sequencing information and base recognition results forming key-value pairs with each set of sequencing information, the each set of sequencing information being determined based on intensity characteristics of signals generated by a corresponding sequencing reaction or a plurality of consecutive sequencing reactions including the corresponding sequencing reaction;
[0183] A processor, configured to execute the application program to implement:
[0184] Get the form;
[0185] At least one set of sequencing information of a specified round of sequencing reaction is obtained, and information query is performed in the table to determine the base recognition result of the specified round of sequencing reaction.
[0186] In another embodiment of the present application, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the base identification method and table generation method as described in any one of the above items are implemented.
[0187] It should be noted that the specific implementation of the processor in this embodiment can refer to the corresponding content in the previous text and will not be described in detail here.
[0188] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0189] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0190] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0191] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A base recognition method, characterized in that: include: Obtaining a table comprising multiple sets of sequencing information and base recognition results forming key-value pairs with each set of sequencing information, wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding sequencing reaction or multiple consecutive sequencing reactions including the corresponding sequencing reaction; At least one set of sequencing information of a specified round of sequencing reaction is obtained, and information query is performed in the table to determine the base recognition result of the specified round of sequencing reaction.
2. The base recognition method according to claim 1, wherein The intensity feature is an intensity feature of the signal intensity data generated by a round of sequencing reaction corresponding to each set of sequencing information or a plurality of consecutive rounds of sequencing reactions including the round of sequencing reaction, after segmentation according to a preset number of segments; Wherein, the intensity data represents the original intensity or the corrected intensity of the signal; The corrected strength is the strength after correcting the original strength; Optionally, the correction includes at least one of crosstalk correction, phase loss correction, and overflow correction.
3. The base recognition method according to claim 2, wherein Each set of sequencing information includes at least one or more combinations of the following parameters: A base type combination for an M+N+1 round of sequencing reaction determined based on the signal intensity characteristics of a current round of sequencing reaction, the signal intensity characteristics of N rounds of sequencing reactions before the current round of sequencing reaction, and the signal intensity characteristics of M rounds of sequencing reactions after the current round of sequencing reaction; Wherein, M and N are integers greater than or equal to 0; A base type combination for a K-round sequencing reaction is determined based on signal intensity characteristics of the K-round sequencing reactions preceding the current-round sequencing reaction; Wherein, K is an integer greater than or equal to 1; The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G; A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction; A second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value, and the preset number of segments, wherein the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction; Optionally, each set of sequencing information includes at least one or more combinations of the following parameters: A base type combination for three rounds of sequencing reactions determined based on a signal intensity characteristic of a current round of sequencing reaction, a signal intensity characteristic of a previous round of sequencing reaction, and a signal intensity characteristic of a subsequent round of sequencing reaction; A base type combination for three rounds of sequencing reactions determined based on the signal intensity characteristics of the current round of sequencing reaction and the signal intensity characteristics of the two rounds of sequencing reactions before the current round of sequencing reaction; A base type combination of three rounds of sequencing reactions determined based on signal intensity characteristics of the three rounds of sequencing reactions preceding the current round of sequencing reaction; The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G; A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction; A second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value, and the preset number of segments, wherein the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction; Optionally, the ranking of the intensity characteristics of the signals generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G is determined based on sorting the original intensities or corrected intensities of the signals generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G according to their intensities; Optionally, the multiple sets of sequencing information are determined based on intensity characteristics of signals generated by multiple rounds of sequencing reactions, the multiple rounds of sequencing reactions being divided into multiple segments, each segment of sequencing reaction comprising a preset number of sequencing rounds, and the preset number of sequencing rounds in the same segment corresponding to the same set of sequencing information; Optionally, each set of sequencing information, the base recognition result, and the preset round number in the table is represented in an integer unit data form, wherein the data form includes at least one of a data form of a target number of bits, a type data form occupying a target number of bytes, or a binary digital form; Optionally, the base recognition result includes a base recognition type and / or a base sequencing quality score; Optionally, the base recognition result is determined based on each set of sequencing information and a base recognition model; The method for constructing the base recognition model comprises: Obtaining a training sample set, wherein each training sample in the training sample set is annotated with a characteristic value and a target value, wherein the characteristic value is each set of sequencing information in the table, and the target value is a base recognition result corresponding to each set of sequencing information; Performing machine learning modeling on each sample in the training sample set based on a specific model structure to obtain a base recognition model; Optionally, the specific model structure includes a lightGBM model structure; Optionally, obtaining at least one set of sequencing information of a specified round of sequencing reaction, performing information query in the table, and determining the base recognition result of the specified round of sequencing reaction, includes: In a case where the table includes at least one subtable, determining a round range corresponding to a specified round of sequencing reaction; wherein each subtable in the table has record information corresponding to the round range, the record information representing multiple sets of sequencing information within the round range and base recognition results forming key-value pairs with each set of sequencing information, and the round range includes a preset number of sequencing reactions; determining a target subtable in the table based on the round number range; An information query is performed in the target subtable based on at least one set of sequencing information of the designated round of sequencing reaction to determine a base recognition result of the designated round of sequencing reaction.
4. A table generation method, characterized in that: include: Inputting each set of sequencing information from the multiple sets of sequencing information into a base recognition model to obtain a base recognition result corresponding to each set of sequencing information; Wherein, each set of sequencing information is determined based on the intensity characteristics of the signal generated by a corresponding round of sequencing reaction or multiple consecutive rounds of sequencing reactions including the corresponding round of sequencing reaction; Each set of sequencing information and the corresponding base recognition results are stored in a preset table in the form of key-value pairs to obtain a table applied to base recognition.
5. The table generation method according to claim 4, characterized in that: The intensity feature is an intensity feature of the signal intensity data generated by a round of sequencing reaction corresponding to each set of sequencing information or a plurality of consecutive rounds of sequencing reactions including the round of sequencing reaction, after segmentation according to a preset number of segments; The intensity data represents the original intensity or the corrected intensity of the signal; the corrected intensity is the intensity after the original intensity is corrected; Optionally, the correction includes at least one of crosstalk correction, phase loss correction, and overflow correction.
6. The table generation method according to claim 5, characterized in that: Each set of sequencing information includes at least one or more combinations of the following parameters: A base type combination for an M+N+1 round of sequencing reaction determined based on the signal intensity characteristics of a current round of sequencing reaction, the signal intensity characteristics of N rounds of sequencing reactions before the current round of sequencing reaction, and the signal intensity characteristics of M rounds of sequencing reactions after the current round of sequencing reaction; Wherein, M and N are integers greater than or equal to 0; A base type combination for a K-round sequencing reaction is determined based on signal intensity characteristics of the K-round sequencing reactions preceding the current-round sequencing reaction; Wherein, K is an integer greater than or equal to 1; The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G; A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction; A second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value, and the preset number of segments, wherein the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction; Optionally, each set of sequencing information includes at least one or more combinations of the following parameters: A base type combination for three rounds of sequencing reactions determined based on a signal intensity characteristic of a current round of sequencing reaction, a signal intensity characteristic of a previous round of sequencing reaction, and a signal intensity characteristic of a subsequent round of sequencing reaction; A base type combination for three rounds of sequencing reactions determined based on the signal intensity characteristics of the current round of sequencing reaction and the signal intensity characteristics of the two rounds of sequencing reactions before the current round of sequencing reaction; A base type combination of three rounds of sequencing reactions determined based on signal intensity characteristics of the three rounds of sequencing reactions preceding the current round of sequencing reaction; The intensity characteristic ranking of the signal generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G; A first target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, a first parameter, and the preset number of segments, wherein the first parameter represents the sum of the maximum signal intensity of the current round of sequencing reaction and the second maximum signal intensity of the current round of sequencing reaction; A second target parameter is determined based on the maximum signal intensity of the current round of sequencing reaction, the second parameter value, and the preset number of segments, wherein the second parameter represents the sum of the signal intensities of the four base channels of the current round of sequencing reaction; Optionally, the ranking of the intensity characteristics of the signals generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G is determined based on sorting the original intensities or corrected intensities of the signals generated by the current round of sequencing reaction in the four base channels of A, T / U, C, and G according to their intensities; Optionally, the multiple sets of sequencing information are determined based on intensity characteristics of signals generated by multiple rounds of sequencing reactions, the multiple rounds of sequencing reactions being divided into multiple segments, each segment of sequencing reaction comprising a preset number of sequencing rounds, and the preset number of sequencing rounds in the same segment corresponding to the same set of sequencing information; Optionally, each set of sequencing information, the base recognition result, and the preset round number in the table is represented in an integer unit data form, wherein the data form includes at least one of a data form of a target number of bits, a type data form occupying a target number of bytes, or a binary digital form; Optionally, the base recognition result includes a base recognition type and / or a base sequencing quality score; Optionally, the method for constructing the base recognition model comprises: Obtaining a training sample set, wherein each training sample in the training sample set is annotated with a characteristic value and a target value, wherein the characteristic value is each set of sequencing information in the table, and the target value is a base recognition result corresponding to each set of sequencing information; Performing machine learning modeling on each sample in the training sample set based on a specific model structure to obtain a base recognition model; Optionally, the specific model structure comprises a lightGMB model structure.
7. A table, characterized in that: The table is generated according to the table method according to any one of claims 4 to 6.
8. A base recognition device, characterized in that: include: a first acquisition unit, configured to acquire a table comprising multiple sets of sequencing information and base recognition results forming key-value pairs with each set of sequencing information, wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding sequencing round or a plurality of consecutive sequencing rounds including the corresponding sequencing round; The query unit is configured to obtain at least one set of sequencing information of a specified round of sequencing reaction, perform information query in the table, and determine a base recognition result of the specified round of sequencing reaction.
9. A table generating device, characterized in that: include: a second acquisition unit, configured to input each set of sequencing information from the plurality of sets of sequencing information into a base recognition model to obtain a base recognition result corresponding to each set of sequencing information; wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding round of sequencing reaction or a plurality of consecutive rounds of sequencing reactions including the corresponding round of sequencing reaction; The storage unit is used to store each set of sequencing information and the corresponding base recognition results in a preset table in the form of key-value pairs to obtain a table applied to base recognition.
10. A sequencing system, characterized in that include: a memory for storing an application and a table, wherein the table includes multiple sets of sequencing information and base recognition results forming key-value pairs with each set of sequencing information, wherein each set of sequencing information is determined based on intensity characteristics of signals generated by a corresponding sequencing reaction or a plurality of consecutive sequencing reactions including the corresponding sequencing reaction; A processor, configured to execute the application program to implement: Get the form; At least one set of sequencing information of a specified round of sequencing reaction is obtained, and information query is performed in the table to determine the base recognition result of the specified round of sequencing reaction.
Citation Information
Patent Citations
Table lookup method and device for voice feature extraction, computer equipment and storage medium
CN110866142A
Base identification method and device, electronic equipment and storage medium
CN118429965A
Basic group identification method and system
CN118429967A
Graph reference genome and base-calling approach using imputed haplotypes
US20230095961A1