Base calling method and system, analysis method and system, and electronic device and storage medium

By analyzing the light intensity signal of each DNB in ​​a time dimension, determining the light intensity distribution interval and using distributed calculations, the problems of poor base recognition quality and large data processing volume in the prior art are solved, and fast and accurate base recognition is achieved.

WO2025138729A1PCT designated stage expired Publication Date: 2025-07-03MGI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/106234
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-25
Filing Date
2024-07-18
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing base interpretation technology has poor base recognition quality, high error rate, large data processing volume, and large storage resource consumption due to factors such as uneven copy number during DNB preparation and different luminous efficiency during sequencing.

Method used

By obtaining the light intensity signal of each DNB in ​​different time dimensions, determining the light intensity distribution interval, and using distributed computing methods for base recognition, reducing data transmission and storage consumption, and improving identification accuracy and efficiency.

Benefits of technology

Fast and accurate base recognition is achieved, reducing the luminescence imbalance caused by the efficiency of DNB rolling ring, and reducing data processing pressure and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024106234_03072025_PF_FP_ABST
    Figure CN2024106234_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A base calling method and system, an analysis method and system, and an electronic device and a storage medium. The base calling method comprises: acquiring first light intensity data corresponding to each target pixel point in a target image under a set number of instances of test sequencing, wherein each target pixel point corresponds to a nucleic acid sequence cluster to be tested; on the basis of the first light intensity data, determining light intensity distribution intervals respectively corresponding to different types of bases; acquiring target light intensity data corresponding to each target pixel point in the target image under any instance of actual sequencing; and on the basis of the target light intensity data and the light intensity distribution intervals, obtaining a target base calling result under the instance of actual sequencing. Base calling can be implemented by using distributed computing, so that data transmission pressure and analysis pressure are reduced, and storage consumption is reduced; and the occurrence of luminescence unevenness caused by the efficiency of DNB rolling circle amplification is also reduced, thereby effectively improving the accuracy and processing efficiency of base calling.
Need to check novelty before this filing date? Find Prior Art

Description

Base recognition method and analysis method, system, electronic device and storage medium

[0001] This disclosure claims priority from Chinese patent application PCT / CN2023 / 141703, filed December 25, 2023. The entire text of the aforementioned Chinese patent application is incorporated herein by reference. Technical Field

[0002] The present disclosure relates to the field of gene sequencing technology, and in particular to a base recognition method and analysis method, system, electronic device and storage medium. Background Art

[0003] The current base calling technology is to divide the entire chip into different areas, consider all the DNB (nanoball sequencing) light intensities in each area using statistical methods, and then calculate the probability of obtaining the base.

[0004] However, due to the uneven copy number during the DNB preparation process, and the sequencing process, factors such as different luminescence efficiency caused by fluids are prone to occur, affecting the overall judgment of the bases, resulting in poor sequencing quality, increased error rate, and low base recognition quality value Q. In addition, all original images need to be stored for simultaneous processing and further converted into signals through complex processing. Each step consumes a lot of computing power and storage space, resulting in large data processing volume, high data processing cost and time, and high storage resource consumption.

[0005] Summary of the Invention

[0006] The technical problem to be solved by the present disclosure is to overcome the above-mentioned defects in the prior art, and the purpose is to provide a base recognition method and analysis method, system, electronic device and storage medium.

[0007] The present disclosure solves the above technical problems through the following technical solutions:

[0008] The present disclosure provides a base recognition method, which comprises:

[0009] Obtaining first light intensity data corresponding to each target pixel in the target image under a set number of test sequencings; wherein each target pixel corresponds to a nucleic acid sequence cluster to be tested;

[0010] Based on the first light intensity data, determining light intensity distribution intervals corresponding to different types of bases;

[0011] Obtaining target light intensity data corresponding to each target pixel in the target image under any actual sequencing condition;

[0012] The target base recognition result under the actual sequencing is obtained according to the target light intensity data and the light intensity distribution interval.

[0013] Preferably, the step of obtaining first light intensity data corresponding to each target pixel in the target image under a set number of test sequencings includes:

[0014] Acquire second light intensity data corresponding to each initial pixel point in the target image under the first number of test sequencings;

[0015] identifying valid pixel points among the initial pixel points according to a plurality of the second light intensity data;

[0016] The first light intensity data corresponding to each of the valid pixel points under the second number of test sequencings is obtained.

[0017] Preferably, the step of identifying effective pixel points among the initial pixel points based on a plurality of the second light intensity data includes:

[0018] If the first light intensity data represents the light intensity at the corresponding initial pixel point, and does not change under the first number of test sequencing, the corresponding initial pixel point is determined to be a non-valid pixel point; otherwise, the corresponding initial pixel point is determined to be the valid pixel point.

[0019] Preferably, the step of determining the light intensity distribution intervals corresponding to different types of bases based on the first light intensity data includes:

[0020] performing clustering processing on a plurality of the first light intensity data to obtain different initial distribution intervals;

[0021] Acquiring reference distribution data that characterizes the differences in light intensity distribution between different types of bases;

[0022] Based on the reference distribution data, different initial distribution intervals are distinguished to obtain the light intensity distribution intervals corresponding to different types of bases.

[0023] Preferably, the step of obtaining the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution interval includes:

[0024] The base type corresponding to the light intensity distribution interval in which the target light intensity data falls is used as the target base recognition result in the actual sequencing.

[0025] Preferably, the step of obtaining the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution interval includes:

[0026] If the target light intensity data does not fall into any of the light intensity distribution intervals, then calculating the distance between the target light intensity data and the centroid of each of the light intensity distribution intervals;

[0027] The base type corresponding to the light intensity distribution interval with the smallest distance is selected as the target base recognition result in the actual sequencing.

[0028] Preferably, after the step of obtaining the first light intensity data corresponding to each target pixel in the target image under the set number of test sequencings, the method further includes:

[0029] determining whether the first light intensity data is within a first preset light intensity range, and if so, adjusting the exposure time or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range;

[0030] and / or,

[0031] After the step of obtaining the first light intensity data corresponding to each target pixel in the target image under the set number of test sequencings, the method further includes:

[0032] Determine whether the light intensity signal value of the first light intensity data is less than a second preset light intensity range. If so, a preset signal amplification circuit is used to amplify the first light intensity data until it reaches the second preset light intensity range.

[0033] The present disclosure also provides a base recognition method, which comprises:

[0034] Acquiring first light intensity data corresponding to a target pixel point of a nucleic acid sequence cluster to be tested in a target image under a set number of test sequencings;

[0035] Based on the first light intensity data, determining light intensity distribution intervals corresponding to different types of bases;

[0036] Obtaining target light intensity data corresponding to a target pixel point of the target nucleic acid sequence cluster in the target image under any actual sequencing condition;

[0037] The target base recognition result under the actual sequencing is obtained according to the target light intensity data and the light intensity distribution interval.

[0038] Preferably, the step of determining the light intensity distribution intervals corresponding to different types of bases based on the first light intensity data includes:

[0039] performing clustering processing on a plurality of the first light intensity data to obtain different initial distribution intervals;

[0040] Acquiring reference distribution data that characterizes the differences in light intensity distribution between different types of bases;

[0041] Based on the reference distribution data, different initial distribution intervals are distinguished to obtain the light intensity distribution intervals corresponding to different types of bases.

[0042] Preferably, the step of obtaining the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution interval includes:

[0043] The base type corresponding to the light intensity distribution interval in which the target light intensity data falls is used as the target base recognition result in the actual sequencing.

[0044] Preferably, the step of obtaining the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution interval includes:

[0045] If the target light intensity data does not fall into any of the light intensity distribution intervals, then calculating the distance between the target light intensity data and the centroid of each of the light intensity distribution intervals;

[0046] The base type corresponding to the light intensity distribution interval with the smallest distance is selected as the target base recognition result in the actual sequencing.

[0047] Preferably, after the step of obtaining the first light intensity data corresponding to the target pixel point of the nucleic acid sequence cluster to be tested in the target image under the set number of test sequencing, the method further includes:

[0048] determining whether the first light intensity data is within a first preset light intensity range, and if so, adjusting the exposure time or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range;

[0049] and / or,

[0050] After the step of obtaining first light intensity data corresponding to target pixel points of the nucleic acid sequence cluster to be tested in the target image under the set number of test sequencings, the method further includes:

[0051] Determine whether the light intensity signal value of the first light intensity data is less than a second preset light intensity range. If so, a preset signal amplification circuit is used to amplify the first light intensity data until it reaches the second preset light intensity range.

[0052] The present disclosure also provides a base analysis method, comprising:

[0053] The target base recognition result obtained based on the above base recognition method is analyzed and processed to obtain a base analysis result.

[0054] The present disclosure also provides a base recognition system, comprising:

[0055] A first light intensity data acquisition module is used to obtain first light intensity data corresponding to each target pixel in the target image under a set number of test sequencings; wherein each target pixel corresponds to a DNB nanosphere;

[0056] a light intensity distribution interval determination module, configured to determine light intensity distribution intervals corresponding to different types of bases based on the first light intensity data;

[0057] A target light intensity data acquisition module is used to acquire target light intensity data corresponding to each target pixel point in the target image under any actual sequencing condition;

[0058] The base recognition module is used to obtain the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution interval.

[0059] Preferably, the first light intensity data acquisition module includes:

[0060] A second light intensity data acquisition unit is configured to acquire second light intensity data corresponding to each initial pixel point in the target image under the first number of test sequencings;

[0061] An effective pixel point identification unit, configured to identify effective pixel points among the initial pixel points based on a plurality of the second light intensity data;

[0062] The first light intensity data acquisition unit is used to acquire the first light intensity data corresponding to each of the valid pixel points under the second number of test sequencings.

[0063] Preferably, the effective pixel point identification unit is used to determine that the corresponding initial pixel point is a non-effective pixel point if the first light intensity data represents the light intensity at the corresponding initial pixel point, and it does not change under the first number of test sequencings; otherwise, the corresponding initial pixel point is determined to be a effective pixel point.

[0064] Preferably, the light intensity distribution interval determination module includes:

[0065] an initial distribution interval acquisition unit, configured to perform clustering processing on a plurality of the first light intensity data to obtain different initial distribution intervals;

[0066] A reference distribution data acquisition unit, configured to acquire reference distribution data characterizing the difference in light intensity distribution between different types of bases;

[0067] The light intensity distribution interval acquisition unit is used to distinguish different initial distribution intervals based on the reference distribution data to obtain the light intensity distribution intervals corresponding to different types of bases.

[0068] Preferably, the base recognition module is used to use the base type corresponding to the light intensity distribution interval in which the target light intensity data falls as the target base recognition result in the actual sequencing.

[0069] The base recognition module is configured to calculate the distance between the target light intensity data and the centroid of each of the light intensity distribution intervals if the target light intensity data does not fall within any of the light intensity distribution intervals;

[0070] The base type corresponding to the light intensity distribution interval with the smallest distance is selected as the target base recognition result in the actual sequencing.

[0071] Preferably, the base recognition system further comprises:

[0072] The first judgment module is used to judge whether the first light intensity data is within a first preset light intensity range. If so, adjust the exposure time and / or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range.

[0073] and / or,

[0074] The second judgment module is used to judge whether the light intensity signal value of the first light intensity data is less than the second preset light intensity range. If so, a preset signal amplification circuit is used to amplify the first light intensity data until it reaches the second preset light intensity range.

[0075] The present disclosure also provides a base recognition system, comprising:

[0076] A first light intensity data acquisition module is used to acquire first light intensity data corresponding to a target pixel point of a nucleic acid sequence cluster to be tested in a target image under a set number of test sequencings;

[0077] a light intensity distribution interval determination module, configured to determine light intensity distribution intervals corresponding to different types of bases based on the first light intensity data;

[0078] A target light intensity data acquisition module is used to acquire target light intensity data corresponding to target pixel points of the target nucleic acid sequence cluster in the target image under any actual sequencing conditions;

[0079] The base recognition module is used to obtain the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution interval.

[0080] Preferably, the light intensity distribution interval determination module includes:

[0081] an initial distribution interval acquisition unit, configured to perform clustering processing on a plurality of the first light intensity data to obtain different initial distribution intervals;

[0082] A reference distribution data acquisition unit, configured to acquire reference distribution data characterizing the difference in light intensity distribution between different types of bases;

[0083] The light intensity distribution interval acquisition unit is used to distinguish different initial distribution intervals based on the reference distribution data to obtain the light intensity distribution intervals corresponding to different types of bases.

[0084] Preferably, the base recognition module is used to use the base type corresponding to the light intensity distribution interval in which the target light intensity data falls as the target base recognition result in the actual sequencing.

[0085] The base recognition module is configured to calculate the distance between the target light intensity data and the centroid of each of the light intensity distribution intervals if the target light intensity data does not fall within any of the light intensity distribution intervals;

[0086] The base type corresponding to the light intensity distribution interval with the smallest distance is selected as the target base recognition result in the actual sequencing.

[0087] Preferably, the base recognition system further comprises:

[0088] The first judgment module is used to judge whether the first light intensity data is within a first preset light intensity range. If so, adjust the exposure time and / or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range.

[0089] and / or,

[0090] The second judgment module is used to judge whether the light intensity signal value of the first light intensity data is less than the second preset light intensity range. If so, a preset signal amplification circuit is used to amplify the first light intensity data until it reaches the second preset light intensity range.

[0091] The present disclosure also provides a base analysis system, comprising:

[0092] The base analysis module is used to analyze and process the target base recognition result obtained based on the above-mentioned base recognition system to obtain a base analysis result.

[0093] The present disclosure also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned base recognition method or the above-mentioned base analysis method when executing the computer program.

[0094] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-mentioned base recognition method or the above-mentioned base analysis method.

[0095] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present disclosure.

[0096] The positive progress of this disclosure is:

[0097] In the present disclosure, the existing method of analyzing all or part of the DNBs in the same time dimension as a whole is no longer adopted. Instead, a DNB is processed as an independent individual, and the changes in the light intensity signal emitted by each DNB in ​​the time dimension (i.e., the cycle in the sequencing process) are observed. The relationship between the light intensity change and the base is further analyzed, so that base recognition can be achieved by using distributed computing, reducing the data transmission pressure and analysis pressure, as well as reducing storage consumption, so that base recognition can be performed quickly; at the same time, it also reduces the occurrence of luminescence imbalance caused by DNB rolling circle efficiency, thereby effectively improving the accuracy of base recognition and processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] FIG1 is a flow chart of the base recognition method of Example 1 of the present disclosure.

[0099] FIG2 is a flow chart of the base recognition method of Example 2 of the present disclosure.

[0100] FIG3 is a schematic diagram of the light intensity distribution range of Example 2 of the present disclosure.

[0101] FIG4 is a schematic diagram of the light intensity distribution under each sequencing in Example 2 of the present disclosure.

[0102] FIG5 is a flow chart of the base recognition method of Example 3 of the present disclosure.

[0103] FIG6 is a flow chart of the base analysis method of Example 5 of the present disclosure.

[0104] FIG7 is a flow chart of the base recognition system of Example 6 of the present disclosure.

[0105] FIG8 is a flow chart of the base recognition system of Example 7 of the present disclosure.

[0106] FIG9 is a flow chart of the base analysis system according to the tenth embodiment of the present disclosure.

[0107] FIG10 is a schematic structural diagram of an electronic device in Embodiment 11 of the present disclosure. DETAILED DESCRIPTION

[0108] The present disclosure is further illustrated below by way of examples, but the present disclosure is not limited to the scope of the examples.

[0109] Example 1

[0110] As shown in FIG1 , the base recognition method of this embodiment includes:

[0111] S101. Obtain first light intensity data corresponding to each target pixel in the target image under a set number of test sequencings; wherein each target pixel corresponds to a nucleic acid sequence cluster to be tested; the set number can be 4, 8, 10, 12, 16, 32, etc., and is determined or adjusted according to specific actual needs.

[0112] Specifically, a nucleic acid sequence cluster is a group of similar or identical nucleotide sequences or DNA strands; for example, a nucleic acid sequence cluster can be an amplified oligonucleotide or polynucleotide having the same or similar sequence. In embodiments, during a nucleic acid sequencing cycle, the cluster can include a DNB that is affixed to a reaction site and / or reaction chamber of a sequencing chip.

[0113] Specifically, when directly using semiconductor sequencing (protein sequencing, nanopore and other single-molecule sequencing, etc.), each pixel corresponds to a DNB, and the position of each DNB is fixed and known. Each DNB to be tested is analyzed and processed as a separate individual, and the light intensity signal emitted by each DNB in ​​different time dimensions (i.e., different cycles in the test sequencing process) is obtained to analyze the light intensity changes based on these light intensity signals.

[0114] S102, determining light intensity distribution intervals corresponding to different types of bases based on the first light intensity data;

[0115] Among them, different types of bases may include base A, base G, base C, base T, etc.

[0116] S103, obtaining target light intensity data corresponding to each target pixel in the target image under any actual sequencing condition;

[0117] S104. Obtain target base recognition results under actual sequencing based on the target light intensity data and the light intensity distribution range.

[0118] In this embodiment, the existing method of analyzing all or part of the DNBs in the same time dimension as a whole is no longer adopted. Instead, a DNB is processed as an independent individual, and the changes in the light intensity signal emitted by each DNB in ​​the time dimension are observed. The relationship between the light intensity change and the base is further analyzed, so that base recognition can be achieved by using distributed computing, which reduces the pressure of data transmission and analysis, as well as the storage consumption, so that base recognition can be performed quickly; at the same time, it also reduces the occurrence of uneven luminescence caused by the DNB rolling circle efficiency, thereby effectively improving the accuracy of base recognition and processing efficiency.

[0119] Example 2

[0120] As shown in FIG2 , the base recognition method of this embodiment is a further improvement of Example 1. Specifically:

[0121] In one feasible solution, step S101 includes:

[0122] S1011, obtaining second light intensity data corresponding to each initial pixel point in the target image under the first number of test sequencings;

[0123] S1012, identifying valid pixel points among the initial pixel points according to the plurality of second light intensity data;

[0124] S1013. Obtain first light intensity data corresponding to each valid pixel point under a second number of test sequencings.

[0125] In this scheme, the first number of test sequencings is the sequencing cycles corresponding to the first N (such as N = 10) test sequencings, and the second number of test sequencings is the sequencing cycles corresponding to the first M (such as M = 100) test sequencings; that is, by obtaining and storing the light intensity data corresponding to each initial pixel point in the target image under the first N test sequencings, the pixel points that meet the preset conditions among all the initial pixel points are extracted as valid pixel points, and the remaining non-valid pixel points are discarded and not analyzed in the subsequent process. That is, by eliminating bad pixels, the accuracy of the subsequent recognition results is guaranteed, and the amount of data analyzed is also reduced, further improving the processing efficiency of the base recognition process.

[0126] In addition, optionally, for any pixel point, the first light intensity data corresponding to each pixel point in the target image is obtained based on the data corresponding to several test sequencings before the pixel point, and then the light intensity distribution intervals corresponding to different types of bases are determined, so as to obtain the target light intensity data corresponding to each pixel point in the target image under the actual sequencing after the pixel point, and finally obtain the target base recognition result under the actual sequencing, so that the data between each pixel point are not mixed and crossed, thereby further improving the accuracy of the base recognition result.

[0127] In one feasible solution, step S1012 includes:

[0128] If the first light intensity data represents the light intensity at the corresponding initial pixel point and does not change under the first number of test sequencings, the corresponding initial pixel point is determined to be a non-valid pixel point; otherwise, the corresponding initial pixel point is determined to be a valid pixel point.

[0129] Specifically, under these actual sequencing conditions, the light intensity data of different pixel points may have different reflection conditions, some of which never emit light, some emit light under partial sequencing, some do not emit light under some sequencing, and some always emit light; in this scheme, the pixel points that never emit light or always emit light under the first number of actual sequencing conditions are eliminated, and the remaining pixel points are regarded as valid pixel points.

[0130] In this scheme, based on the light intensity changes of the initial pixels in the target image in the time dimension, the pixels whose light intensity has not changed and the pixels whose light intensity has changed under multiple continuous test sequencing are identified, so as to eliminate the pixels whose light intensity has not changed, and only analyze and process the pixels whose light intensity has changed, so as to ensure the accuracy and processing speed of the subsequent recognition results.

[0131] In one feasible solution, step S102 includes:

[0132] S1021, performing clustering processing on a plurality of first light intensity data to obtain different initial distribution intervals;

[0133] The first light intensity data corresponding to each valid pixel point under M (eg, M=100) test sequencings are clustered to obtain 4 clusters, each of which corresponds to an initial distribution interval.

[0134] Specifically, the first light intensity data refers to the light intensity value of a pixel point in two or four channels; for example, for the case of two channels, the first light intensity data of the i-th cycle is (a, b); for the case of four channels, the first light intensity data of the i-th cycle is (a, b, c, d).

[0135] In addition, there is no specific requirement for the specific clustering algorithm to be used, as long as it can achieve the corresponding clustering purpose.

[0136] S1022, obtaining reference distribution data characterizing the difference in light intensity distribution between different types of bases;

[0137] The differences in the light intensity between different types of bases are well known to those skilled in the art and will not be elaborated here.

[0138] S1023. Based on the reference distribution data, different initial distribution intervals are distinguished to obtain light intensity distribution intervals corresponding to different types of bases.

[0139] For example, referring to the table below, clustering is performed for the first light intensity data corresponding to each valid pixel point under M (M is 100) test sequencing, and the light intensity distribution intervals corresponding to different types of bases are finally determined:

[0140] Among them, under certain sequencing conditions, if the light intensity range of channel X is 49-150 and the light intensity range of channel Y is 55-410, the corresponding base type is base A; if the light intensity range of channel X is 148-251 and the light intensity range of channel Y is 310-690, the corresponding base type is base G; if the light intensity range of channel X is 275-370 and the light intensity range of channel Y is 0-350, the corresponding base type is base C; if the light intensity range of channel X is 40-140 and the light intensity range of channel Y is 550-910, the corresponding base type is base T.

[0141] As shown in Figure 3, the horizontal axis corresponds to the light intensity data corresponding to channel X, and the vertical axis corresponds to the light intensity data corresponding to channel Y; the four circular areas correspond to different types of bases A, G, C, and T, respectively; among them, each point in base A, G, C, and T represents the distribution of light intensity under the first 10 test sequencings.

[0142] As shown in Figure 4, the light intensity data corresponding to channel X and channel Y under each test sequencing are M (M is 100), showing the change of the light intensity signal over time (i.e., different test sequencing). Among them, each column in the upper half corresponds to the light intensity data of channel X under the test sequencing, and each column in the lower half corresponds to the light intensity data of channel Y under the test sequencing.

[0143] Specifically, different light intensity distribution intervals can be obtained based on a standard sample with a fixed copy number; of course, different light intensity distribution intervals can also be determined based on actual conditions by using a standard sample with a smaller copy number.

[0144] In this scheme, by clustering a number of first light intensity data, automatic analysis of the light intensity distribution under different test sequencing is achieved. By comparing with the reference distribution data of the light intensity distribution differences between different types of bases, the light intensity distribution intervals corresponding to different bases are automatically determined. The correlation relationship between the light intensity distribution of each DNB and the base in different time dimensions is constructed, ensuring the feasibility, reliability and efficiency of the base recognition process.

[0145] In one feasible solution, step S104 includes:

[0146] S1041. The base type corresponding to the light intensity distribution interval in which the target light intensity data falls is used as the target base recognition result in actual sequencing.

[0147] In this scheme, for any actual sequencing, the target light intensity data under the sequencing is obtained, including the first target light intensity data of channel X and the second target light intensity data of channel Y; the light intensity distribution interval to which the target light intensity data belongs is automatically determined. Once it falls into the corresponding light intensity distribution interval, the base type corresponding to the light intensity distribution interval is used as the base recognition result under the current sequencing, thereby achieving efficiency and accuracy in base recognition.

[0148] In one feasible solution, step S104 includes:

[0149] If the target light intensity data does not fall into any light intensity distribution interval, the distance between the target light intensity data and the centroid of each light intensity distribution interval is calculated;

[0150] The base type corresponding to the light intensity distribution interval with the minimum distance is selected as the target base recognition result in actual sequencing.

[0151] In this scheme, considering that in most cases, the number M of test sequencing is taken as a smaller value, such as M is taken as 50, the obtained light intensity distribution range can be used, and the light intensity distribution range can also be determined based on the light intensity data under the condition of a smaller number M of test sequencing, thereby reducing the data processing volume to a certain extent and reducing the data transmission pressure and analysis pressure.

[0152] However, in some cases, the number M of test sequencing is too small, such as M taking a value of 50, which will lead to inaccurate determination of the light intensity distribution interval, resulting in the target light intensity data not belonging to any light intensity distribution interval. At this time, based on the target light intensity data under the sequencing, including the first target light intensity data of channel X and the second target light intensity data of channel Y, the distance between the target light intensity data and the center of mass of each light intensity distribution interval is calculated, and then included in the nearest light intensity distribution interval, and the base type corresponding to the light intensity distribution interval is automatically used as the base recognition result under the current sequencing, which realizes remedial treatment of abnormal situations, ensures that base recognition can be achieved for any actual sequencing, and can ensure the efficiency and accuracy of base recognition, thereby further optimizing the base recognition process.

[0153] In one feasible solution, after step S101, the following steps are further included:

[0154] Determine whether the first light intensity data is within a first preset light intensity range; if so, adjust the exposure time and / or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range.

[0155] In this solution, the acquired light intensity data is monitored in real time. If the intensity of the light intensity data is too strong and exceeds the first preset light intensity range, the working parameters of the corresponding light-emitting device, such as exposure time, exposure gain, etc., are automatically driven and controlled so that the acquired light intensity data meets the data analysis requirements, thereby ensuring the feasibility of subsequent base recognition and the accuracy of the final base recognition results.

[0156] In one feasible solution, after step S101, the following steps are further included:

[0157] It is determined whether the light intensity signal value of the first light intensity data is less than a second preset light intensity range. If so, a preset signal amplification circuit is used to amplify the first light intensity data until it reaches the second preset light intensity range.

[0158] The preset signal amplification circuit includes but is not limited to a logarithmic signal amplification circuit.

[0159] In this solution, the acquired light intensity data is monitored in real time. If the light intensity signal value is too small and exceeds the first preset light intensity range, the corresponding preset signal amplification circuit is automatically driven and controlled to amplify it, so as to amplify the difference between similar signals, so that the acquired light intensity data meets the data analysis requirements, thereby ensuring the feasibility of subsequent base recognition and the accuracy of the final base recognition result.

[0160] Example 3

[0161] As shown in FIG5 , the base recognition method of this embodiment includes:

[0162] S201, obtaining first light intensity data corresponding to a target pixel point of a nucleic acid sequence cluster to be tested in a target image under a set number of test sequencings;

[0163] Among them, based on the nucleic acid sequence cluster to be tested, the light intensity signals emitted by the target pixel points of the nucleic acid sequence cluster to be tested in the image in different time dimensions (i.e., cycles in different test sequencing processes) are obtained to analyze the light intensity changes based on these light intensity signals.

[0164] S202, determining light intensity distribution intervals corresponding to different types of bases based on the first light intensity data;

[0165] Among them, different types of bases may include base A, base G, base C, base T, etc.

[0166] S203, obtaining target light intensity data corresponding to a target pixel point of a target nucleic acid sequence cluster in a target image under any actual sequencing condition;

[0167] S204: Obtain target base recognition results in actual sequencing based on the target light intensity data and the light intensity distribution range.

[0168] In this embodiment, the existing method of analyzing all or part of the DNBs in the same time dimension as a whole is no longer adopted. Instead, the nucleic acid sequence cluster to be tested is used as the basis, and the light intensity signals emitted by the target pixel points of the nucleic acid sequence cluster to be tested in the image in different time dimensions are obtained. The light intensity changes are analyzed based on these light intensity signals, and the relationship between the light intensity changes and the bases is further analyzed. In this way, base recognition can be achieved by using distributed computing, which reduces the pressure of data transmission and analysis, as well as storage consumption, so that base recognition can be performed quickly, and the accuracy and processing efficiency of base recognition are effectively improved.

[0169] Example 4

[0170] The base recognition method of this embodiment is a further improvement of Example 3. Specifically:

[0171] In one feasible solution, step S202 includes:

[0172] performing clustering processing on a plurality of first light intensity data to obtain different initial distribution intervals;

[0173] The first light intensity data corresponding to each valid pixel point under M (eg, M=100) test sequencings are clustered to obtain 4 clusters, each of which corresponds to an initial distribution interval.

[0174] Specifically, the first light intensity data refers to the light intensity value of a pixel point in two or four channels; for example, for the case of two channels, the first light intensity data of the i-th cycle is (a, b); for the case of four channels, the first light intensity data of the i-th cycle is (a, b, c, d).

[0175] In addition, there is no specific requirement for the specific clustering algorithm to be used, as long as it can achieve the corresponding clustering purpose.

[0176] Acquiring reference distribution data that characterizes the differences in light intensity distribution between different types of bases;

[0177] The differences in the light intensity between different types of bases are well known to those skilled in the art and will not be elaborated here.

[0178] Based on the reference distribution data, different initial distribution intervals are distinguished to obtain the light intensity distribution intervals corresponding to different types of bases.

[0179] In one feasible solution, step S204 includes:

[0180] The base type corresponding to the light intensity distribution interval in which the target light intensity data falls is used as the target base recognition result in actual sequencing.

[0181] In this scheme, for any actual sequencing, the target light intensity data under the sequencing is obtained, including the first target light intensity data of channel X and the second target light intensity data of channel Y; the light intensity distribution interval to which the target light intensity data belongs is automatically determined. Once it falls into the corresponding light intensity distribution interval, the base type corresponding to the light intensity distribution interval is used as the base recognition result under the current sequencing, thereby achieving efficiency and accuracy in base recognition.

[0182] In one feasible solution, step S204 includes:

[0183] If the target light intensity data does not fall into any light intensity distribution interval, the distance between the target light intensity data and the centroid of each light intensity distribution interval is calculated;

[0184] The base type corresponding to the light intensity distribution interval with the minimum distance is selected as the target base recognition result in actual sequencing.

[0185] In this scheme, considering that in most cases, the number M of test sequencing is taken as a smaller value, such as M is taken as 50, the obtained light intensity distribution range can be used, and the light intensity distribution range can also be determined based on the light intensity data under the condition of a smaller number M of test sequencing, thereby reducing the data processing volume to a certain extent and reducing the data transmission pressure and analysis pressure.

[0186] However, in some cases, the number M of test sequencing is too small, such as M taking a value of 50, which will lead to inaccurate determination of the light intensity distribution interval, resulting in the target light intensity data not belonging to any light intensity distribution interval. At this time, based on the target light intensity data under the sequencing, including the first target light intensity data of channel X and the second target light intensity data of channel Y, the distance between the target light intensity data and the center of mass of each light intensity distribution interval is calculated, and then included in the nearest light intensity distribution interval, and the base type corresponding to the light intensity distribution interval is automatically used as the base recognition result under the current sequencing, which realizes remedial treatment of abnormal situations, ensures that base recognition can be achieved for any actual sequencing, and can ensure the efficiency and accuracy of base recognition, thereby further optimizing the base recognition process.

[0187] In one feasible solution, after step S201, the following steps are further included:

[0188] Determine whether the first light intensity data is within a first preset light intensity range; if so, adjust the exposure time and / or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range.

[0189] In this solution, the acquired light intensity data is monitored in real time. If the intensity of the light intensity data is too strong and exceeds the first preset light intensity range, the working parameters of the corresponding light-emitting device, such as exposure time, exposure gain, etc., are automatically driven and controlled so that the acquired light intensity data meets the data analysis requirements, thereby ensuring the feasibility of subsequent base recognition and the accuracy of the final base recognition results.

[0190] In one feasible solution, after step S201, the following steps are further included:

[0191] It is determined whether the light intensity signal value of the first light intensity data is less than a second preset light intensity range. If so, a preset signal amplification circuit is used to amplify the first light intensity data until it reaches the second preset light intensity range.

[0192] The preset signal amplification circuit includes but is not limited to a logarithmic signal amplification circuit.

[0193] In this solution, the acquired light intensity data is monitored in real time. If the light intensity signal value is too small and exceeds the first preset light intensity range, the corresponding preset signal amplification circuit is automatically driven and controlled to amplify it, so as to amplify the difference between similar signals, so that the acquired light intensity data meets the data analysis requirements, thereby ensuring the feasibility of subsequent base recognition and the accuracy of the final base recognition result.

[0194] In addition, the steps in this embodiment are the same as those in Example 2, and the corresponding implementation processes and principles are also the same. Please refer to Example 2 for details and will not be repeated here.

[0195] Example 5

[0196] The base analysis method of this embodiment is implemented based on the base recognition method in any one of the above embodiments 1-4.

[0197] As shown in FIG6 , the base analysis method of this embodiment includes:

[0198] S301 , analyzing and processing the target base recognition result obtained by the base recognition method to obtain a base analysis result.

[0199] Wherein, step S301 includes:

[0200] Analyzing and processing the target base recognition results obtained by the base recognition method to calculate the corresponding base recognition quality value;

[0201] For example, if the target intensity data for the 11th sequence is used to identify the corresponding base type as base T, then the base call score is calculated by comparing the base call result with the overall signal variation pattern of the DNB. For example, if the target intensity data for the 11th sequence is 208 for channel X and 722 for channel Y, the base call quality (Q) is calculated to be 97%.

[0202] Based on the base calling quality value, the corresponding base analysis text is generated.

[0203] Specifically, the base analysis text is a FastQ base sequence file (a sequencing file) and a corresponding report.

[0204] In this solution, the target base recognition results obtained by the base recognition method are analyzed and processed to obtain base analysis results, and the corresponding base recognition quality value and base analysis text are output, thereby achieving the accuracy of the base analysis results and ensuring the comprehensiveness of the base analysis process.

[0205] Example 6

[0206] As shown in FIG7 , the base recognition system of this embodiment includes:

[0207] The first light intensity data acquisition module 1 is used to obtain the first light intensity data corresponding to each target pixel in the target image under a set number of test sequencings; wherein, each target pixel corresponds to a nucleic acid sequence cluster to be tested; the value of the set number can be 4, 8, 10, 12, 16, 32, etc., which is determined or adjusted according to specific actual needs.

[0208] Specifically, a nucleic acid sequence cluster is a group of similar or identical nucleotide sequences or DNA strands; for example, a nucleic acid sequence cluster can be an amplified oligonucleotide or polynucleotide having the same or similar sequence. In embodiments, during a nucleic acid sequencing cycle, the cluster can include a DNB that is affixed to a reaction site and / or reaction chamber of a sequencing chip.

[0209] Specifically, when directly using semiconductor sequencing (protein sequencing, nanopore and other single-molecule sequencing, etc.), each pixel corresponds to a DNB, and the position of each DNB is fixed and known. Each DNB to be tested is analyzed and processed as a separate individual, and the light intensity signal emitted by each DNB in ​​different time dimensions (i.e., different cycles in the test sequencing process) is obtained to analyze the light intensity changes based on these light intensity signals.

[0210] Light intensity distribution interval determination module 2, for determining light intensity distribution intervals corresponding to different types of bases based on the first light intensity data;

[0211] Among them, different types of bases may include base A, base G, base C, base T, etc.

[0212] The target light intensity data acquisition module 3 is used to obtain the target light intensity data corresponding to each target pixel in the target image under any actual sequencing condition;

[0213] The base recognition module 4 is used to obtain the target base recognition result under actual sequencing according to the target light intensity data and the light intensity distribution range.

[0214] In this embodiment, the existing method of analyzing all or part of the DNBs in the same time dimension as a whole is no longer adopted. Instead, a DNB is processed as an independent individual, and the changes in the light intensity signal emitted by each DNB in ​​the time dimension are observed. The relationship between the light intensity change and the base is further analyzed, so that base recognition can be achieved by using distributed computing, which reduces the pressure of data transmission and analysis, as well as the storage consumption, so that base recognition can be performed quickly; at the same time, it also reduces the occurrence of uneven luminescence caused by the DNB rolling circle efficiency, thereby effectively improving the accuracy of base recognition and processing efficiency.

[0215] Example 7

[0216] As shown in FIG8 , the base recognition system of this embodiment is a further improvement of Example 6. Specifically:

[0217] In one feasible solution, the first light intensity data acquisition module 1 includes:

[0218] A second light intensity data acquisition unit 5 is configured to acquire second light intensity data corresponding to each initial pixel point in the target image under the first number of test sequencings;

[0219] An effective pixel point identification unit 6 is used to identify effective pixel points in the initial pixel points according to a plurality of second light intensity data;

[0220] The first light intensity data acquisition unit 7 is used to acquire first light intensity data corresponding to each valid pixel point under the second number of test sequencing.

[0221] In this scheme, the first number of test sequencings is the sequencing cycles corresponding to the first N (such as N = 10) test sequencings, and the second number of test sequencings is the sequencing cycles corresponding to the first M (such as M = 100) test sequencings; that is, by obtaining and storing the light intensity data corresponding to each initial pixel point in the target image under the first N test sequencings, the pixel points that meet the preset conditions among all the initial pixel points are extracted as valid pixel points, and the remaining non-valid pixel points are discarded and not analyzed in the subsequent process. That is, by eliminating bad pixels, the accuracy of the subsequent recognition results is guaranteed, and the amount of data analyzed is also reduced, further improving the processing efficiency of the base recognition process.

[0222] In addition, optionally, for any pixel point, the first light intensity data corresponding to each pixel point in the target image is obtained based on the data corresponding to several test sequencings before the pixel point, and then the light intensity distribution intervals corresponding to different types of bases are determined, so as to obtain the target light intensity data corresponding to each pixel point in the target image under the actual sequencing after the pixel point, and finally obtain the target base recognition result under the actual sequencing, so that the data between each pixel point are not mixed and crossed, thereby further improving the accuracy of the base recognition result.

[0223] In one feasible solution, the effective pixel point identification unit 6 is used to determine that the corresponding initial pixel point is a non-effective pixel point if the first light intensity data represents the light intensity at the corresponding initial pixel point, and it does not change under a first number of test sequencings; otherwise, the corresponding initial pixel point is determined to be a effective pixel point.

[0224] Specifically, under these actual sequencing conditions, the light intensity data of different pixel points may have different reflection conditions, some of which never emit light, some emit light under partial sequencing, some do not emit light under some sequencing, and some always emit light; in this scheme, the pixel points that never emit light or always emit light under the first number of actual sequencing conditions are eliminated, and the remaining pixel points are regarded as valid pixel points.

[0225] In this scheme, based on the light intensity changes of the initial pixels in the target image in the time dimension, the pixels whose light intensity has not changed and the pixels whose light intensity has changed under multiple continuous test sequencing are identified, so as to eliminate the pixels whose light intensity has not changed, and only analyze and process the pixels whose light intensity has changed, so as to ensure the accuracy and processing speed of the subsequent recognition results.

[0226] In one feasible solution, the light intensity distribution interval determination module 2 includes:

[0227] An initial distribution interval acquisition unit 8 is used to perform clustering processing on a plurality of first light intensity data to obtain different initial distribution intervals;

[0228] The first light intensity data corresponding to each valid pixel point under M (eg, M=100) test sequencings are clustered to obtain 4 clusters, each of which corresponds to an initial distribution interval.

[0229] Specifically, the first light intensity data refers to the light intensity value of a pixel point in two or four channels; for example, for the case of two channels, the first light intensity data of the i-th cycle is (a, b); for the case of four channels, the first light intensity data of the i-th cycle is (a, b, c, d).

[0230] In addition, there is no specific requirement for the specific clustering algorithm to be used, as long as it can achieve the corresponding clustering purpose.

[0231] A reference distribution data acquisition unit 9 is used to acquire reference distribution data characterizing the difference in light intensity distribution between different types of bases;

[0232] The differences in the light intensity between different types of bases are well known to those skilled in the art and will not be elaborated here.

[0233] The light intensity distribution interval acquisition unit 10 is used to distinguish different initial distribution intervals based on the reference distribution data to obtain light intensity distribution intervals corresponding to different types of bases.

[0234] For example, referring to the table below, clustering is performed for the first light intensity data corresponding to each valid pixel point under M (M is 100) test sequencing, and the light intensity distribution intervals corresponding to different types of bases are finally determined:

[0235] Among them, under certain sequencing conditions, if the light intensity range of channel X is 49-150 and the light intensity range of channel Y is 55-410, the corresponding base type is base A; if the light intensity range of channel X is 148-251 and the light intensity range of channel Y is 310-690, the corresponding base type is base G; if the light intensity range of channel X is 275-370 and the light intensity range of channel Y is 0-350, the corresponding base type is base C; if the light intensity range of channel X is 40-140 and the light intensity range of channel Y is 550-910, the corresponding base type is base T.

[0236] As shown in Figure 3, the horizontal axis corresponds to the light intensity data corresponding to channel X, and the vertical axis corresponds to the light intensity data corresponding to channel Y; the four circular areas correspond to different types of bases A, G, C, and T, respectively; among them, each point in base A, G, C, and T represents the distribution of light intensity under the first 10 test sequencings.

[0237] As shown in Figure 4, the light intensity data corresponding to channel X and channel Y under each test sequencing are M (M is 100), showing the change of the light intensity signal over time (i.e., different test sequencing). Among them, each column in the upper half corresponds to the light intensity data of channel X under the test sequencing, and each column in the lower half corresponds to the light intensity data of channel Y under the test sequencing.

[0238] Specifically, different light intensity distribution intervals can be obtained based on a standard sample with a fixed copy number; of course, different light intensity distribution intervals can also be determined based on actual conditions by using a standard sample with a smaller copy number.

[0239] In this scheme, by clustering a number of first light intensity data, automatic analysis of the light intensity distribution under different test sequencing is achieved. By comparing with the reference distribution data of the light intensity distribution differences between different types of bases, the light intensity distribution intervals corresponding to different bases are automatically determined. The correlation relationship between the light intensity distribution of each DNB and the base in different time dimensions is constructed, ensuring the feasibility, reliability and efficiency of the base recognition process.

[0240] In an implementable solution, the base recognition module 4 is configured to use the base type corresponding to the light intensity distribution interval into which the target light intensity data falls as the target base recognition result in actual sequencing.

[0241] In this scheme, for any actual sequencing, the target light intensity data under the sequencing is obtained, including the first target light intensity data of channel X and the second target light intensity data of channel Y; the light intensity distribution interval to which the target light intensity data belongs is automatically determined. Once it falls into the corresponding light intensity distribution interval, the base type corresponding to the light intensity distribution interval is used as the base recognition result under the current sequencing, thereby achieving efficiency and accuracy in base recognition.

[0242] In one feasible solution, the base recognition module 4 is configured to calculate the distance between the target light intensity data and the centroid of each light intensity distribution interval if the target light intensity data does not fall into any light intensity distribution interval;

[0243] The base type corresponding to the light intensity distribution interval with the minimum distance is selected as the target base recognition result in actual sequencing.

[0244] In this scheme, considering that in most cases, the number M of test sequencing is taken as a smaller value, such as M is taken as 50, the obtained light intensity distribution range can be used, and the light intensity distribution range can also be determined based on the light intensity data under the condition of a smaller number M of test sequencing, thereby reducing the data processing volume to a certain extent and reducing the data transmission pressure and analysis pressure.

[0245] However, in some cases, the number M of test sequencing is too small, such as M taking a value of 50, which will lead to inaccurate determination of the light intensity distribution interval, resulting in the target light intensity data not belonging to any light intensity distribution interval. At this time, based on the target light intensity data under the sequencing, including the first target light intensity data of channel X and the second target light intensity data of channel Y, the distance between the target light intensity data and the center of mass of each light intensity distribution interval is calculated, and then included in the nearest light intensity distribution interval, and the base type corresponding to the light intensity distribution interval is automatically used as the base recognition result under the current sequencing, which realizes remedial treatment of abnormal situations, ensures that base recognition can be achieved for any actual sequencing, and can ensure the efficiency and accuracy of base recognition, thereby further optimizing the base recognition process.

[0246] In one feasible solution, the base recognition system further comprises:

[0247] The first judgment module 11 is used to judge whether the first light intensity data is within a first preset light intensity range, and if so, adjust the exposure time and / or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range.

[0248] In this solution, the acquired light intensity data is monitored in real time. If the intensity of the light intensity data is too strong and exceeds the first preset light intensity range, the working parameters of the corresponding light-emitting device, such as exposure time, exposure gain, etc., are automatically driven and controlled so that the acquired light intensity data meets the data analysis requirements, thereby ensuring the feasibility of subsequent base recognition and the accuracy of the final base recognition results.

[0249] In one feasible solution, the base recognition system further comprises:

[0250] The second judgment module 12 is used to judge whether the light intensity signal value of the first light intensity data is less than the second preset light intensity range. If so, a preset signal amplification circuit is used to amplify the first light intensity data until it reaches the second preset light intensity range.

[0251] The preset signal amplification circuit includes but is not limited to a logarithmic signal amplification circuit.

[0252] In this solution, the acquired light intensity data is monitored in real time. If the light intensity signal value is too small and exceeds the first preset light intensity range, the corresponding preset signal amplification circuit is automatically driven and controlled to amplify it, so as to amplify the difference between similar signals, so that the acquired light intensity data meets the data analysis requirements, thereby ensuring the feasibility of subsequent base recognition and the accuracy of the final base recognition result.

[0253] Example 8

[0254] The base recognition system of this embodiment includes:

[0255] A first light intensity data acquisition module is used to acquire first light intensity data corresponding to a target pixel point of a nucleic acid sequence cluster to be tested in a target image under a set number of test sequencings;

[0256] Among them, based on the nucleic acid sequence cluster to be tested, the light intensity signals emitted by the target pixel points of the nucleic acid sequence cluster to be tested in the image in different time dimensions (i.e., cycles in different test sequencing processes) are obtained to analyze the light intensity changes based on these light intensity signals.

[0257] A light intensity distribution interval determination module, configured to determine light intensity distribution intervals corresponding to different types of bases based on the first light intensity data;

[0258] Among them, different types of bases may include base A, base G, base C, base T, etc.

[0259] A target light intensity data acquisition module is used to obtain target light intensity data corresponding to target pixel points of a target nucleic acid sequence cluster in a target image under any actual sequencing condition;

[0260] The base recognition module is used to obtain the target base recognition results under actual sequencing based on the target light intensity data and light intensity distribution range.

[0261] Among them, this embodiment is a system corresponding to the base recognition method of Example 3, and the corresponding implementation process and principles are the same. Please refer to Example 3 for details and will not be repeated here.

[0262] In this embodiment, the existing method of analyzing all or part of the DNBs in the same time dimension as a whole is no longer adopted. Instead, the nucleic acid sequence cluster to be tested is used as the basis, and the light intensity signals emitted by the target pixel points of the nucleic acid sequence cluster to be tested in the image in different time dimensions are obtained. The light intensity changes are analyzed based on these light intensity signals, and the relationship between the light intensity changes and the bases is further analyzed. In this way, base recognition can be achieved by using distributed computing, which reduces the pressure of data transmission and analysis, as well as storage consumption, so that base recognition can be performed quickly, and the accuracy and processing efficiency of base recognition are effectively improved.

[0263] Example 9

[0264] The base recognition system of this embodiment is a further improvement of Example 8. Specifically:

[0265] In one feasible solution, the light intensity distribution interval determination module includes:

[0266] An initial distribution interval acquisition unit, configured to perform clustering processing on a plurality of first light intensity data to obtain different initial distribution intervals;

[0267] A reference distribution data acquisition unit, configured to acquire reference distribution data characterizing the difference in light intensity distribution between different types of bases;

[0268] The light intensity distribution interval acquisition unit is used to distinguish different initial distribution intervals based on the reference distribution data to obtain light intensity distribution intervals corresponding to different types of bases.

[0269] In one feasible solution, the base recognition module is used to use the base type corresponding to the light intensity distribution interval in which the target light intensity data falls as the target base recognition result in actual sequencing.

[0270] The base recognition module is used to calculate the distance between the target light intensity data and the centroid of each light intensity distribution interval if the target light intensity data does not fall into any light intensity distribution interval;

[0271] The base type corresponding to the light intensity distribution interval with the minimum distance is selected as the target base recognition result in actual sequencing.

[0272] In one feasible solution, the base recognition system further comprises:

[0273] The first judgment module is used to judge whether the first light intensity data is within a first preset light intensity range. If so, adjust the exposure time and / or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range.

[0274] In one feasible solution, the second judgment module is used to judge whether the light intensity signal value of the first light intensity data is less than a second preset light intensity range, and if so, amplify the first light intensity data using a preset signal amplification circuit until it reaches the second preset light intensity range;

[0275] The preset signal amplification circuit includes but is not limited to a logarithmic signal amplification circuit.

[0276] Among them, this embodiment is a system corresponding to the base recognition method of Example 4, and the corresponding implementation process and principles are the same. Please refer to Example 4 for details and will not be repeated here.

[0277] Example 10

[0278] The base analysis system of this embodiment is implemented based on the base recognition system in any one of the above-mentioned embodiments 6-9.

[0279] As shown in FIG9 , the base analysis system of this embodiment includes:

[0280] The base analysis module 13 is used to analyze and process the target base recognition result obtained by the base recognition method to obtain a base analysis result.

[0281] The base analysis module 13 includes:

[0282] The quality value calculation unit 14 is used to analyze and process the target base recognition result obtained by the base recognition method and calculate the corresponding base recognition quality value;

[0283] For example, if the target intensity data for the 11th sequence is used to identify the corresponding base type as base T, then the base call score is calculated by comparing the base call result with the overall signal variation pattern of the DNB. For example, if the target intensity data for the 11th sequence is 208 for channel X and 722 for channel Y, the base call quality (Q) is calculated to be 97%.

[0284] The analysis text generating unit 15 is configured to generate corresponding base analysis text based on the base recognition quality value.

[0285] Specifically, the base analysis text is a FastQ base sequence file and a corresponding report.

[0286] In this solution, the target base recognition results obtained by the base recognition method are analyzed and processed to obtain base analysis results, and the corresponding base recognition quality value and base analysis text are output, thereby achieving the accuracy of the base analysis results and ensuring the comprehensiveness of the base analysis process.

[0287] Example 11

[0288] Figure 10 is a schematic diagram of the structure of an electronic device provided in Example 7 of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable by the processor. When the processor executes the program, the method described in the above embodiments is implemented. The electronic device 30 shown in Figure 10 is merely an example and should not limit the functionality or scope of application of the embodiments of the present disclosure.

[0289] As shown in FIG10 , the electronic device 30 may be a general-purpose computing device, such as a server device. Components of the electronic device 30 may include, but are not limited to, the at least one processor 31, the at least one memory 32, and a bus 33 connecting various system components (including the memory 32 and the processor 31).

[0290] The bus 33 includes a data bus, an address bus, and a control bus.

[0291] The memory 32 may include a volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322 , and may further include a read-only memory (ROM) 323 .

[0292] The memory 32 may also include a program / utility 325 having a set (at least one) of program modules 324, such program modules 324 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0293] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32, such as the method in the above-mentioned embodiment of the present disclosure.

[0294] The electronic device 30 can also communicate with one or more external devices 34 (e.g., a keyboard, pointing device, etc.). This communication can occur via an input / output (I / O) interface 35. Furthermore, the model-generating device 30 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 36. As shown in FIG10 , the network adapter 36 communicates with other modules of the model-generating device 30 via a bus 33. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the model-generating device 30, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.

[0295] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0296] Example 12

[0297] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method in the above embodiment are implemented.

[0298] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0299] In a possible implementation manner, the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps in the method of the above embodiment.

[0300] The program code for executing the present disclosure may be written in any combination of one or more programming languages, and the program code may be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on the remote device.

[0301] While specific embodiments of the present disclosure have been described above, those skilled in the art will appreciate that these are merely illustrative and that the scope of protection of the present disclosure is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present disclosure, and such changes and modifications are intended to fall within the scope of protection of the present disclosure.

Claims

1. A base recognition method, characterized in that, The base recognition method includes: Obtaining first light intensity data corresponding to each target pixel point in a target image under a set number of test sequencing; wherein, each of the target pixel points corresponds to a nucleic acid sequence cluster to be detected; Based on the first light intensity data, determining light intensity distribution intervals corresponding to different types of bases respectively; Obtaining target light intensity data corresponding to each of the target pixel points in the target image under any actual sequencing; According to the target light intensity data and the light intensity distribution intervals, obtaining a target base recognition result under the actual sequencing.

2. The base recognition method according to claim 1, wherein The step of obtaining first light intensity data corresponding to each target pixel point in a target image under a set number of test sequencing includes: Obtaining second light intensity data corresponding to each initial pixel point in the target image under a first number of the test sequencing; Identifying valid pixel points among the initial pixel points according to a plurality of the second light intensity data; Obtaining the first light intensity data corresponding to each of the valid pixel points under a second number of the test sequencing.

3. The base recognition method according to claim 2, wherein The step of identifying valid pixel points among the initial pixel points according to a plurality of the second light intensity data includes: If the light intensity represented by the first light intensity data does not change under the first number of the test sequencing at the corresponding initial pixel point, determining that the corresponding initial pixel point is an invalid pixel point; otherwise, determining that the corresponding initial pixel point is the valid pixel point.

4. The base recognition method according to at least one of claims 1-3, characterized in that, The step of determining light intensity distribution intervals corresponding to different types of bases respectively based on the first light intensity data includes: Performing clustering processing on a plurality of the first light intensity data to obtain different initial distribution intervals; Obtaining reference distribution data representing the light intensity distribution difference between different types of bases; Based on the reference distribution data, differentiating different ones of the initial distribution intervals to obtain the light intensity distribution intervals corresponding to different types of bases respectively.

5. The base recognition method according to at least one of claims 1-3, characterized in that, The step of obtaining a target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution intervals includes: Taking the base type corresponding to the light intensity distribution interval into which the target light intensity data falls as the target base recognition result under the actual sequencing under.

6. The base recognition method according to claim 1, wherein The step of obtaining a target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution intervals includes: If the target light intensity data does not fall into any of the light intensity distribution intervals, calculating the distances between the target light intensity data and the centroids of each of the light intensity distribution intervals; Selecting the base type corresponding to the light intensity distribution interval with the smallest of the distances as the target base recognition result under the actual sequencing.

7. The base recognition method according to at least one of claims 1-6, characterized in that, After the step of obtaining first light intensity data corresponding to each target pixel point in a target image under a set number of test sequencing, it further includes: Judging whether the first light intensity data is within a first preset light intensity range, and if so, adjusting the exposure time and / or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range; and / or After the step of obtaining the first light intensity data corresponding to each target pixel point in the target image under a set number of test sequencing, the following steps are further included: Determine whether the light intensity signal value of the first light intensity data is less than a second preset light intensity range. If so, use a preset signal amplification circuit to amplify the first light intensity data until it reaches the second preset light intensity range.

8. A base recognition method, characterized in that, The base recognition method includes: Obtain the first light intensity data corresponding to the target pixel points of the nucleic acid sequence cluster to be tested in the target image under a set number of test sequencing; Based on the first light intensity data, determine the light intensity distribution intervals corresponding to different types of bases; Obtain the target light intensity data corresponding to the target pixel points of the target nucleic acid sequence cluster in the target image under any actual sequencing; According to the target light intensity data and the light intensity distribution intervals, obtain the target base recognition result under the actual sequencing.

9. The base recognition method according to claim 8, wherein, The step of determining the light intensity distribution intervals corresponding to different types of bases based on the first light intensity data includes: Perform clustering processing on a plurality of the first light intensity data to obtain different initial distribution intervals; Obtain reference distribution data characterizing the light intensity distribution differences between different types of bases; Based on the reference distribution data, distinguish different initial distribution intervals to obtain the light intensity distribution intervals corresponding to different types of bases.

10. The base recognition method according to claim 8 or 9, characterized in that, The step of obtaining the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution intervals includes: Use the base type corresponding to the light intensity distribution interval into which the target light intensity data falls as the target base recognition result under the actual sequencing.

11. The base recognition method according to claim 8 or 9, characterized in that, The step of obtaining the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution intervals includes: If the target light intensity data does not fall into any of the light intensity distribution intervals, calculate the distances between the target light intensity data and the centroids of each light intensity distribution interval; Select the base type corresponding to the light intensity distribution interval with the minimum distance as the target base recognition result under the actual sequencing.

12. The base recognition method according to at least one of claims 8-11, characterized in that, After the step of obtaining the first light intensity data corresponding to the target pixel points of the nucleic acid sequence cluster to be tested in the target image under a set number of test sequencing, the following steps are further included: Determine whether the first light intensity data is within a first preset light intensity range. If so, adjust the exposure time and / or exposure gain of the light emitting device until the first light intensity data is within the first preset light intensity range; and / or, After the step of obtaining the first light intensity data corresponding to the target pixel points of the nucleic acid sequence cluster to be tested in the target image under a set number of test sequencing, the following steps are further included: Determine whether the light intensity signal value of the first light intensity data is less than a second preset light intensity range. If so, use a preset signal amplification circuit to amplify the first light intensity data until it reaches the second preset light intensity range.

13. A base analysis method, characterized in that, The base analysis method includes: Analyze and process the obtained target base recognition result based on the base recognition method described in at least one of claims 1-7 or at least one of claims 8-12 to obtain a base analysis result.

14. A base recognition system, characterized in that, The base recognition system includes: A first light intensity data acquisition module, configured to acquire first light intensity data corresponding to each target pixel point in a target image under a set number of test sequencing; wherein, each of the target pixel points corresponds to a nucleic acid sequence cluster to be tested; A light intensity distribution interval determination module, configured to determine light intensity distribution intervals corresponding to different types of bases respectively based on the first light intensity data; A target light intensity data acquisition module, configured to acquire target light intensity data corresponding to each of the target pixel points in the target image under any actual sequencing; A base recognition module, configured to obtain the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution interval sequence.

15. A base recognition system, characterized in that, The base recognition system includes: A first light intensity data acquisition module, configured to acquire first light intensity data corresponding to the target pixel points of the nucleic acid sequence clusters to be tested in a target image under a set number of test sequencing; A light intensity distribution interval determination module, configured to determine light intensity distribution intervals corresponding to different types of bases respectively based on the first light intensity data; A target light intensity data acquisition module, configured to acquire target light intensity data corresponding to the target pixel points of the target nucleic acid sequence clusters in the target image under any actual sequencing; A base recognition module, configured to obtain the target base recognition result under the actual sequencing according to the target light intensity data and the light intensity distribution interval.

16. A base analysis system, characterized in that, The base analysis system includes: A base analysis module, configured to analyze and process the obtained target base recognition result based on the base recognition system described in claim 14 or 15 to obtain a base analysis result.

17. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes a computer program, it implements the base recognition method described in at least one of claims 1-7, or implements the base recognition method described in at least one of claims 8-12, or the base analysis method described in claim 13.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the base recognition method described in at least one of claims 1-7, or implements the base recognition method described in at least one of claims 8-12, or the base analysis method described in claim 13.

Citation Information

Patent Citations

  • Base cluster detection method and device, gene sequencer and storage medium

    CN116596933A

  • Base recognition method and training set construction method thereof, gene sequencer and medium

    CN117274739A

  • Systems and methods for identifying microparticles

    US20110096975A1

  • High-Throughput Sequencing with Semiconductor-Based Detection

    US20190212294A1

  • Equalization-Based Image Processing and Spatial Crosstalk Attenuator

    US20210350163A1