Gene sequencing method and system, electronic equipment and storage medium

CN120418445APending Publication Date: 2025-08-01QINGDAO HUADA ZHIZAO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102833.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing gene sequencing technology suffers from the problems of low base identification accuracy and high requirements on computing resources. Especially when there is a large difference in signal strength, the base sequence cannot be effectively identified and is affected by the number of reads.

Method used

By collecting signal groups to be sequenced from multiple nucleotide sequence clusters, the signal intensity value of each signal to be sequenced under different preset sequencing conditions is obtained, a two-dimensional space is constructed, and the Euclidean distance and arm angle values ​​are calculated, using the preset Analyze the processing rules to identify signal spatial distribution points, automatically identify nucleotide sequence clusters, simplify the sequencing operation process, and reduce computing resource requirements.

Benefits of technology

It improves the accuracy and efficiency of gene sequencing results, significantly reduces the error rate, increases the output and mapping rate, simplifies the sequencing operation process, and reduces the requirements for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120418445A_ABST
    Figure CN120418445A_ABST
Patent Text Reader

Abstract

The invention provides a gene sequencing method and system, electronic equipment and a storage medium. The gene sequencing method comprises the following steps: collecting a plurality of to-be-sequenced signal groups of a plurality of nucleotide sequence clusters; the plurality of to-be-sequenced signal groups are in one-to-one correspondence with the plurality of nucleotide sequence clusters, each to-be-sequenced signal group comprises a plurality of to-be-sequenced signals, and each to-be-sequenced signal corresponds to one sequencing cycle; under each sequencing cycle, acquiring a signal intensity value of the corresponding to-be-sequenced signal; and according to the signal intensity value of each to-be-sequenced signal, identifying to obtain a gene sequencing result of the nucleotide sequence cluster corresponding to the to-be-sequenced signal group. According to the method, the signal to be sequenced only needs to be independently analyzed to identify and obtain the corresponding base sequence, the sequencing operation process is simple and efficient, the requirement for computing resources is greatly reduced, the efficiency and precision of obtaining the gene sequencing result are effectively improved, the error rate is greatly reduced, and the yield and the mapping rate are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Gene sequencing method, system, electronic device and storage medium Technical Field

[0001] The present disclosure relates to the field of gene sequencing technology, and in particular to a gene sequencing method, system, electronic device, and storage medium. Background Art

[0002] Existing gene sequencing solutions generally identify each sequencing cycle one by one. After reading all the signal data in a sequencing cycle on the sequencing chip, they perform multiple processes such as pre-classification, normalization, clustering, etc. After all sequencing cycles are processed, the bases at the corresponding positions are concatenated to obtain the final base sequence. As a result, there are problems such as low base recognition accuracy and high requirements for computing resources in gene sequencing.

[0003] Summary of the Invention

[0004] The technical problem to be solved by the present disclosure is to overcome the defects of gene sequencing in the prior art, such as low base recognition accuracy and high requirements for computing resources, and the purpose is to provide a gene sequencing method, system, electronic device and storage medium.

[0005] The present disclosure solves the above technical problems through the following technical solutions:

[0006] In a first aspect, the present disclosure provides a gene sequencing method, comprising:

[0007] collecting multiple signal groups to be sequenced from multiple nucleotide sequence clusters;

[0008] The plurality of signal groups to be sequenced correspond one-to-one to the plurality of nucleotide sequence clusters, each signal group to be sequenced includes a plurality of signals to be sequenced, and each signal to be sequenced corresponds to one sequencing cycle;

[0009] In each sequencing cycle, respectively obtaining a signal intensity value corresponding to the signal to be sequenced;

[0010] According to the signal intensity value of each signal to be sequenced, a gene sequencing result of the nucleotide sequence cluster corresponding to the signal group to be sequenced is identified and obtained.

[0011] Preferably, the step of respectively obtaining the signal intensity values ​​corresponding to the signals to be sequenced comprises:

[0012] The signal strength values ​​of different preset sequencing channels corresponding to the signal to be sequenced are respectively obtained; or, the signal strength values ​​of different preset sequencing time periods corresponding to the signal to be sequenced are respectively obtained.

[0013] Preferably, the step of identifying and obtaining the gene sequencing result of the nucleotide sequence cluster corresponding to the group of signals to be sequenced according to the signal intensity value of each signal to be sequenced comprises:

[0014] Determining a corresponding number of signal spatial distribution points according to the signal intensity value of each signal to be sequenced;

[0015] Analyzing the signal spatial distribution points using preset processing rules to obtain target analysis results;

[0016] The gene sequencing result of the nucleotide sequence cluster corresponding to the signal group to be sequenced is obtained based on the target analysis result.

[0017] Preferably, the step of determining a corresponding number of signal spatial distribution points according to the signal intensity value of each signal to be sequenced includes:

[0018] Construct a preset two-dimensional space;

[0019] The spatial position information corresponding to the signal strength value in the preset two-dimensional space is determined to obtain a plurality of signal spatial distribution points distributed in the preset two-dimensional space.

[0020] Preferably, the step of analyzing the signal spatial distribution points using preset processing rules to obtain target analysis results includes:

[0021] Calculating the Euclidean distance between each of the signal space distribution points and a preset reference point in the preset two-dimensional space;

[0022] Performing statistical analysis using a first processing rule according to different Euclidean distances to obtain a first analysis result;

[0023] performing statistical analysis based on the first analysis result using a second processing rule to obtain a second analysis result;

[0024] The first analysis result corresponds to the identification result of a part of the groups, the second analysis result corresponds to the identification result of the remaining parts of the groups, and the target analysis result includes the first analysis result and the second analysis result.

[0025] Preferably, the step of performing statistical analysis using a first processing rule according to different Euclidean distances to obtain a first analysis result includes:

[0026] generating corresponding first statistical content according to different Euclidean distances;

[0027] Wherein, the first statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text;

[0028] The first statistical content is analyzed using the first processing rule to obtain the first analysis result.

[0029] Preferably, the step of using the second processing rule to perform statistical analysis based on the first analysis result to obtain the second analysis result includes:

[0030] determining an identified group based on the first analysis result;

[0031] Calculate the arm angle value between the center of the identified group and the center of the other groups to be identified;

[0032] generating corresponding second statistical content according to different arm angle values;

[0033] Wherein, the second statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text;

[0034] The second statistical content is analyzed using the second processing rule to obtain the second analysis result.

[0035] Preferably, when the first statistical content includes a first statistical histogram, the step of analyzing the first statistical content using the first processing rule to obtain the first analysis result includes:

[0036] Selecting a first longest zero sequence that meets a first screening condition in the first statistical histogram;

[0037] The first longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero;

[0038] Obtaining first starting position information of the first longest zero sequence;

[0039] Determining radius information of a first preset group according to the first starting position information;

[0040] The signal space distribution points whose Euclidean distance is smaller than the radius information are identified as the first preset group, and the remaining signal space distribution points are used as other groups to be identified.

[0041] Preferably, when the second statistical content includes a second statistical histogram, the step of analyzing the second statistical content using the second processing rule to obtain the second analysis result includes:

[0042] Selecting the second longest zero sequence and the second longest zero sequence that meet the second screening condition in the second statistical histogram;

[0043] The second longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero, and the second longest zero sequence is a sequence with the second longest horizontal length and all vertical values ​​being zero;

[0044] Based on the second longest zero sequence, the second longest zero sequence and the arm angle value, the other groups to be identified are distinguished and identified respectively.

[0045] Preferably, the first preset group includes a G group, and the other groups to be identified include an A group, a T group, and a C group.

[0046] Preferably, before the step of selecting the first longest zero sequence satisfying the first screening condition in the first statistical histogram, the step includes:

[0047] Obtaining a first number of histograms whose vertical values ​​are zero and are arranged continuously in the first statistical histogram;

[0048] If the first number is less than a set threshold, the current signal to be sequenced is determined to be a bad pixel, the current signal to be sequenced is filtered out, and the next sequencing cycle of the signal to be sequenced is continued.

[0049] Preferably, the step of selecting the first longest zero sequence that meets the first screening condition in the first statistical histogram includes:

[0050] If the first number is greater than or equal to the set threshold, the current signal to be sequenced is determined to be a good point, and the first longest zero sequence is obtained according to the initial position and the end position corresponding to the histogram of the first number.

[0051] Preferably, the gene sequencing method further comprises:

[0052] Using preset identification information to identify the signal to be sequenced that belongs to a good point;

[0053] After all the signals to be sequenced in the plurality of signal groups to be sequenced are processed, the signals to be sequenced marked with the preset identification information are stored in a preset storage space of the gene sequencing chip.

[0054] Preferably, the gene sequencing method is implemented using multi-core parallel processing hardware.

[0055] In a second aspect of the present disclosure, a gene sequencing system is provided, comprising:

[0056] A signal group collection module to be sequenced is used to collect multiple signal groups to be sequenced of multiple nucleotide sequence clusters;

[0057] The plurality of signal groups to be sequenced correspond one-to-one to the plurality of nucleotide sequence clusters, each signal group to be sequenced includes a plurality of signals to be sequenced, and each signal to be sequenced corresponds to one sequencing cycle;

[0058] A signal strength value acquisition module, configured to obtain the signal strength value corresponding to the signal to be sequenced in each sequencing cycle;

[0059] The gene sequencing result acquisition module is used to identify and obtain the gene sequencing result of the nucleotide sequence cluster corresponding to the group of signals to be sequenced according to the signal intensity value of each signal to be sequenced.

[0060] Preferably, the signal strength value acquisition module is used to respectively acquire the signal strength values ​​of different preset sequencing channels corresponding to the signal to be sequenced; or, respectively acquire the signal strength values ​​of different preset sequencing time periods corresponding to the signal to be sequenced.

[0061] Preferably, the gene sequencing result acquisition module includes:

[0062] a spatial distribution point determination unit, configured to determine a corresponding number of signal spatial distribution points according to the signal intensity value of each signal to be sequenced;

[0063] a target analysis result acquisition unit, configured to analyze the signal space distribution points using a preset processing rule to obtain a target analysis result;

[0064] A gene sequencing result acquisition unit is used to acquire the gene sequencing result of the nucleotide sequence cluster corresponding to the signal group to be sequenced based on the target analysis result.

[0065] Preferably, the spatial distribution point determination unit includes:

[0066] A two-dimensional space construction subunit, used to construct a preset two-dimensional space;

[0067] The distribution point determination subunit is configured to determine spatial position information corresponding to the signal strength value in the preset two-dimensional space, so as to obtain a plurality of signal spatial distribution points distributed in the preset two-dimensional space.

[0068] Preferably, the target analysis result acquisition unit includes:

[0069] a distance calculation subunit, configured to calculate a Euclidean distance between each of the signal space distribution points and a preset reference point in the preset two-dimensional space;

[0070] A first result acquisition subunit is configured to perform statistical analysis using a first processing rule according to different Euclidean distances to obtain a first analysis result;

[0071] A second result acquisition subunit is configured to perform statistical analysis based on the first analysis result using a second processing rule to obtain a second analysis result;

[0072] a target result acquisition subunit, configured to use the first analysis result and the second analysis result as the target analysis result;

[0073] The first analysis result corresponds to the identification result of a part of the groups, and the second analysis result corresponds to the identification result of the remaining parts of the groups.

[0074] Preferably, the first result obtaining subunit is further used for:

[0075] generating corresponding first statistical content according to different Euclidean distances;

[0076] Wherein, the first statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text;

[0077] The first statistical content is analyzed using the first processing rule to obtain the first analysis result.

[0078] Preferably, the second result obtaining subunit is further used for:

[0079] determining an identified group based on the first analysis result;

[0080] Calculate the arm angle value between the center of the identified group and the center of the other groups to be identified;

[0081] generating corresponding second statistical content according to different arm angle values;

[0082] Wherein, the second statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text;

[0083] The second statistical content is analyzed using the second processing rule to obtain the second analysis result.

[0084] Preferably, the first result obtaining subunit is further used for:

[0085] When the first statistical content includes a first statistical histogram, selecting a first longest zero sequence in the first statistical histogram that meets a first screening condition;

[0086] The first longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero;

[0087] Obtaining first starting position information of the first longest zero sequence;

[0088] Determining radius information of a first preset group according to the first starting position information;

[0089] The signal space distribution points whose Euclidean distance is smaller than the radius information are identified as the first preset group, and the remaining signal space distribution points are used as other groups to be identified.

[0090] Preferably, the second result obtaining subunit is further used for:

[0091] When the second statistical content includes a second statistical histogram, selecting the second longest zero sequence and the second longest zero sequence in the second statistical histogram that meet the second screening condition;

[0092] The second longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero, and the second longest zero sequence is a sequence with the second longest horizontal length and all vertical values ​​being zero;

[0093] Based on the second longest zero sequence, the second longest zero sequence and the arm angle value, the other groups to be identified are distinguished and identified respectively.

[0094] Preferably, the first preset group includes a G group, and the other groups to be identified include an A group, a T group, and a C group.

[0095] Preferably, the gene sequencing system further comprises:

[0096] A first quantity acquisition module, configured to acquire a first quantity of histograms whose vertical values ​​are zero and are arranged continuously in the first statistical histogram;

[0097] The judgment module is configured to determine that the current signal to be sequenced is a bad pixel if the first number is less than a set threshold, filter out the current signal to be sequenced, and continue to execute a next sequencing cycle of the signal to be sequenced.

[0098] Preferably, the judgment module is also used to determine that the current signal to be sequenced is a good point if the first number is greater than or equal to the set threshold, and call the first result acquisition subunit to obtain the first longest zero sequence according to the initial position and end position corresponding to the histogram of the first number.

[0099] Preferably, the gene sequencing system further comprises:

[0100] An information identification module, configured to identify the signal to be sequenced belonging to a good point using preset identification information;

[0101] The signal storage module is configured to store the signals to be sequenced marked with the preset identification information in a preset storage space of the gene sequencing chip after processing all the signals to be sequenced in the plurality of signal groups to be sequenced.

[0102] Preferably, the gene sequencing system further includes a processor, and the processor integrates multi-core parallel processing hardware.

[0103] In a third aspect of the present disclosure, a gene sequencing chip is provided, wherein the gene sequencing chip includes the above-mentioned gene sequencing system.

[0104] In a fourth aspect of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned gene sequencing method when executing the computer program.

[0105] In a fifth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned gene sequencing method is implemented.

[0106] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present disclosure.

[0107] The positive progress of this disclosure is:

[0108] In the present disclosure, a signal group to be sequenced is collected for each nucleotide sequence cluster, and the signal intensity values ​​of the signals to be sequenced in each signal group to be sequenced under corresponding sequencing cycles and different preset sequencing conditions are obtained respectively, so as to automatically identify the gene sequencing results of the nucleotide sequence cluster corresponding to the signal group to be sequenced. There is no need to consider the impact of signal strengths between different signal groups to be sequenced on the sequencing results, and it is not affected by the number of reads. There is no need to perform separate quality value filtering. Only independent analysis of the signals to be sequenced is required to identify the corresponding base sequence. The sequencing operation process is simple and efficient, greatly reducing the requirements for computing resources, effectively improving the efficiency and accuracy of obtaining gene sequencing results, and significantly reducing the error rate, thereby achieving an increase in yield and mapping rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0109] FIG1 is a flow chart of the gene sequencing method of Example 1 of the present disclosure.

[0110] FIG2 is a first flow chart of the gene sequencing method of Example 2 of the present disclosure.

[0111] FIG3 is a second flow chart of the gene sequencing method of Example 2 of the present disclosure.

[0112] FIG4 is a third flow chart of the gene sequencing method of Example 2 of the present disclosure.

[0113] FIG5 is a schematic diagram of a statistical histogram formed based on Euclidean distance values ​​according to Example 2 of the present disclosure.

[0114] FIG6 is a schematic diagram of a statistical histogram formed based on arm angle values ​​according to Embodiment 2 of the present disclosure.

[0115] FIG7 is a schematic diagram of the gene sequencing results of Example 2 of the present disclosure.

[0116] FIG8 is a flow chart of an example of the gene sequencing method of Example 2 of the present disclosure.

[0117] FIG9 is a module diagram of the gene sequencing system of Example 3 of the present disclosure.

[0118] FIG10 is a module diagram of the gene sequencing system of Example 4 of the present disclosure.

[0119] FIG11 is a schematic structural diagram of an electronic device according to Embodiment 5 of the present disclosure. DETAILED DESCRIPTION

[0120] The present disclosure is further illustrated below by way of examples, but the present disclosure is not limited to the scope of the examples.

[0121] Gene sequencing refers to the analysis of the base sequence of a specific DNA fragment, namely the arrangement of adenine (A), thymine (T), cytosine (C), and guanine (G). The core principle of second-generation sequencing is sequencing by synthesis, which basically includes library construction, the generation of monoclonal DNA clusters, and sequencing reactions, allowing the simultaneous analysis of DNA samples on an array. The reaction carrier for loading DNA is a chip, where the DNA reacts on the chip. The base sequence at that position is determined by measuring the fluorescent groups attached to the reaction. The fluorescent signals of the DNA on the chip are collected by an optical system and a scientific camera and converted into digital signals.

[0122] The core principle of next-generation sequencing is sequencing-by-synthesis. For each biochemical reaction that synthesizes a base, the sequencer collects a fluorescent signal from the chip. This process is called a cycle. As a result, each cycle measures a single base in the DNA sequence on the chip array. After several cycles, the DNA sequence on the chip array is obtained.

[0123] In existing gene sequencing schemes, the identification is generally carried out one by one in sequencing cycles. After reading all the signal data in a sequencing cycle on the sequencing chip, pre-classification, normalization, clustering and other processes are carried out in sequence. After all the sequencing cycles are processed, the bases at the corresponding positions are connected in series to obtain the final base sequence. This gene sequencing scheme has the following defects: (1) When there is a large difference in signal strength, interference will occur between the bases, and there will be a significant deviation in base identification, resulting in low base identification accuracy; (2) It is impossible to identify whether the signal is normal or not, and thus it is impossible to distinguish between normal and abnormal signals, which leads to a reduction in the final base identification accuracy; (3) When the number of reads (sequencing fragments) is small, base identification cannot be performed, that is, it is affected by the number of reads; (4) The overall execution process is complex and requires high computing resources.

[0124] Example 1

[0125] As shown in FIG1 , the gene sequencing method of this embodiment includes:

[0126] S101, collecting multiple signal groups to be sequenced from multiple nucleotide sequence clusters;

[0127] The plurality of signal groups to be sequenced correspond one-to-one to the plurality of nucleotide sequence clusters, that is, each nucleotide sequence cluster corresponds to one signal group to be sequenced, each signal group to be sequenced includes a plurality of signals to be sequenced, and each signal to be sequenced corresponds to one sequencing cycle;

[0128] Specifically, a cluster is a group of similar or identical molecules or nucleotide sequences or DNA strands; for example, a cluster can be any other group of amplified oligonucleotides or polynucleotides or polypeptides having the same or similar sequences. In other embodiments, a cluster can be any element or group of elements that occupies a physical area on a sample surface. In embodiments, clusters are fixed to reaction sites and / or reaction chambers during base sequencing cycles.

[0129] S102, obtaining the signal intensity value of the corresponding signal to be sequenced in each sequencing cycle;

[0130] In each sequencing cycle, the signal strength of the corresponding signal to be sequenced under different preset sequencing conditions is analyzed separately, that is, each signal to be sequenced is analyzed and processed independently, instead of the existing method of identifying different signals to be sequenced one by one according to sequencing cycle. This avoids the impact of mixed processing of different signal strengths, normal and abnormal signals, etc. on the gene sequencing results from the source, greatly improving the accuracy of the gene sequencing results and the processing efficiency of the gene sequencing process.

[0131] S103 , identifying and obtaining gene sequencing results of nucleotide sequence clusters corresponding to the corresponding signal group to be sequenced based on the signal intensity value of each signal to be sequenced.

[0132] In this embodiment, a signal group to be sequenced is collected for each nucleotide sequence cluster, and the signal intensity values ​​of the signals to be sequenced in each signal group to be sequenced under corresponding sequencing cycles and different preset sequencing conditions are obtained to automatically identify the gene sequencing results of the nucleotide sequence cluster corresponding to the signal group to be sequenced. There is no need to consider the impact of signal strengths between different signal groups to be sequenced on the sequencing results, and it is not affected by the number of reads. There is no need to perform separate quality value filtering. Only independent analysis of the signals to be sequenced is required to identify the corresponding base sequence. The sequencing operation process is simple and efficient, greatly reducing the requirements for computing resources, effectively improving the efficiency and accuracy of obtaining gene sequencing results, and significantly reducing the error rate, thereby achieving an increase in yield and mapping rate.

[0133] Example 2

[0134] The gene sequencing method of this embodiment is a further improvement of Example 1. Specifically:

[0135] In one feasible solution, as shown in FIG2 , step S102 includes:

[0136] S1021. In each sequencing cycle, respectively obtain signal intensity values ​​of different preset sequencing channels corresponding to the signal to be sequenced; or respectively obtain signal intensity values ​​of different preset sequencing time periods corresponding to the signal to be sequenced.

[0137] Of course, in addition to the preset sequencing conditions for different preset sequencing channels and different preset sequencing time periods, other sequencing conditions that can independently analyze the sequencing signals can also be used, which will not be described in detail here.

[0138] In one feasible solution, step S103 includes:

[0139] S1031. Determine a corresponding number of signal spatial distribution points according to the signal intensity value of each signal to be sequenced;

[0140] S1032. Analyze the signal spatial distribution points using a preset processing rule to obtain a target analysis result;

[0141] S1033. Obtain gene sequencing results of the nucleotide sequence cluster corresponding to the corresponding signal group to be sequenced based on the target analysis results.

[0142] In one feasible solution, as shown in FIG3 , step S1031 includes:

[0143] S10311. Construct a preset two-dimensional space;

[0144] S10312. Determine spatial position information corresponding to the signal strength value in the preset two-dimensional space to obtain a plurality of signal spatial distribution points distributed in the preset two-dimensional space.

[0145] Specifically, when different preset sequencing channels correspond to a first preset sequencing channel and a second preset sequencing channel, the signal intensity value corresponds to a first signal intensity in the first preset sequencing channel and a second signal intensity in the second preset sequencing channel.

[0146] The constructed preset two-dimensional space is a two-dimensional coordinate system, the first preset sequencing channel is used as the abscissa, the second preset sequencing channel is used as the ordinate, and a plurality of signal space distribution points in the two-dimensional coordinate system are obtained.

[0147] In one feasible solution, step S1032 includes:

[0148] S10321. Calculate the Euclidean distance between each signal space distribution point and a preset reference point in a preset two-dimensional space;

[0149] Among them, the Euclidean distance between each signal space distribution point and the coordinate origin of the two-dimensional coordinate system is calculated; the calculation method of the Euclidean distance between any two coordinate points in space is a mature technology in this field, so it will not be repeated here.

[0150] S10322. Perform statistical analysis using a first processing rule according to different Euclidean distances to obtain a first analysis result;

[0151] The first processing rule includes but is not limited to: a longest zero sequence search rule based on a Euclidean distance statistics table.

[0152] S10323. Perform statistical analysis based on the first analysis result using a second processing rule to obtain a second analysis result;

[0153] The first processing rule includes but is not limited to: a longest zero sequence search rule based on an angle statistics table.

[0154] S10324. Use the first analysis result and the second analysis result as target analysis results;

[0155] The first analysis result corresponds to the identification result of a part of the groups, and the second analysis result corresponds to the identification result of the remaining parts of the groups.

[0156] In one feasible solution, as shown in FIG4 , step S10322 includes:

[0157] S103221. Generate corresponding first statistical content according to different Euclidean distances;

[0158] The first statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text;

[0159] S103222. Analyze the first statistical content using a first processing rule to obtain a first analysis result.

[0160] In one feasible solution, step S10323 includes:

[0161] S103231. Determine the identified group based on the first analysis result;

[0162] S103232. Calculate the arm angle value between the center of the identified group and the center of the other groups to be identified;

[0163] S103233. Generate corresponding second statistical content according to different arm angle values;

[0164] The second statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text;

[0165] S103234. Analyze the second statistical content using a second processing rule to obtain a second analysis result.

[0166] In one feasible solution, the first statistical content includes a first statistical histogram, and the number of the first statistical histograms is equal to the number of sequencing cycles. Step S103222 includes:

[0167] Selecting the first longest zero sequence that meets the first screening condition in the first statistical histogram;

[0168] The first longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero;

[0169] Specifically, as shown in FIG5 , which is a first statistical histogram obtained based on statistics of all Euclidean distance values, the first longest zero sequence to be found is shown in the black frame S1 in the figure.

[0170] Obtain first starting position information of the first longest zero sequence;

[0171] Determining radius information of a first preset group according to the first starting position information;

[0172] Signal space distribution points whose Euclidean distance is smaller than the radius information are identified as a first preset group, and the remaining signal space distribution points are used as other groups to be identified.

[0173] Among them, the first preset group includes the G group, and the other groups to be identified include the A group, the T group and the C group.

[0174] Based on the first statistical histogram, the first longest zero sequence is found to identify the base G cluster, the radius and center of the base G cluster are determined, and then the recognition results of other bases ACT are obtained, thereby ensuring the accuracy of base recognition.

[0175] Specifically, before the step of selecting the first longest zero sequence that meets the first screening condition in the first statistical histogram, the method includes:

[0176] Obtaining a first number of histograms whose vertical values ​​are zero and are arranged continuously in the first statistical histogram;

[0177] If the first number is less than the set threshold, the current signal to be sequenced is determined to be a bad pixel, the current signal to be sequenced is filtered out, and the sequencing cycle of the next signal to be sequenced is continued.

[0178] In this solution, all abnormal signals to be sequenced that do not meet the conditions are filtered out, leaving only normal signals to be sequenced to determine the final gene sequencing results, further ensuring the accuracy and reliability of the sequencing results.

[0179] The step of selecting the first longest zero sequence that meets the first screening condition in the first statistical histogram includes:

[0180] If the first number is greater than or equal to the set threshold, the current signal to be sequenced is determined to be a good point, and the first longest zero sequence is obtained according to the initial position and the end position corresponding to the histogram of the first number.

[0181] In this solution, the traditional Q-value method of filtering signal bad points is abandoned, and the length of the first longest zero sequence is used for filtering, which improves the accuracy of bad point identification and simplifies the filtering process, thereby effectively improving the final gene sequencing accuracy and efficiency.

[0182] In one feasible solution, when the second statistical content includes a second statistical histogram, step S103234 includes:

[0183] Selecting the second longest zero sequence and the second longest zero sequence that meet the second screening condition in the second statistical histogram;

[0184] The second longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero, and the second longest zero sequence is a sequence with the second longest horizontal length and all vertical values ​​being zero.

[0185] Specifically, as shown in FIG6 , a second statistical histogram is obtained based on statistics of all arm angle values, and the second longest zero sequence and the second longest zero sequence to be found are shown in the black boxes S2 and S3 in the figure, respectively.

[0186] Based on the second longest zero sequence, the second longest zero sequence and the arm angle value, other groups to be identified are distinguished and identified respectively.

[0187] In this scheme, the two target arm angle values ​​between the center of the first preset group under the second longest zero sequence and the second longest zero sequence and any two signal space distribution points in other preset groups are calculated respectively; based on the two target arm angle values, the group type corresponding to each other preset group is determined.

[0188] Specifically, among the two target arm angle values, if one target arm angle value is greater than a preset angle and the other target arm angle value is less than a preset angle, it is determined that both target arm angle values ​​are valid, and based on the two target arm angle values, the group type corresponding to each other preset group is determined;

[0189] Otherwise, obtain the third long zero sequences corresponding to several second statistical histograms; calculate two new target arm angle values ​​between the center of the first preset group under the first longest zero sequence and the third long zero sequence and any two signal space distribution points in other preset groups respectively; based on the two new target arm angle values, determine the group type corresponding to each other preset group.

[0190] Wherein, after determining that the arm angle value is valid, the step of determining the group type corresponding to each other group based on the two target arm angle values ​​includes:

[0191] Selecting a signal space distribution point whose arm angle value with the center of the first predetermined group is greater than the larger angle value of the two target arm angle values ​​to form an adenine A group;

[0192] Selecting a signal space distribution point whose arm angle value with the center of the first predetermined group is smaller than the smaller angle value of the two target arm angle values ​​to form a cytosine T group;

[0193] Thymine C clusters are formed based on the remaining signal spatial distribution points.

[0194] In one feasible solution, the gene sequencing method further comprises:

[0195] Use preset identification information to identify the signals to be sequenced that belong to good points;

[0196] After all the signals to be sequenced in the plurality of signal groups to be sequenced are processed, the signals to be sequenced marked with preset identification information are stored in a preset storage space of the gene sequencing chip.

[0197] As shown in FIG7 , the gene sequencing method according to this embodiment reads the signal intensity values ​​of the corresponding signal to be sequenced in different preset sequencing channels under one sequencing cycle, and the distribution diagram of the signal intensity after normalization shows that there is a good clustering phenomenon, thereby ensuring the accuracy of the gene sequencing results.

[0198] In one feasible solution, the gene sequencing method is implemented using multi-core parallel processing hardware.

[0199] Specifically, the gene sequencing process of the present embodiment has lower requirements for computing resources compared to existing gene sequencing solutions, is more convenient to be used in actual products, and reduces costs. If multi-core parallel processing hardware such as GPU (graphics processing unit), FPGA (field programmable gate array) are used to accelerate computing, efficiency can be significantly better than existing gene sequencing solutions. The prerequisite for parallel computing is that a task can be divided into several unrelated subtasks, and for this method, Each DNB (a nanoball sequencing method based on a patterned array of in situ sequencing) can be used as an independent computing subtask. In addition, the advantage of GPU compared with CPU is that the number of processing cores of GPU is almost hundreds of times that of CPU, and each processing core can execute tasks in parallel, and all DNB data processing can be assigned to the processing cores of GPU for parallel execution, which can obtain a processing speed higher than that on CPU.

[0200] The gene sequencing method of this embodiment can be applied to sequencing platforms such as NGS (high-throughput sequencing technology) second-generation sequencing platforms.

[0201] The following describes the implementation principle of the gene sequencing method of this embodiment in detail with reference to an example (see FIG8 ):

[0202] (1) collecting multiple signal groups to be sequenced from multiple nucleotide sequence clusters;

[0203] The plurality of signal groups to be sequenced correspond one-to-one to the plurality of nucleotide sequence clusters, each signal group to be sequenced includes n signals to be sequenced (n is an integer), and each signal to be sequenced corresponds to one sequencing cycle;

[0204] (2) in each sequencing cycle, respectively obtaining a first signal intensity signal1 in the first preset sequencing channel image2 and a second signal intensity signal2 in the second preset sequencing channel image2 corresponding to the signal to be sequenced;

[0205] (3) constructing a two-dimensional coordinate system, using the first preset sequencing channel image1 as the x-coordinate and the second preset sequencing channel image2 as the y-coordinate, to obtain a corresponding number of signal space distribution points; wherein the x-coordinate of each signal space distribution point corresponds to the first signal intensity signal1, and the y-coordinate corresponds to the second signal intensity signal2;

[0206] (4) The signal strength and coordinate values ​​are normalized to ensure the feasibility of subsequent data calculation and simplify the complexity of data processing; the Euclidean distance between each signal space distribution point and the coordinate origin of the two-dimensional coordinate system is obtained based on the numerical calculation after normalization. The Euclidean distance calculation formula is as follows:

[0207]

[0208] (5) For the Euclidean distances obtained in step (4) (distance values ​​between 1% and 99%), a first statistical histogram is obtained by counting the distances;

[0209] The number of histograms num_hist in the first statistical histogram is the same as the number of sequencing cycles. For example, when there are 100 sequencing cycles, a first statistical histogram containing 100 histograms is obtained.

[0210] Here, removing distance values ​​within the range of less than 1% and greater than 99% is equivalent to filtering out possible interference values, ensuring the accuracy of the subsequent acquisition of the first longest zero sequence, thereby further improving the accuracy of the final gene sequencing results;

[0211] (6) Select the first longest zero sequence in the first statistical histogram that meets the first screening condition; wherein the first longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero; the detailed steps are as follows:

[0212] 1) Determine whether the vertical value of each histogram is 0 in turn. If the vertical value of the i-th histogram is 0, then number is increased by 1; if the vertical value of the i-th histogram is not 0, then save the signal space distribution point i at the current position into the array location, save number into the array location_number, and clear number to 0;

[0213] 2) Take the maximum value max_value from the array location_number and obtain the maximum subscript max_subscript at the same time; if max_value is less than the set value (for example, 7), set the flag corresponding to the signal space distribution point to 1, indicating that the signal to be sequenced is a bad point and needs to be discarded, thereby achieving the purpose of directly filtering out abnormal signals; if max_value is greater than or equal to 7, set the flag corresponding to the signal space distribution point to 0, indicating that the signal to be sequenced is a good point, then obtain the values ​​of the subscripts max_subscript-1 and max_subscript from the array location_number, use the value of max_subscript-1 as the initial position x1 of the first longest zero sequence, and use the value of max_subscript as the end position x2 of the first longest zero sequence;

[0214] (7) According to the initial position x1 and the end position x2 of the first longest zero sequence obtained in step (6), the radius of the base G cluster RadiusG is calculated. The corresponding calculation formula is as follows:

[0215]

[0216] Wherein, P1 refers to the value at the 1% position among all the distance values ​​obtained in step (4), and P99 refers to the value at the 99% position among all the distance values ​​obtained in step (4);

[0217] (8) Identify the base G group (i.e., the first preset group) based on the distance value obtained in step (4)

[0218] Among them, the signal space distribution points with a Euclidean distance less than RadiusG in step (4) are identified as G, and the remaining signal space distribution points are temporarily identified as other group types, namely ACT groups (i.e., other groups to be identified);

[0219] (9) Calculate the center of the base G group

[0220] According to the recognition results of step (8), the signal intensity values ​​and numbers of all G-groups identified in the first preset sequencing channel image1 and the second preset sequencing channel image2 are calculated respectively. The corresponding calculation formulas are as follows:

[0221]

[0222]

[0223] Wherein, m is the number of G bases identified in any channel, and the number of G bases identified in the first preset sequencing channel image1 and the second preset sequencing channel image2 is the same, centerGx and centerGy are the x-coordinate and y-coordinate of the center of the base G cluster, respectively; the number of base G clusters in different preset sequencing channels is directly obtained by counting the identification results of step (8), and will not be repeated here.

[0224] (10) Calculate the clustering angles of other ACT groups to be identified

[0225] Based on the center of the base G cluster obtained in step (9), the arm angle value slop between the signal space distribution point identified as the ACT group in step (8) and the center of the base G cluster is calculated. The corresponding arm angle value slop is calculated as follows:

[0226]

[0227] (11) For the arm angle slop value obtained in step (10) (the arm angle value between 1% and 99%), a second statistical histogram is obtained by counting, e.g., the number of the second statistical histogram is 90;

[0228] Here, eliminating arm angle values ​​within the range of less than 1% and greater than 99% is equivalent to filtering out possible interference values, ensuring the accuracy of the subsequent second longest zero sequence and the second longest zero sequence, thereby further improving the accuracy of the final gene sequencing results.

[0229] (12) Determine the second longest zero sequence and the second longest zero sequence in the second statistical histogram in the same manner as step (6). The implementation process is the same and will not be repeated here;

[0230] (13) Combined with step (10), the two arm angle values ​​slop1 and slop2 corresponding to the second longest zero sequence and the second longest zero sequence are calculated respectively;

[0231] (14) Obtain the ACT boundary angles slopAC and slopTC. The corresponding calculation formula is as follows:

[0232] slopAC=max{slop1,slop2}

[0233] slopTC=min{slop1,slop2}

[0234] (15) Based on the recognition result of the base G group in step (8), the remaining signal space distribution points corresponding to other groups to be recognized are re-identified, specifically:

[0235] The signal space distribution points with arm angle values ​​greater than slopAC in step (11) are identified as A clusters, the signal space distribution points with arm angle values ​​less than slopTC in step (11) are identified as T clusters, and the remaining signal space distribution points are identified as C clusters;

[0236] It should be noted that, the arm angle values ​​slop1 and slop2 need to have one greater than 45 degrees and one less than 45 degrees. Otherwise, it means that the current arm angle value is unavailable, and it is necessary to continue to search for the third longest zero sequence in the second statistical histogram, use the third longest zero sequence to replace the second longest zero sequence, and repeat the above steps (12)-(15) until the new arm angle values ​​slop1 and slop2 are calculated to meet the conditions that one is greater than 45 degrees and the other is less than 45 degrees. If the newly obtained zero sequence still cannot meet the condition that the arm angle values ​​slop1 and slop2 need to have one greater than 45 degrees and the other less than 45 degrees, it means that no effective sequencing results can be obtained, and the gene sequencing operation is stopped.

[0237] (16) Return to step (1), read the next signal to be sequenced, and repeat all the above steps until all the signals to be sequenced in the signal group to be sequenced corresponding to each nucleotide sequence cluster are processed;

[0238] (17) After all the signals to be sequenced have been processed, when writing fq, if the flag corresponding to the signal to be sequenced in step (6) is 0, that is, the signal to be sequenced is a good point representing a normal signal, then the write fq operation is performed to store it in the preset storage space of the gene sequencing chip; if the flag corresponding to the signal to be sequenced is 1, that is, the signal to be sequenced is a bad point representing an abnormal signal, then fq is not written, thereby achieving the purpose of automatically filtering abnormal signals.

[0239] In addition, to test the effectiveness of the gene sequencing solution of this embodiment, offline data from two gene sequencing chips of the NGS sequencing platform were randomly selected. The data for gene sequencing chips 1 and 2 were 817x765 rectangular areas randomly selected from the two chips, respectively. Chip 3 was the offline data for the entire chip corresponding to chip 2. The test results are shown in the following table:

[0240]

[0241] After testing, it was found that the gene sequencing solution of this embodiment had a mapping rate of over 99.6% and an error rate of under 0.15%, regardless of whether the entire chip was tested or a test area was randomly selected from the chip. It also showed consistency across different chips and demonstrated high robustness. Therefore, it can be shown that the gene sequencing solution of this embodiment is feasible and can complete gene sequencing operations with high quality.

[0242] In this embodiment, a signal group to be sequenced is collected for each nucleotide sequence cluster, and the signal intensity values ​​of the signals to be sequenced in each signal group to be sequenced under corresponding sequencing cycles and different preset sequencing conditions are obtained to automatically identify the gene sequencing results of the nucleotide sequence cluster corresponding to the signal group to be sequenced. There is no need to consider the impact of signal strengths between different signal groups to be sequenced on the sequencing results, and it is not affected by the number of reads. There is no need to perform separate quality value filtering. Only independent analysis of the signals to be sequenced is required to identify the corresponding base sequence. The sequencing operation process is simple and efficient, greatly reducing the requirements for computing resources, effectively improving the efficiency and accuracy of obtaining gene sequencing results, and significantly reducing the error rate, thereby achieving an increase in yield and mapping rate.

[0243] Example 3

[0244] As shown in FIG9 , the gene sequencing system of this embodiment includes:

[0245] The signal group collection module 1 to be sequenced is used to collect multiple signal groups to be sequenced of multiple nucleotide sequence clusters;

[0246] The plurality of signal groups to be sequenced correspond one-to-one to the plurality of nucleotide sequence clusters, that is, each nucleotide sequence cluster corresponds to one signal group to be sequenced, each signal group to be sequenced includes a plurality of signals to be sequenced, and each signal to be sequenced corresponds to one sequencing cycle;

[0247] Specifically, a cluster is a group of similar or identical molecules or nucleotide sequences or DNA strands; for example, a cluster can be any other group of amplified oligonucleotides or polynucleotides or polypeptides having the same or similar sequences. In other embodiments, a cluster can be any element or group of elements that occupies a physical area on a sample surface. In embodiments, clusters are fixed to reaction sites and / or reaction chambers during base sequencing cycles.

[0248] Signal intensity value acquisition module 2, used to obtain the signal intensity value of the corresponding signal to be sequenced in each sequencing cycle;

[0249] In each sequencing cycle, the signal strength of the corresponding signal to be sequenced under different preset sequencing conditions is analyzed separately, that is, each signal to be sequenced is analyzed and processed independently, instead of the existing method of identifying different signals to be sequenced one by one according to sequencing cycle. This avoids the impact of mixed processing of different signal strengths, normal and abnormal signals, etc. on the gene sequencing results from the source, greatly improving the accuracy of the gene sequencing results and the processing efficiency of the gene sequencing process.

[0250] The gene sequencing result acquisition module 3 is used to identify and obtain the gene sequencing result of the nucleotide sequence cluster corresponding to the corresponding signal group to be sequenced based on the signal intensity value of each signal to be sequenced.

[0251] It should be noted that the implementation principle of the gene sequencing system of this embodiment is the same as the implementation principle of the gene sequencing method of Example 1, so it will not be repeated here.

[0252] In this embodiment, a signal group to be sequenced is collected for each nucleotide sequence cluster, and the signal intensity values ​​of the signals to be sequenced in each signal group to be sequenced under corresponding sequencing cycles and different preset sequencing conditions are obtained to automatically identify the gene sequencing results of the nucleotide sequence cluster corresponding to the signal group to be sequenced. There is no need to consider the impact of signal strengths between different signal groups to be sequenced on the sequencing results, and it is not affected by the number of reads. There is no need to perform separate quality value filtering. Only independent analysis of the signals to be sequenced is required to identify the corresponding base sequence. The sequencing operation process is simple and efficient, greatly reducing the requirements for computing resources, effectively improving the efficiency and accuracy of obtaining gene sequencing results, and significantly reducing the error rate, thereby achieving an increase in yield and mapping rate.

[0253] Example 4

[0254] As shown in FIG10 , the gene sequencing system of this embodiment is a further improvement of Example 3. Specifically:

[0255] In one feasible solution, the signal strength value acquisition module 2 is used to respectively acquire the signal strength values ​​of different preset sequencing channels corresponding to the signal to be sequenced in each sequencing cycle; or, respectively acquire the signal strength values ​​of different preset sequencing time periods corresponding to the signal to be sequenced.

[0256] Of course, in addition to the preset sequencing conditions for different preset sequencing channels and different preset sequencing time periods, other sequencing conditions that can independently analyze the sequencing signals can also be used, which will not be described in detail here.

[0257] In one feasible solution, the gene sequencing result acquisition module 3 includes:

[0258] The spatial distribution point determination unit 4 is used to determine a number of corresponding signal spatial distribution points according to the signal intensity value of each signal to be sequenced;

[0259] The target analysis result acquisition unit 5 is used to analyze the signal space distribution points using a preset processing rule to obtain the target analysis result;

[0260] The gene sequencing result acquisition unit 6 is used to acquire the gene sequencing result of the nucleotide sequence cluster corresponding to the corresponding signal group to be sequenced based on the target analysis result.

[0261] In one feasible solution, the spatial distribution point determination unit 4 includes:

[0262] A two-dimensional space construction subunit 7, used to construct a preset two-dimensional space;

[0263] The distribution point determination subunit 8 is configured to determine spatial position information corresponding to the signal strength value in the preset two-dimensional space, so as to obtain a plurality of signal spatial distribution points distributed in the preset two-dimensional space.

[0264] Specifically, when different preset sequencing channels correspond to a first preset sequencing channel and a second preset sequencing channel, the signal intensity value corresponds to a first signal intensity in the first preset sequencing channel and a second signal intensity in the second preset sequencing channel.

[0265] The constructed preset two-dimensional space is a two-dimensional coordinate system, the first preset sequencing channel is used as the abscissa, the second preset sequencing channel is used as the ordinate, and a plurality of signal space distribution points in the two-dimensional coordinate system are obtained.

[0266] In one feasible solution, the target analysis result acquisition unit 5 includes:

[0267] The distance calculation subunit 9 is used to calculate the Euclidean distance between each signal space distribution point and a preset reference point in a preset two-dimensional space;

[0268] Among them, the Euclidean distance between each signal space distribution point and the coordinate origin of the two-dimensional coordinate system is calculated; the calculation method of the Euclidean distance between any two coordinate points in space is a mature technology in this field, so it will not be repeated here.

[0269] A first result acquisition subunit 10 is configured to perform statistical analysis using a first processing rule according to different Euclidean distances to obtain a first analysis result;

[0270] The first processing rule includes but is not limited to: a longest zero sequence search rule based on a Euclidean distance statistics table.

[0271] A second result acquisition subunit 11 is configured to perform statistical analysis based on the first analysis result using a second processing rule to obtain a second analysis result;

[0272] The first processing rule includes but is not limited to: a longest zero sequence search rule based on an angle statistics table.

[0273] A target result acquisition subunit 12 is configured to use the first analysis result and the second analysis result as target analysis results;

[0274] The first analysis result corresponds to the identification result of a part of the groups, and the second analysis result corresponds to the identification result of the remaining parts of the groups.

[0275] In one feasible solution, the first result obtaining subunit 10 is further configured to:

[0276] Generate corresponding first statistical content according to different Euclidean distances;

[0277] The first statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text;

[0278] The first statistical content is analyzed using a first processing rule to obtain a first analysis result.

[0279] In one feasible solution, the second result obtaining subunit 11 is further configured to:

[0280] determining an identified group based on the first analysis result;

[0281] Calculate the arm angle value between the center of the identified group and the center of the other groups to be identified;

[0282] generating corresponding second statistical content according to different arm angle values;

[0283] The second statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text;

[0284] The second statistical content is analyzed using a second processing rule to obtain a second analysis result.

[0285] In one feasible solution, the first statistical content includes first statistical histograms, and the number of the first statistical histograms is equal to the number of sequencing cycles.

[0286] The first result obtaining subunit 10 is further configured to:

[0287] Selecting the first longest zero sequence that meets the first screening condition in the first statistical histogram;

[0288] The first longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero;

[0289] Specifically, as shown in FIG5 , which is a first statistical histogram obtained based on statistics of all Euclidean distance values, the first longest zero sequence to be found is shown in the black frame S1 in the figure.

[0290] Obtain first starting position information of the first longest zero sequence;

[0291] Determining radius information of a first preset group according to the first starting position information;

[0292] Signal space distribution points whose Euclidean distance is smaller than the radius information are identified as a first preset group, and the remaining signal space distribution points are used as other groups to be identified.

[0293] Among them, the first preset group includes the G group, and the other groups to be identified include the A group, the T group and the C group.

[0294] Based on the first statistical histogram, the first longest zero sequence is found to identify the base G cluster, the radius and center of the base G cluster are determined, and then the recognition results of other bases ACT are obtained, thereby ensuring the accuracy of base recognition.

[0295] In one feasible solution, the gene sequencing system further includes:

[0296] A first number acquisition module 13 is configured to acquire a first number of histograms whose vertical values ​​are zero and are arranged continuously in the first statistical histogram;

[0297] The judgment module 14 is configured to determine that the current signal to be sequenced is a bad pixel if the first number is less than a set threshold, filter out the current signal to be sequenced, and continue to execute a sequencing cycle for the next signal to be sequenced.

[0298] In this solution, all abnormal signals to be sequenced that do not meet the conditions are filtered out, leaving only normal signals to be sequenced to determine the final gene sequencing results, further ensuring the accuracy and reliability of the sequencing results.

[0299] In one feasible solution, the judgment module 14 is further used to determine that the current signal to be sequenced is a good point if the first number is greater than or equal to a set threshold, and to call the first result acquisition subunit to obtain the first longest zero sequence according to the initial position and end position corresponding to the histogram of the first number.

[0300] In this solution, the traditional Q-value method of filtering signal bad points is abandoned, and the length of the first longest zero sequence is used for filtering, which improves the accuracy of bad point identification and simplifies the filtering process, thereby effectively improving the final gene sequencing accuracy and efficiency.

[0301] In one feasible solution, the second result obtaining subunit 11 is further configured to:

[0302] When the second statistical content includes a second statistical histogram, selecting the second longest zero sequence and the second longest zero sequence in the second statistical histogram that meet the second screening condition;

[0303] The second longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero, and the second longest zero sequence is a sequence with the second longest horizontal length and all vertical values ​​being zero.

[0304] Specifically, as shown in FIG6 , a second statistical histogram is obtained based on statistics of all arm angle values, and the second longest zero sequence and the second longest zero sequence to be found are shown in the black boxes S2 and S3 in the figure, respectively.

[0305] Based on the second longest zero sequence, the second longest zero sequence and the arm angle value, other groups to be identified are distinguished and identified respectively.

[0306] In this scheme, the two target arm angle values ​​between the center of the first preset group under the second longest zero sequence and the second longest zero sequence and any two signal space distribution points in other preset groups are calculated respectively; based on the two target arm angle values, the group type corresponding to each other preset group is determined.

[0307] Specifically, among the two target arm angle values, if one target arm angle value is greater than a preset angle and the other target arm angle value is less than a preset angle, it is determined that both target arm angle values ​​are valid, and based on the two target arm angle values, the group type corresponding to each other preset group is determined;

[0308] Otherwise, obtain the third long zero sequences corresponding to several second statistical histograms; calculate two new target arm angle values ​​between the center of the first preset group under the first longest zero sequence and the third long zero sequence and any two signal space distribution points in other preset groups respectively; based on the two new target arm angle values, determine the group type corresponding to each other preset group.

[0309] Wherein, after determining that the arm angle value is valid, the step of determining the group type corresponding to each other group based on the two target arm angle values ​​includes:

[0310] Selecting a signal space distribution point whose arm angle value with the center of the first predetermined group is greater than the larger angle value of the two target arm angle values ​​to form an adenine A group;

[0311] Selecting a signal space distribution point whose arm angle value with the center of the first predetermined group is smaller than the smaller angle value of the two target arm angle values ​​to form a cytosine T group;

[0312] Thymine C clusters are formed based on the remaining signal spatial distribution points.

[0313] In one feasible solution, the gene sequencing system further includes:

[0314] An information identification module 15 is used to identify the signals to be sequenced that belong to good points using preset identification information;

[0315] The signal storage module 16 is configured to store the signals to be sequenced marked with preset identification information in a preset storage space of the gene sequencing chip after processing all the signals to be sequenced in the plurality of signal groups to be sequenced.

[0316] As shown in FIG7 , the gene sequencing method according to this embodiment reads the signal intensity values ​​of the corresponding signal to be sequenced in different preset sequencing channels under one sequencing cycle, and the distribution diagram of the signal intensity after normalization shows that there is a good clustering phenomenon, thereby ensuring the accuracy of the gene sequencing results.

[0317] In one feasible solution, the gene sequencing system further includes a processor 17 , which integrates multi-core parallel processing hardware.

[0318] Specifically, the gene sequencing process of this embodiment has lower requirements for computing resources than existing gene sequencing solutions, is more convenient to use in actual products, and reduces costs. If multi-core parallel processing hardware such as GPU (graphics processing unit), FPGA (field programmable gate array) and the like are used to accelerate computing, the efficiency will be significantly better than existing gene sequencing solutions. The prerequisite for parallel computing is that a task can be divided into several unrelated subtasks, and for this method, each DNB can be used as an independent computing subtask. In addition, the advantage of GPU over CPU is that the number of processing cores of GPU is almost hundreds or thousands of times that of CPU, and each processing core can execute tasks in parallel. All DNB data processing can be assigned to the processing cores of GPU for parallel execution, which can achieve a higher processing speed than on CPU.

[0319] The gene sequencing method of this embodiment can be applied to sequencing platforms such as the NGS second-generation sequencing platform.

[0320] It should be noted that the implementation principle of the gene sequencing system of this embodiment is the same as the implementation principle of the gene sequencing method of Example 2, so it will not be repeated here.

[0321] In this embodiment, a signal group to be sequenced is collected for each nucleotide sequence cluster, and the signal intensity values ​​of the signals to be sequenced in each signal group to be sequenced under corresponding sequencing cycles and different preset sequencing conditions are obtained to automatically identify the gene sequencing results of the nucleotide sequence cluster corresponding to the signal group to be sequenced. There is no need to consider the impact of signal strengths between different signal groups to be sequenced on the sequencing results, and it is not affected by the number of reads. There is no need to perform separate quality value filtering. Only independent analysis of the signals to be sequenced is required to identify the corresponding base sequence. The sequencing operation process is simple and efficient, greatly reducing the requirements for computing resources, effectively improving the efficiency and accuracy of obtaining gene sequencing results, and significantly reducing the error rate, thereby achieving an increase in yield and mapping rate.

[0322] Example 5

[0323] Figure 11 is a schematic diagram of the structure of an electronic device provided in Example 5 of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable by the processor. When the processor executes the program, the method described in the above embodiment is implemented. The electronic device 30 shown in Figure 11 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.

[0324] As shown in FIG11 , the electronic device 30 may be a general-purpose computing device, such as a server device. Components of the electronic device 30 may include, but are not limited to, the at least one processor 31, the at least one memory 32, and a bus 33 connecting various system components (including the memory 32 and the processor 31).

[0325] The bus 33 includes a data bus, an address bus, and a control bus.

[0326] The memory 32 may include a volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322 , and may further include a read-only memory (ROM) 323 .

[0327] The memory 32 may also include a program / utility 325 having a set (at least one) of program modules 324, such program modules 324 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0328] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32, such as the method in the above embodiments of the present disclosure.

[0329] The electronic device 30 can also communicate with one or more external devices 34 (e.g., a keyboard, pointing device, etc.). This communication can occur via an input / output (I / O) interface 35. Furthermore, the model-generating device 30 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 36. As shown in FIG11 , the network adapter 36 communicates with other modules of the model-generating device 30 via a bus 33. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the model-generating device 30, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.

[0330] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0331] Example 6

[0332] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method in the above embodiment are implemented.

[0333] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0334] In a possible implementation manner, the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps in the method of the above embodiment.

[0335] The program code for executing the present disclosure may be written in any combination of one or more programming languages, and the program code may be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on the remote device.

[0336] Although the above describes specific embodiments of the present invention, it should be understood by those skilled in the art that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention. Therefore, the scope of protection of the present invention is defined by the appended claims.

Claims

1. A gene sequencing method, It is characterized in that The gene sequencing method comprises: Collecting multiple signal groups to be sequenced of multiple nucleotide sequence clusters; Wherein, the plurality of signal groups to be sequenced correspond one-to-one to the plurality of nucleotide sequence clusters, each of the signal groups to be sequenced includes a plurality of signals to be sequenced, and each of the signals to be sequenced corresponds to a sequencing cycle; In each sequencing cycle, respectively obtaining the signal intensity value corresponding to the signal to be sequenced; According to the signal intensity value of each of the signals to be sequenced, the gene sequencing result of the nucleotide sequence cluster corresponding to the group of signals to be sequenced is identified.

2. The gene sequencing method according to claim 1, It is characterized in that The step of respectively obtaining the signal strength values ​​corresponding to the signals to be sequenced comprises: The signal strength values ​​of different preset sequencing channels corresponding to the signal to be sequenced are respectively obtained; or, the signal strength values ​​of different preset sequencing time periods corresponding to the signal to be sequenced are respectively obtained.

3. The gene sequencing method according to claim 1 or 2, It is characterized in that The step of identifying and obtaining the gene sequencing result of the nucleotide sequence cluster corresponding to the signal group to be sequenced according to the signal intensity value of each signal to be sequenced comprises: Determining a corresponding number of signal space distribution points according to the signal intensity value of each signal to be sequenced; Analyzing the signal spatial distribution points using preset processing rules to obtain target analysis results; The gene sequencing result of the nucleotide sequence cluster corresponding to the signal group to be sequenced is obtained based on the target analysis result.

4. The gene sequencing method according to at least one of claims 1 to 3, It is characterized in that The step of determining a corresponding number of signal space distribution points according to the signal intensity value of each signal to be sequenced comprises: Construct a preset two-dimensional space; The spatial position information corresponding to the signal strength value in the preset two-dimensional space is determined to obtain a plurality of signal spatial distribution points distributed in the preset two-dimensional space.

5. The gene sequencing method according to claim 4, It is characterized in that The step of analyzing the signal spatial distribution points using a preset processing rule to obtain a target analysis result comprises: Calculate the Euclidean distance between each of the signal space distribution points and a preset reference point in the preset two-dimensional space; Perform statistical analysis using a first processing rule according to different Euclidean distances to obtain a first analysis result; Performing statistical analysis based on the first analysis result using a second processing rule to obtain a second analysis result; Using the first analysis result and the second analysis result as the target analysis result; The first analysis result corresponds to the identification result of a part of the groups, and the second analysis result corresponds to the identification result of the remaining parts of the groups.

6. The gene sequencing method according to claim 5, It is characterized in that The step of performing statistical analysis using a first processing rule according to different Euclidean distances to obtain a first analysis result includes: Generate corresponding first statistical content according to different Euclidean distances; Wherein, the first statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text; The first statistical content is analyzed using the first processing rule to obtain the first analysis result.

7. The gene sequencing method according to claim 5 or 6, It is characterized in that The step of using the second processing rule to perform statistical analysis based on the first analysis result to obtain the second analysis result includes: determining an identified group based on the first analysis result; Calculate the arm angle value between the center of the identified group and the center of other groups to be identified; Generate corresponding second statistical content according to different arm angle values; Wherein, the second statistical content includes a preset statistical graph, a preset statistical table or a preset statistical text; The second statistical content is analyzed using the second processing rule to obtain the second analysis result.

8. The gene sequencing method according to claim 6 or 7, It is characterized in that When the first statistical content includes a first statistical histogram, the step of analyzing the first statistical content using the first processing rule to obtain the first analysis result includes: Selecting a first longest zero sequence that meets a first screening condition in the first statistical histogram; The first longest zero sequence is a sequence with the longest horizontal length and all vertical values ​​being zero; Obtaining first starting position information of the first longest zero sequence; Determine radius information of a first preset group according to the first starting position information; The signal space distribution points whose Euclidean distance is smaller than the radius information are identified as the first preset group, and the other remaining signal space distribution points are used as other groups to be identified.

9. The gene sequencing method according to claim 7 or 8, It is characterized in that When the second statistical content includes a second statistical histogram, the step of analyzing the second statistical content using the second processing rule to obtain the second analysis result includes: Selecting the second longest zero sequence and the second longest zero sequence that meet the second screening condition in the second statistical histogram; The second longest zero sequence is a sequence with the longest horizontal length and zero vertical values, and the second longest zero sequence is a sequence with the second longest horizontal length and zero vertical values; Based on the second longest zero sequence, the second longest zero sequence and the arm angle value, the other groups to be identified are distinguished and identified respectively.

10. The gene sequencing method according to claim 8 or 9, It is characterized in that The first preset group includes the G group, and the other groups to be identified include the A group, the T group and the C group.

11. The gene sequencing method according to at least one of claims 8 to 10, It is characterized in that Before the step of selecting the first longest zero sequence that meets the first screening condition in the first statistical histogram, the method includes: Obtaining a first number of histograms whose vertical values ​​are zero and are arranged continuously in the first statistical histogram; If the first number is less than a set threshold, the current signal to be sequenced is determined to be a bad pixel, and the current signal to be sequenced is filtered out, and the next sequencing cycle of the signal to be sequenced is continued.

12. The gene sequencing method according to claim 11, It is characterized in that The step of selecting the first longest zero sequence that meets the first screening condition in the first statistical histogram includes: If the first number is greater than or equal to the set threshold, the current signal to be sequenced is determined to be a good point, and the first longest zero sequence is obtained according to the initial position and the end position corresponding to the histogram of the first number.

13. The gene sequencing method according to claim 12, It is characterized in that The gene sequencing method further comprises: Using preset identification information to identify the signal to be sequenced belonging to the good point; After all the signals to be sequenced in the plurality of signal groups to be sequenced are processed, the signals to be sequenced marked with the preset identification information are stored in a preset storage space of the gene sequencing chip.

14. The gene sequencing method according to at least one of claims 1 to 13, It is characterized in that The gene sequencing method is implemented using multi-core parallel processing hardware.

15. A gene sequencing system, It is characterized in that The gene sequencing system comprises: A signal group collection module to be sequenced is used to collect multiple signal groups to be sequenced of multiple nucleotide sequence clusters; Wherein, the plurality of signal groups to be sequenced correspond one-to-one to the plurality of nucleotide sequence clusters, each of the signal groups to be sequenced includes a plurality of signals to be sequenced, and each of the signals to be sequenced corresponds to a sequencing cycle; A signal strength value acquisition module, used to respectively acquire the signal strength value corresponding to the signal to be sequenced in each sequencing cycle; The gene sequencing result acquisition module is used to identify and obtain the gene sequencing result of the nucleotide sequence cluster corresponding to the signal group to be sequenced according to the signal intensity value of each signal to be sequenced.

16. A gene sequencing chip, It is characterized in that The gene sequencing chip includes the gene sequencing system described in claim 14.

17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the gene sequencing method according to any one of claims 1 to 14 is implemented.

18. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the gene sequencing method according to any one of claims 1 to 14 is implemented.