Sequencing Depth Estimation Method, Device, Equipment and Storage Medium for Dual Sequencing
By assigning tags and IDs to the DNA templates, generating reads, counting mutation detection frequency, and determining the probability of stable detection, the problem of stable detection of low-frequency mutations in dual sequencing is solved, and stable detection of low-frequency mutations in dual sequencing technology is achieved.
Patent Information
- Application Number
- CN202310118400.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-10
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-02-10
AI Technical Summary
The existing dual sequencing technology is difficult to ensure that low-frequency mutations are stable detected in multiple experiments, and the existing sequencing depth estimation algorithm can only be performed separately at the read segment level or the DNA template level, which cannot effectively ensure the stable detection of low-frequency mutations.
By assigning tags and template IDs to multiple DNA templates, saturated sequencing data are generated, read counts are generated using zero-truncated negative binomial distribution, mutation templates and mutation support reads are generated, mutation detection frequency is set after sub-downsampling, and these steps are repeated multiple times to estimate the probability of stable detection, and the sequencing depth required for stable detection of mutations is determined.
A sequencing depth estimation method that can be stably detected under dual sequencing technology is provided to ensure that the low-frequency mutations are stably detected in dual sequencing, which is suitable for sequencing depth estimation in different laboratories, improving the accuracy of detection.
Smart Images

Figure CN116189768B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technologies, and in particular, to a method, device, equipment and storage medium for estimating the sequencing depth of dual sequencing. Background Art
[0002] The dual sequencing technology has been widely used in the field of low-frequency mutation detection of ctDNA.
[0003] The dual sequencing technology first clusters reads with the same UMI sequence by using the unique molecule identifier (UMI) technology and constructs single strand consensus sequences (SSCSs), and then integrates SSCSs with complementary UMIs into duplex consensus sequences (DCSs). Mutations that consistently appear in the DCSs are identified as true mutations, thus achieving the distinction from PCR (polymerase chain reaction) errors and sequencing errors. For low-frequency mutation detection, there is still a difficulty with dual sequencing: it cannot ensure that low-frequency mutations are repeatedly and stably detected in multiple experiments. This problem can be solved to a certain extent by increasing the sequencing depth, and the lower the variant allele frequency (VAF) of the mutation, the higher the sequencing depth required for stable detection. Therefore, determining the minimum depth requirement for sequencing is necessary to ensure the correctness of clinical detection.
[0004] Commonly used sequencing depth estimation algorithms estimate at the read level or DNA template level using a simple binomial distribution. Taking the sequencing depth D (total number of reads / total number of DNA templates) and the variant allele frequency VAF as the total number of experiments and the success probability of the binomial distribution (Binomial Distribution) respectively, the number of mutant reads / mutant templates X follows a Binom(D, VAF) distribution. If it is denoted that the mutation is detected when at least k mutant reads / mutant templates are detected, then the corresponding mutation detection probability is P(X≥k). Conversely, if the mutation detection probability P(X≥k) is known to be 95%, then the total number of experiments D corresponding to it is the sequencing depth required to ensure the stable detection of this mutation. The above sequencing depth estimation using the binomial distribution can only be carried out separately at the read level or at the DNA template level. However, when judging whether a mutation is detected based on dual sequencing data, constraints are imposed on both the number of mutant templates and the minimum number of mutant-supporting reads corresponding to each forward and reverse mutant template, that is, the minimum cluster size (family size) corresponding to the mutant template strand. Therefore, in order to ensure the stable detection of low-frequency mutations, a sequencing depth estimation method suitable for dual sequencing technology needs to be designed.
[0005] How to determine experimental parameters such as sequencing depth to ensure the stable detection of low-frequency mutations remains an urgent problem to be solved. Summary of the Invention
[0006] The present invention provides a sequencing depth estimation method, device, equipment and storage medium for dual sequencing, aiming to estimate the sequencing depth for detecting low-frequency mutations under dual sequencing technology.
[0007] In a first aspect, an embodiment of the present invention provides a sequencing depth estimation method for dual sequencing, including:
[0008] According to the proportion of double-stranded templates, forward single-stranded templates and reverse single-stranded templates, tags are assigned to multiple DNA templates in the same proportion, and a template ID is assigned to each of the DNA templates; wherein the number of the DNA templates is the number of detected templates in the saturated sequencing state;
[0009] Generate saturated sequencing data, wherein, based on the zero-truncated negative binomial distribution, the number of reads corresponding to each of the DNA templates is generated, and a read ID is assigned to the reads supporting each of the DNA templates according to the quantitative relationship of the number of reads;
[0010] Generate mutant templates and mutant-supporting reads, wherein, according to the mutation frequency, a corresponding number of the DNA templates are selected as mutant templates, and a mutant template tag is assigned to the mutant templates, and the number of mutant reads corresponding to the mutant templates is counted as the number of mutant-supporting reads;
[0011] Set the mutation detection frequency after secondary downsampling. Specifically, after secondary downsampling the reads to a specified sequencing depth, the mutation detection frequency under the mutation detection rule is statistically calculated;
[0012] Repeat the steps of generating saturated sequencing data, generating mutation templates and mutation-supporting reads, and setting the mutation detection frequency after secondary downsampling multiple times. Take the average value of the mutation detection frequency as the estimated value of the detection probability at the specified sequencing depth. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required for stable mutation detection.
[0013] In a second aspect, an embodiment of the present invention provides a sequencing depth estimation device for dual sequencing, including:
[0014] A template marking module, configured to assign tags to multiple DNA templates in the same proportion according to the proportion of double-stranded templates, forward single-stranded templates, and reverse single-stranded templates, and assign a template ID to each of the DNA templates; wherein, the number of the DNA templates is the number of detected templates in the saturated sequencing state;
[0015] A saturated sequencing data generation module, configured to generate saturated sequencing data. Specifically, generate the number of reads corresponding to each DNA template based on the zero-truncated negative binomial distribution, and assign a read ID to the reads supporting each DNA template according to the quantitative relationship of the number of reads;
[0016] A mutation template read generation module, configured to generate mutation templates and mutation-supporting reads. Specifically, select a corresponding number of the DNA templates as mutation templates according to the mutation frequency, assign a mutation template tag to the mutation templates, and count the number of mutation reads corresponding to the mutation templates as the number of mutation-supporting reads;
[0017] A mutation detection frequency statistics module, configured to set the mutation detection frequency after secondary downsampling. Specifically, after secondary downsampling the reads to a specified sequencing depth, the mutation detection frequency under the mutation detection rule is statistically calculated
[0018] A sequencing depth determination module, configured to repeat the steps of generating saturated sequencing data, generating mutation templates and mutation-supporting reads, and setting the mutation detection frequency after secondary downsampling multiple times. Take the average value of the mutation detection frequency as the estimated value of the detection probability at the specified sequencing depth. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required for stable mutation detection.
[0019] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0020] One or more processors;
[0021] A memory, configured to store one or more programs;
[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement the sequencing depth estimation method for dual sequencing provided in any embodiment of the present invention.
[0023] In a fourth aspect, an embodiment of the present invention provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the sequencing depth estimation method for dual sequencing provided in any embodiment of the present invention when executed by a computer processor.
[0024] A sequencing depth estimation method, device, equipment and storage medium for dual sequencing provided by an embodiment of the present invention, through the identity correspondence and quantity relationship between DNA templates and reads, recommend the sequencing depth during dual sequencing when the mutation frequency and mutation detection rules are known, propose a depth estimation for stable detection of low-frequency mutations for dual sequencing, do not need to generate real base sequences, can recommend the sequencing depth that can ensure stable detection of mutations during dual sequencing, and has great application value in guiding the experimental parameter setting of dual sequencing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flowchart of a sequencing depth estimation method for dual sequencing provided in Embodiment 1 of the present invention;
[0026] Figure 2 It is a schematic structural diagram of a sequencing depth estimation device for dual sequencing provided in Embodiment 2 of the present invention;
[0027] Figure 3 It is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention;
[0028] Figure 4 It is a graph of the fitted asymptotic exponential function curve in an embodiment of the present invention;
[0029] Figure 5 It is a flowchart of the sequencing depth estimation method for dual sequencing in an embodiment of the present invention;
[0030] Figure 6 It is a line graph of the result of sequencing depth estimation in an embodiment of the present invention;
[0031] Figure 7 It is a heat map of the result of sequencing depth estimation in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the convenience of description, only the parts related to the present invention rather than all the structures are shown in the accompanying drawings.
[0033] Embodiment 1
[0034] Figure 1 FIG. is a flowchart of a method for estimating sequencing depth of dual sequencing provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation where the sequencing depth needs to be determined for low-frequency mutation detection in dual sequencing technology. This method can be executed by a sequencing depth estimation device for dual sequencing, which can be implemented by hardware and / or software and is generally integrated in an electronic device, such as a computer device. Specifically, this method includes:
[0035] Step 110: Assign tags to multiple DNA templates in the same proportion according to the proportion of double-stranded templates, forward single-stranded templates, and reverse single-stranded templates, and assign a template ID to each of the DNA templates;
[0036] Among them, the number of the DNA templates is the number of detected templates in the saturated sequencing state. Before assigning tags to multiple DNA templates in the same proportion according to the proportion of double-stranded templates, forward single-stranded templates, and reverse single-stranded templates, and assigning a template ID to each of the DNA templates, it may further include: obtaining preset parameters, where the preset parameters include: the sequencing depth of the current site in the saturated sequencing state is represented as DEP satu 、the number of the detected templates is represented as T satu 、the proportion of double-stranded templates is represented as s, the strand bias degree at the read level is represented as R bias and the strand bias degree at the template level is represented as T bias ; the mutation frequency is represented as VAF and the mutation detection rule is represented as c - n1 + n2; the target depth of downsampling at the current site is represented as DEP target .
[0037] The meanings and values of the above preset parameters are as follows:
[0038] (Ⅰ) The sequencing depth DEP satu in the saturated sequencing state needs to ensure that the sequencing has reached the saturated state at this time, that is, the number of templates that can be detected has reached the saturation value T satu , so the value of DEP satu has no upper limit, but has requirements for the lower limit. Usually, in order to reduce the calculation amount, it is recommended that DEP satu take a smaller value after reaching the saturated state and must be greater than the target depth DEP targetSince the sequencing depth and effective depth of the sample are the averages of the sequencing depths and effective depths of all loci respectively, the saturated sequencing depth and effective depth at the sample level can be used to replace those at the locus level.
[0039] The effective depth T of the plasma sample in the saturated sequencing state satu is determined by the DNA input amount MASS and the capture efficiency of the sequencing technology, where T satu The calculation method: Given the DNA input amount MASS, use the asymptotic exponential function to fit the relationship between the effective depth y and the sequencing depth x of the real sequencing sample, so as to obtain the effective depth T in the saturated sequencing state satu = a. As Figure 4 shown, the effective depths and the original sequencing depths of plasma samples with DNA input amounts of 30(±5) ng, 50(±5) ng, and 80(±5) ng are respectively fitted with the asymptotic exponential function (the curve represents the fitted asymptotic exponential function curve, and the circles represent the points taken from the data loess regression line, and the fitting effects of the three functions all have p - value < 2e - 16), and the corresponding effective depths T in the saturated sequencing state are obtained satu which are 4474, 6635, and 10654 respectively. Here, the y - axis is the average effective depth of the sample estimated for the clustered read pairs (reads pair). Since there may be an overlap between the paired reads of PE100, the effective depth of the overlapping part is recorded as 2 instead of 1, resulting in an over - estimation of the effective depth. Through statistics, a stable proportional relationship is found between the estimated effective depth and the real effective depth: Therefore, the real effective depth T in the saturated sequencing state satu should be 3892, 5772, and 9268 respectively.
[0040] In the saturated sequencing state, the proportion s of double - stranded templates, the strand bias degree R at the read level bias and the strand bias degree T at the template level bias all remain stable, so the values of these parameters can be determined in advance. Denote the total number of reads corresponding to T1 forward templates at the current locus as R1, and the total number of reads corresponding to T2 reverse templates as R2. Since R1 + R2 = DEP satu and DEP satu is known, and considering the existence of strand bias at the read level: Then there is Given the proportion s of double - stranded templates and the strand bias degree T at the template level bias , the proportion of forward single - stranded templates can be obtained The proportion of reverse single - stranded templates is Thus, the total number of forward templates T1 = T satu ·(s + t); the total number of reverse templates T2 = T satu ·(1 - t). Thus, the expected number of reads corresponding to each forward template can be obtained as The expected number of reads corresponding to each reverse template is The value range of the double-stranded template ratio s is [0, 1], and from the calculation formulas of t and 1 - s - t, it can be known that s ≤ T bias And Usually, the strand bias T at the template level bias The default value is 1, and the strand bias at the read level varies more than that at the template level and R bias ≥ 0.
[0041] (II) It is known that there is a mutation with a mutation frequency of VAF at this locus, and the requirement for successful detection of this mutation is to be able to detect at least c mutant templates that meet the minimum variant support read count of n1 + n2, abbreviated as "c - n1 + n2". The mutation detection rule simultaneously constrains the number of mutant templates (≥ c) and the number of mutant reads corresponding to each forward template (≥ n1 or ≥ n2) and the number of mutant reads corresponding to the reverse template (≥ n2 or ≥ n1). Here, it is assumed that n1 ≥ n2.
[0042] (III) It is known that the target depth of downsampling at this locus is DEP target , by repeatedly downsampling DEP satu reads multiple times to DEP target reads to simulate the situation under different sequencing depths.
[0043] Step 120: Generate saturated sequencing data;
[0044] Among them, based on the zero-truncated negative binomial distribution, generate the number of reads corresponding to each of the DNA templates, and assign read IDs to the reads supporting each of the DNA templates according to the quantity relationship of the number of reads.
[0045] Step 130: Generate mutant templates and mutant support reads;
[0046] Among them, select a corresponding number of the DNA templates as mutant templates according to the mutation frequency, assign mutant template tags to the mutant templates, and count the number of mutant reads corresponding to the mutant templates as the mutant support read count.
[0047] Step 140: Set the mutation detection frequency after sub-downsampling;
[0048] Among them, set the mutation detection frequency under the mutation detection rule after sub-downsampling the reads to the specified sequencing depth; the specified depth is the target depth of downsampling at the current locus.
[0049] Step 150: Repeat the steps of generating saturated sequencing data, generating mutant templates and mutant supporting reads, and statistically calculating the mutant detection frequency after the set subsampling multiple times. Take the average value of the mutant detection frequency as an estimated value of the detection probability at the specified sequencing depth. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required to stably detect mutants.
[0050] Among them, steps 120, 130, and 140 are respectively repeated N1, N2, and N3 times. The repetitions of steps 120, 130, and 140 have a nested loop relationship (see Figure 5 ). The innermost N3 - loop is to calculate the mutant detection frequency by repeated subsampling The middle - layer N2 - loop is to repeatedly generate mutant templates and mutant reads; the outermost N1 - loop is to repeat the entire process, including the simulation generation of saturated sequencing data at the beginning, the generation of mutant templates and mutant reads, and then subsampling.
[0051] Optionally, according to the proportion of double - stranded templates, forward single - stranded templates, and negative single - stranded templates, assigning labels to multiple DNA templates in the same proportion and assigning a template ID to each DNA template includes:
[0052] Initializing the number of DNA templates to be the number of detected templates;
[0053] According to the proportion of double - stranded templates, forward single - stranded templates, and negative single - stranded templates, assigning labels representing double - stranded templates, forward single - stranded templates, and negative single - stranded templates to the same proportion of DNA templates among the DNA templates respectively;
[0054] Assigning a unique template ID to all the DNA templates.
[0055] Among them, initialize T satu templates, assign the label "D" representing double - stranded templates, the label "P" representing forward single - stranded templates, and the label "N" representing negative single - stranded templates according to the ratio s:t:(1 - s - t) respectively, and assign a unique template ID to all templates.
[0056] Optionally, generating the number of reads corresponding to each DNA template based on the zero - truncated negative binomial distribution and assigning a read ID to the reads supporting each DNA template according to the quantitative relationship of the number of reads includes:
[0057] Using the zero - truncated negative binomial distribution to generate the number of reads corresponding to each DNA template;
[0058] Using the expected value E1 of the number of reads corresponding to each forward DNA template and the expected value E2 of the number of reads corresponding to each reverse DNA template as the expected values of the zero-truncated negative binomial distribution, the number of reads corresponding to each forward template and the number of reads corresponding to each reverse template are generated respectively;
[0059] Adopting the representation form of the expected value and the dispersion parameter α of the negative binomial distribution, by solving the expected value relationship formula between the zero-truncated negative binomial distribution and the standard negative binomial distribution:
[0060]
[0061] The expected value of the standard negative binomial distribution can be obtained and where the α parameter of the standard negative binomial distribution is obtained through pre-statistics and parameter adjustment. For example, the value of the α parameter is 0.3;
[0062] Generate T1 random numbers for the double-stranded forward template and the forward single-stranded template according to the standard negative binomial distribution When the generated random number is 0, continue to generate a non-zero random number to obtain a random number that follows the zero-truncated negative binomial distribution;
[0063] Generate T2 random numbers for the double-stranded reverse template and the reverse single-stranded template according to the standard negative binomial distribution The above random numbers represent the number of reads corresponding to each of the DNA templates. According to the number of reads corresponding to each of the DNA templates, read IDs are assigned to the reads supporting each of the DNA templates. In this way, the DNA templates and the corresponding reads in the saturated state are generated, and at the same time, the identity correspondence relationship between them is obtained.
[0064] Optionally, the step of selecting a corresponding number of the DNA templates as mutant templates according to the mutation frequency, assigning a mutant template label to the mutant templates, and counting the number of mutant reads corresponding to the mutant templates as the number of mutant supporting reads includes:
[0065] According to Determine the number of mutant templates selected, where VAF represents the mutation frequency, and T satu represents the number of detected templates in the saturated sequencing state, represents the number of mutant templates selected;
[0066] Randomly select pieces of the DNA templates from the DNA templates as the mutant templates and assign the mutant template labels;
[0067] The reads corresponding to the mutant templates are the mutant reads, and the number of the mutant reads is counted and recorded as
[0068] Among them, the number of mutant templates obtained by calculation is usually not an integer but a positive real number, including decimal places. By controlling the value of the number of mutant templates in the repeated execution of this step and the number of times, the purpose of improving the model accuracy can be achieved.
[0069] Optionally, after subsampling the reads a set number of times to a specified depth, the mutation detection frequency under the mutation detection rule is statistically analyzed, including:
[0070] Subsample the reads a set number of times to the target depth DEP of the current site downsampling target After that, statistically analyze the frequency of mutations detected when the detection rule is c - n1 + n2 Set the requirement for successful mutation detection to be able to detect at least c mutant templates that meet the minimum number of mutant reads of n1 + n2, and the detection rule is denoted as c - n1 + n2;
[0071] Randomly sample DEP satu reads from DEP target reads;
[0072] Retain the sampled DEP target reads. Correspondingly, the number of DNA templates, the number of mutant templates, and the number of mutant reads are reduced;
[0073] Statistically analyze the number of mutant templates x that meet the n1 + n2 condition among T m mutant templates. If x ≥ c, it means the mutation is detected; otherwise, the mutation is not detected;
[0074] Statistically analyze the frequency of mutations detected when the detection rule is c - n1 + n2 in the set number of downsamplings
[0075] Optionally, repeat the steps of generating saturated sequencing data, generating mutant templates and mutant supporting reads, and statistically analyzing the mutation detection frequency after the set number of downsamplings multiple times, and take the average value of the mutation detection frequency as an estimated value of the detection probability at the specified sequencing depth. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required for stable mutation detection, including:
[0076] Repeat the steps of generating saturated sequencing data, generating mutant templates and mutant supporting reads, and statistically analyzing the mutation detection frequency after the set number of downsamplings N1, N2, and N3 times respectively, and for the detection frequency obtained from the step of statistically analyzing the mutation detection frequency Take the mean value as the estimated value of the mutation detection probability. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required to stably detect mutations. By way of example, the set threshold can be set to 95%.
[0077] Among them, as Figure 5 shown in the repetition of simulating the generation of saturated sequencing data, simulating the generation of mutant templates and mutant support reads, and determining whether the mutation detection conditions are met after downsampling the reads is a nested loop relationship. By calculating the average value of N1*N2 detection frequencies the estimated detection probability of a mutation with a mutation frequency of VAF in a DNA input amount of MASS and a sequencing depth of DEP target and a detection rule of c - n1 + n2 can be obtained. The more the repetition times N1, N2, and N3 take values, the more accurate the model output result is.
[0078] According to the above method, the relationship curve between the detection probability and the sequencing depth of duplex sequencing for mutations with different VAFs can be obtained under the known DNA input amount and mutation detection rule. For example, using the following input parameters:
[0079] (Ⅰ) In the saturated sequencing state, the effective depths T of the sites corresponding to DNA input amounts of 30, 50, and 80 ng are satu 3892, 5772, and 9268 respectively; the proportion s of double-stranded templates is 0.55; the strand bias degree R at the read level bias is 1.0; and the strand bias degree T at the template level bias is 1.0.
[0080] (Ⅱ) Mutation frequencies VAF: 0.0002, 0.0005, 0.001, 0.002, 0.005, 0.008; mutation detection rules c - n1 + n2: 2 - 1 + 0, 4 - 1 + 0, 2 - 2 + 0, 1 - 1 + 1.
[0081] (Ⅲ) The target depth DEP of downsampling target : 2000, 5000, 8000, 10000, 15000, 20000, 25000, 30000, 35000, 40000, 45000.
[0082] The curve graph output by this embodiment shows the results of this embodiment under different combinations of input parameters ( Figure 6) Among them, the graphs in the same row represent the same DNA input amount, the graphs in the same column represent the same detection rule, and the lines of different colors in each sub-graph correspond to different VAFs. It can be seen from the figure that the detection frequencies of mutations with all VAFs either remain 0 all the time or increase to a certain value and then remain stable as the sequencing depth increases; at the same sequencing depth, generally, the higher the VAF of a mutation, the higher its detection frequency. The above rules are consistent with experience.
[0083] When the DNA input amount (effective depth in the saturated state), detection rule, VAF, and target depth DEP target are determined, the corresponding detection frequency value can be obtained. Taking the detection rule 4-1+0 as an example, for a mutation with a VAF of 0.001, when the DNA input amount is 50 ng, the sequencing depth reaches 18000X to ensure the stable detection of this mutation (the detection probability reaches 95%); if the DNA input amount is increased to 80 ng, the stable detection can be achieved when the sequencing depth is 12000x. However, at the same but limited sequencing depth, whether increasing the DNA input amount can bring an increase in the detection frequency needs to be discussed case by case. When the detection rule is 1-1+1 and the sequencing depth is 10000X, for a mutation with a VAF of 0.001, when the DNA input amount is increased from 30 ng to 80 ng, it does not bring an increase in the detection frequency but instead decreases it. Based on the results in the graph, it can be found that for the two cases of the detection rule c-n1+n2, given the previous assumption that n1≥n2, when n2 = 0, increasing the DNA input amount corresponds to an increase in the effective depth, and the possibility of detecting the mutant template increases, thus increasing the detection frequency; but when n2>0, the influence of the DNA input amount on the detection frequency needs to be determined according to the output results of this embodiment because a balance needs to be struck between increasing the effective depth and increasing the expected value of the number of reads corresponding to each template by reducing the effective depth to more easily meet the strict detection rule.
[0084] The detection frequencies in the above results are the average values of N1*N2 detection frequencies When the value of N1*N2 is larger, it is closer to the theoretical detection probability. For the N1*N2 detection frequencies In addition to their mean values, their standard deviations were also statistically analyzed, and the relationship between the standard deviation and the sequencing depth was observed. As shown, the standard deviation of the detection frequency shows the characteristic of first increasing and then decreasing, and when the mean value of the detection frequency reaches more than 95%, the standard deviation of the detection frequency tends to 0, indicating that the mutation can be stably detected at the current sequencing depth. Figure 6
[0085] The technical solution of this embodiment can obtain the detection probabilities of mutations with different VAFs at different sequencing depths under the known DNA input amount and mutation detection rules; under the condition of ensuring stable detection of mutations, it can recommend the combination of DNA input amount and sequencing depth to be used during duplex sequencing under the known mutation frequency and mutation detection rules. As Figure 7 shown, a heat map of the results under different combinations of input parameters. Under the known mutation frequencies (A - F correspond to 0.0002, 0.0005, 0.001, 0.002, 0.005, 0.008 respectively) and mutation detection rules (4 - 1 + 0), it is possible to recommend the combination of DNA input amount (y - axis) and sequencing depth (x - axis) to be used during duplex sequencing.
[0086] The technical solution of this embodiment, to solve the minimum depth requirement for sequencing in the detection of low - frequency mutations and ensure that low - frequency mutations can be stably detected, considering the read - segment clustering process, the sequencing depth estimation method for duplex sequencing can provide an adapted sequencing depth estimation scheme for duplex sequencing technologies from different laboratories, thereby achieving stable detection of mutations, especially low - frequency mutations.
[0087] Embodiment Two
[0088] Figure 2 is a schematic structural diagram of a sequencing depth estimation device for duplex sequencing provided by the second embodiment of the present invention. As Figure 2 shown, the sequencing depth estimation device for duplex sequencing includes: a template tagging module 210, a saturated sequencing data generation module 220, a mutant template read - segment generation module 230, a mutant detection frequency statistics module 240, and a sequencing depth determination module 250, where
[0089] The template tagging module 210 is used to assign tags to multiple DNA templates in the same proportion according to the proportion of double - stranded templates, forward single - stranded templates, and reverse single - stranded templates, and assign a template ID to each of the DNA templates; where the number of the DNA templates is the number of detected templates in the saturated sequencing state;
[0090] The saturated sequencing data generation module 220 is used to generate saturated sequencing data, where the number of read - segments corresponding to each DNA template is generated based on the zero - truncated negative binomial distribution, and a read - segment ID is assigned to the read - segments supporting each DNA template according to the quantitative relationship of the number of read - segments;
[0091] The mutant template read - segment generation module 230 is used to generate mutant templates and mutant - supporting read - segments, where a corresponding number of the DNA templates are selected as mutant templates according to the mutation frequency, a mutant template tag is assigned to the mutant templates, and the number of mutant read - segments corresponding to the mutant templates is counted as the number of mutant - supporting read - segments;
[0092] The mutation detection frequency statistics module 240 is used to set the mutation detection frequency after secondary downsampling. Specifically, it calculates the mutation detection frequency under the mutation detection rule after downsampling the reads to a specified sequencing depth; the specified depth is the target depth of downsampling at the current locus.
[0093] The sequencing depth determination module 250 is used to repeat the steps of generating saturated sequencing data, generating mutation templates and mutation-supporting reads, and calculating the mutation detection frequency after secondary downsampling multiple times. It takes the average value of the mutation detection frequencies as an estimated value of the detection probability at the specified sequencing depth, and the sequencing depth corresponding to this value reaching the set threshold is the sequencing depth required for stable mutation detection.
[0094] Optionally, the sequencing depth estimation device for duplex sequencing further includes:
[0095] A parameter acquisition module, which is used to obtain preset parameters before assigning tags to multiple DNA templates in the same proportion according to the proportion of double-stranded templates, forward single-stranded templates, and reverse single-stranded templates, and assigning a template ID to each DNA template. The preset parameters include: the sequencing depth of the current locus in the saturated sequencing state, the number of detected templates, the proportion of double-stranded templates, the strand bias at the read level, and the strand bias at the template level, the mutation frequency, and the mutation detection rule; the target depth of downsampling at the current locus.
[0096] Optionally, the template marking module 210 is specifically used for:
[0097] Initializing the number of DNA templates to be the number of detected templates;
[0098] According to the proportion of double-stranded templates, forward single-stranded templates, and reverse single-stranded templates, assigning tags representing double-stranded templates, forward single-stranded templates, and reverse single-stranded templates to the same proportion of the DNA templates respectively;
[0099] Assigning a unique template ID to all the DNA templates.
[0100] Optionally, the saturated sequencing data generation module 220 is specifically used for:
[0101] Using the zero-truncated negative binomial distribution to generate the number of reads corresponding to each DNA template;
[0102] Taking the expected value E1 of the number of reads corresponding to each forward DNA template and the expected value E2 of the number of reads corresponding to each reverse DNA template as the expected values of the zero-truncated negative binomial distribution to generate the number of reads corresponding to each forward template and the number of reads corresponding to each reverse template respectively;
[0103] Using the representation of the expectation and the dispersion parameter α of the negative binomial distribution, by solving the formula for the expectation relationship between the zero-truncated negative binomial distribution and the standard negative binomial distribution:
[0104]
[0105] The expectation of the standard negative binomial distribution can be obtained and where the α parameter of the standard negative binomial distribution is obtained through pre-statistics and parameter adjustment;
[0106] Generate T1 random numbers for the forward template and the forward single-stranded template of the double-stranded according to the standard negative binomial distribution When the generated random number is 0, continue to generate a non-zero random number to obtain a random number that follows the zero-truncated negative binomial distribution;
[0107] Generate T2 random numbers for the reverse template and the reverse single-stranded template of the double-stranded according to the standard negative binomial distribution The above random numbers represent the number of reads corresponding to each of the DNA templates, and read IDs are assigned to the reads supporting each of the DNA templates according to the number of reads corresponding to each of the DNA templates.
[0108] Optionally, the mutant template read generation module 230 is specifically configured to:
[0109] According to Determine the number of mutant templates selected, where VAF represents the mutation frequency, T satu represents the number of detected templates in the saturated sequencing state, represents the number of mutant templates selected;
[0110] Randomly select pieces of the DNA templates as the mutant templates and assign labels to the mutant templates;
[0111] The reads corresponding to the mutant templates are the mutant reads, and the number of the mutant reads is counted and recorded as
[0112] Optionally, the mutant detection frequency statistics module 240 is specifically configured to:
[0113] Perform a set number of downsamplings on the reads to the target depth DEP at the current site target After that, count the frequency of mutant detection when the detection rule is c - n1 + n2 Set the requirement for successful mutant detection to be able to detect at least c mutant templates that meet the minimum number of mutant reads of n1 + n2, and the detection rule is denoted as c - n1 + n2;
[0114] From DEPsatu Random sampling of DEP in read segments target Read segments;
[0115] Retain the sampled DEP target Read segments. Correspondingly, the number of DNA templates, the number of mutant templates, and the number of mutant read segments are reduced;
[0116] Statistical analysis at T m The number of mutant templates x that meet the condition of n1 + n2 in the T mutant templates is counted. If x ≥ c, it indicates that the mutation is detected; otherwise, the mutation is not detected;
[0117] Statistical analysis of the frequency of mutation detection when the detection rule is c - n1 + n2 in the set subsampling
[0118] Optionally, the sequencing depth determination module 250 is specifically configured to:
[0119] Repeat the steps of generating saturated sequencing data, generating mutant templates and mutant supporting read segments, and statistically analyzing the mutation detection frequency after the set subsampling N1, N2, and N3 times respectively, and for the detection frequency obtained from the step of statistically analyzing the mutation detection frequency Take the average value as the estimated value of the mutation detection probability at the specified sequencing depth. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required to stably detect mutations.
[0120] The sequencing depth estimation device for dual sequencing provided by the embodiments of the present invention can execute the sequencing depth estimation method for dual sequencing provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0121] Embodiment III
[0122] Figure 3 It is a schematic structural diagram of an electronic device provided by Embodiment III of the present invention. As Figure 3 shown, the electronic device includes a processor 310, a memory 320, an input device 330, and an output device 340; the number of processors 310 in the electronic device can be one or more, Figure 3 Taking one processor 310 as an example; the processor 310, the memory 320, the input device 330, and the output device 340 in the electronic device can be connected through a bus or other means, Figure 3 Taking connection through a bus as an example.
[0123] The memory 320, being a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the sequencing depth estimation method of duplex sequencing in the embodiments of the present invention (for example, the template tagging module 210, the saturated sequencing data generation module 220, the mutant template read segment generation module 230, the mutant detection frequency statistics module 240, and the sequencing depth determination module 250 in the duplex sequencing depth estimation device). The processor 310 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 320, thereby implementing the above-mentioned sequencing depth estimation method of duplex sequencing.
[0124] The memory 320 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 320 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 320 may further include a memory remotely set relative to the processor 310, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0125] The input device 330 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device. The output device 340 may include a display device such as a display screen.
[0126] Embodiment Four
[0127] Embodiment Four of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a sequencing depth estimation method of duplex sequencing when executed by a computer processor, including:
[0128] According to the proportion of double-stranded templates, forward single-stranded templates, and reverse single-stranded templates, tags are assigned to multiple DNA templates in the same proportion, and a template ID is assigned to each of the DNA templates; wherein, the number of the DNA templates is the number of detected templates in the saturated sequencing state;
[0129] Generate saturated sequencing data, wherein, based on the zero-truncated negative binomial distribution, the number of read segments corresponding to each of the DNA templates is generated, and a read segment ID is assigned to the read segments supporting each of the DNA templates according to the quantitative relationship of the number of read segments;
[0130] Generate mutant templates and mutant-supporting reads. Among them, select a corresponding number of the DNA templates as mutant templates according to the mutation frequency, assign mutant template tags to the mutant templates, and count the number of mutant reads corresponding to the mutant templates as the number of mutant-supporting reads;
[0131] Set the mutation detection frequency after secondary downsampling. Among them, set the secondary downsampling of the reads to a specified sequencing depth and then count the mutation detection frequency under the mutation detection rule; the specified depth is the target depth of downsampling at the current site;
[0132] Repeat the steps of generating saturated sequencing data, generating mutant templates and mutant-supporting reads, and setting the mutation detection frequency after secondary downsampling multiple times. Take the average value of the mutation detection frequencies as the estimated value of the detection probability at the specified sequencing depth. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required to stably detect mutations.
[0133] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the method operations described above, and can also execute related operations in the sequencing depth estimation method of double sequencing provided by any embodiment of the present invention.
[0134] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disc of a computer, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0135] It should be noted that in the embodiments of the above-mentioned sequencing depth estimation device for double sequencing, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0136] Although the present invention has been described in detail above with general descriptions, specific embodiments and experiments, modifications or improvements can be made to it on the basis of the present invention, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A method for estimating sequencing depth of dual sequencing, characterized in that, Including: According to the proportion of double-stranded templates, forward single-stranded templates, and negative single-stranded templates, assign labels to DNA templates and assign a template ID to each of the DNA templates; wherein, the number of the DNA templates is the number of detected templates in the saturated sequencing state; Generate saturated sequencing data, wherein, based on the zero-truncated negative binomial distribution, generate the number of reads corresponding to each of the DNA templates, and assign a read ID to the reads supporting each of the DNA templates according to the quantitative relationship of the number of reads; Generate mutant templates and mutant-supporting reads, wherein, select a corresponding number of the DNA templates as mutant templates according to the mutation frequency, assign a mutant template label to the mutant templates, and count the number of mutant reads corresponding to the mutant templates as the number of mutant-supporting reads; Set the mutation detection frequency after secondary subsampling, wherein, set the secondary subsampling of the reads to a specified sequencing depth and then count the mutation detection frequency under the mutation detection rule; Repeat the steps of generating saturated sequencing data, generating mutant templates and mutant-supporting reads, and setting the mutation detection frequency after secondary subsampling multiple times, take the average value of the mutation detection frequencies as an estimated value of the detection probability at the specified sequencing depth, and the sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required to stably detect mutations; Among them, the generating the number of reads corresponding to each of the DNA templates based on the zero-truncated negative binomial distribution and assigning a read ID to the reads supporting each of the DNA templates according to the quantitative relationship of the number of reads includes: Use the zero-truncated negative binomial distribution to generate the number of reads corresponding to each of the DNA templates; The expectation of the number of reads corresponding to each forward DNA template and the expectation of the number of reads corresponding to each reverse DNA template are used as the expected values of the zero-truncated negative binomial distribution to generate the number of reads corresponding to each forward template and the number of reads corresponding to each reverse template, respectively; Using the expectation and dispersion parameters of the negative binomial distribution In the representation form, by solving the expectation relationship formula between the zero-truncated negative binomial distribution and the standard negative binomial distribution: The expectation of the standard negative binomial distribution can be obtained and , where the parameters of the standard negative binomial distribution are obtained through pre-statistics and parameter tuning; Generate random numbers for the forward template of the double strand and the forward single-strand template according to the standard negative binomial distribution Generate random numbers; when the generated random number is 0, continue to generate a non-zero random number to obtain a random number that follows the zero-truncated negative binomial distribution; Assign the negative-strand template of the double-strand and the negative-strand single-strand template according to the standard negative binomial distribution Generate random numbers; the above random numbers represent the number of reads corresponding to each of the DNA templates, and read IDs are assigned to the reads supporting each of the DNA templates according to the number of reads corresponding to each of the DNA templates; The selecting a corresponding number of the DNA templates as mutant templates according to the mutation frequency, assigning a mutant template label to the mutant templates, and counting the number of mutant reads corresponding to the mutant templates as the number of mutant-supporting reads includes: According to , determine the number of mutant templates selected, where represents the mutation frequency,[[]] represents the number of detected templates in the saturation sequencing state,[[]] represents the number of mutant templates selected; Randomly select of the DNA templates as the mutation template, and assign a label to the mutation template; The read segment corresponding to the mutation template is the mutation read segment, and the number of the mutation read segments is counted and denoted as .
2. The method according to claim 1, wherein Before assigning labels to multiple DNA templates in the same proportion according to the proportion of double-stranded templates, forward single-stranded templates, and negative single-stranded templates and assigning a template ID to each of the DNA templates, it further includes: Obtain preset parameters, wherein, the preset parameters include: the sequencing depth of the current site in the saturated sequencing state, the number of detected templates, the proportion of double-stranded templates, the strand bias degree at the read level, and the strand bias degree at the template level, the mutation frequency, and the mutation detection rule; the target depth of subsampling at the current site.
3. The method according to claim 2, wherein The assigning labels to multiple DNA templates in the same proportion according to the proportion of double-stranded templates, forward single-stranded templates, and negative single-stranded templates and assigning a template ID to each of the DNA templates includes: Initialize the DNA templates with the number of detected templates; According to the proportion of double-stranded templates, forward single-stranded templates, and negative single-stranded templates, assign labels representing double-stranded templates, forward single-stranded templates, and negative single-stranded templates to the DNA templates in the same proportion among the DNA templates; Assign a unique template ID to all the DNA templates.
4. The method according to claim 1, wherein The setting the secondary subsampling of the reads to a specified sequencing depth and then counting the mutation detection frequency under the mutation detection rule includes: Perform the set number of subsamplings on the read segments to the target depth at the current site After that, count the frequency at which mutations are detected when the detection rule is - + ; Set the requirement for successful mutation detection to be that at least mutation templates satisfy + + the minimum number of mutant read segments, and the detection rule is denoted as - + ; Sample randomly from reads reads; Retain the sampled reads, and correspondingly, the number of the DNA templates, the number of the mutant templates, and the number of the mutant reads are reduced; Count the number of mutant templates that meet the conditions in mutant templates, and if + condition is met, it means the mutation is detected; otherwise, the mutation is not detected. , if Count the frequency of mutations detected when the detection rule is - + in the set number of downsamplings .
5. The method according to claim 4, characterized in that, Repeatedly perform the steps of generating saturated sequencing data, generating mutant templates and mutant-supported reads, and statistically calculating the mutant detection frequency after the set number of downsamplings. Take the average value of the mutant detection frequencies as an estimated value of the detection probability at a specified sequencing depth. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required for stable mutant detection, including: Repeat the steps of generating the saturation sequencing data, generating the mutant templates and mutant supporting reads, and statistically analyzing the mutant detection frequency after the set sub-sampling respectively , and times, and take the mean value of the detection frequencies obtained from the step of statistically analyzing the mutant detection frequency as the estimated value of the mutant detection probability at the specified sequencing depth. The sequencing depth corresponding to the value reaching the set threshold is the sequencing depth required for stable mutant detection.
6. A sequencing depth estimation device for dual sequencing, characterized in that, Including: A template tagging module, which, according to the proportion of double-stranded templates, forward single-stranded templates, and reverse single-stranded templates, assigns tags to multiple DNA templates in the same proportion and assigns a template ID to each of the DNA templates; wherein, the number of the DNA templates is the number of detected templates in the saturated sequencing state; A saturated sequencing data generation module, which is used to generate saturated sequencing data. Among them, based on the zero-truncated negative binomial distribution, the number of reads corresponding to each of the DNA templates is generated, and read IDs are assigned to the reads supporting each of the DNA templates according to the quantitative relationship of the number of reads; A mutant template read generation module, which is used to generate mutant templates and mutant-supported reads. Among them, according to the mutation frequency, a corresponding number of the DNA templates are selected as mutant templates, and mutant template tags are assigned to the mutant templates, and the number of mutant reads corresponding to the mutant templates is statistically calculated as the number of mutant-supported reads; A mutant detection frequency statistics module, which is used to statistically calculate the mutant detection frequency after the set number of downsamplings. Among them, after downsampling the reads to the specified sequencing depth for the set number of times, the mutant detection frequency under the mutant detection rule is statistically calculated; A sequencing depth determination module, which repeatedly performs the steps of generating saturated sequencing data, generating mutant templates and mutant-supported reads, and statistically calculating the mutant detection frequency after the set number of downsamplings. Take the average value of the mutant detection frequencies as an estimated value of the detection probability at a specified sequencing depth. The sequencing depth corresponding to when this value reaches the set threshold is the sequencing depth required for stable mutant detection; Among them, the saturated sequencing data generation module is specifically used for: Using the zero-truncated negative binomial distribution to generate the number of reads corresponding to each of the DNA templates; The expectation of the number of reads corresponding to each forward DNA template and the expectation of the number of reads corresponding to each reverse DNA template are used as the expected values of the zero-truncated negative binomial distribution to generate the number of reads corresponding to each forward template and the number of reads corresponding to each reverse template, respectively; Using the expectation and dispersion parameters of the negative binomial distribution in the form of representation, by solving the expectation relationship formula between the zero-truncated negative binomial distribution and the standard negative binomial distribution: The expectation of the standard negative binomial distribution can be obtained and , where the parameters of the standard negative binomial distribution are obtained through pre-statistics and parameter tuning; Generate random numbers for the forward template and forward single-strand template of the double strand according to the standard negative binomial distribution Generate random numbers; when the generated random number is 0, continue to generate a non-zero random number to obtain a random number that follows the zero-truncated negative binomial distribution; Assign the negative-strand template of the double-strand and the negative-strand single-strand template according to the standard negative binomial distribution Generate random numbers; the above random numbers represent the number of reads corresponding to each of the DNA templates, and read IDs are assigned to the reads supporting each of the DNA templates according to the number of reads corresponding to each of the DNA templates; The mutant template read generation module, specifically used for: According to , determine the number of mutant templates selected, where represents the mutation frequency, represents the number of detected templates in the saturation sequencing state, represents the number of mutant templates selected; Randomly select of the DNA templates as the mutation template, and assign a label to the mutation template; The read segment corresponding to the mutation template is the mutation read segment, and the number of the mutation read segments is counted and denoted as .
7. An electronic device, characterized in that, Including: One or more processors; A memory, which is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the sequencing depth estimation method for dual sequencing as described in any one of claims 1-5.
8. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the sequencing depth estimation method for dual sequencing as described in any one of claims 1-5 when executed by a computer processor.
Citation Information
Patent Citations
Molecular label counting adjustment methods
CN109074430A
Single-cell gene fusion detection method
CN112509639A