Copy Number Variant Cola
A hidden Markov model-based method addresses NGS limitations in detecting copy number variants by adjusting for sequencing depth artifacts and spurious probes, ensuring accurate copy number determination and variant aberration detection, thereby improving disease diagnosis.
Patent Information
- Application Number
- JP2024042482
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-20
- Filing Date
- 2024-03-18
- Publication Date
- 2025-09-08
- Estimated Expiration
- 2039-05-31
AI Technical Summary
Next-generation sequencing (NGS) technologies face challenges in accurately detecting copy number variants (CNVs) due to analytical limitations from sample preparation, sequencing, mapping, target GC content, target size, and sequence complexity, which affect read depth and accuracy, making it difficult to use for detecting exon-level copy number variants or rearrangements, and existing methods lack effective evaluation of copy number variant callers without sufficient positive controls.
A method for evaluating sample-specific performance of copy number variant models using a hidden Markov model (HMM) that parameterizes the model based on actual sequence read numbers, generates synthetic variants, and adjusts for sequencing depth artifacts and spurious capture probes to determine accurate copy numbers and variant aberrations.
The method provides accurate and reliable determination of copy numbers and variant aberrations, improving diagnostic yield and enabling better genetic susceptibility analysis for diseases like cancer by enhancing the accuracy of copy number variant detection.
Smart Images

Figure 0007735457000047 
Figure 0007735457000048 
Figure 0007735457000049
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 681,517, filed June 6, 2018, and U.S. Provisional Patent Application No. 62 / 733,842, filed September 20, 2018, each of which is incorporated herein in its entirety, including all tables, figures, and claims.
[0002] The present invention relates to a method for determining the copy number of a gene region of interest. [Background technology]
[0003] Many important advances have been made in understanding genetic susceptibility to cancer and other diseases. Identifying mutations associated with hereditary cancer syndromes and other diseases can lead to reduced morbidity and mortality through targeted risk management options. The traditional approach for germline testing has been to test for mutations in a single gene or a limited panel of genes using Sanger sequencing. Advances in next-generation sequencing technology and bioinformatics analysis have made it possible to simultaneously test many genes (panel-based testing) at costs comparable to traditional testing. Panel-based testing can offer improved accuracy compared to traditional methods and improved diagnostic yield due to analytical concordance between next-generation sequencing (NGS) results and traditional Sanger sequencing for detecting small mutations such as single nucleotide variants, small deletions, and small insertions.
[0004] Despite advances in NGS technology over the past few years, NGS panels have analytical limitations arising from sample preparation, sequencing, mapping, target GC content, target size, and sequence complexity. These factors affect the relationship between read depth and copy number, which is key to copy number variant calling, and consequently, the accuracy of using NGS technology to detect copy number variants. Such limitations make NGS technology difficult to use for the detection of copy number variants (CNVs), such as exon-level copy number variants, larger insertion or deletion variants, or rearrangements. Scientific studies suggest that many cancers and complex diseases, such as schizophrenia, are at least partially related to copy number variants. Therefore, greater accuracy and consideration of the impact of sequencing depth on noise when relating copy number are particularly desirable. To address this concern, some laboratories complement NGS with microarrays, which introduce a unique level of complexity and bias into calls. Copy number variants provide valuable information needed to improve our understanding and characterization of genetic susceptibility to cancer and other diseases. Therefore, a method for detecting CNVs with high accuracy is desirable.
[0005] Typically, the performance of genetic variant screening is evaluated for concordance with known reference samples. The inherent variability of sequence data, coupled with inadequate quality control (QC) measures, can compromise high CNV calling accuracy. Screen evaluation can be performed using a large number of positive controls with known genetic variants, and screening performance statistics (such as sensitivity or specificity) can be determined. However, if a large number of positive controls, such as controls with rare genetic variant events, are not available, the performance of the genetic variant calling algorithm (i.e., "Cola") or assay cannot be accurately evaluated. While a large number of positive controls with single nucleotide variants (SNVs) are commonly available, the frequency of positive control samples with copy number variants is lower.
[0006] The disclosures of all publications, patents, and patent applications mentioned herein are each incorporated by reference in their entirety. To the extent that any reference incorporated by reference conflicts with the present disclosure, the present disclosure shall control. Summary of the Invention
[0007] Disclosed herein are methods for assessing sample-specific performance of copy number variant models, methods for determining the copy number of a queried segment within a region of interest, and methods for determining copy number variant aberrations within a region of interest.
[0008] In some embodiments, a method for evaluating sample-specific performance of a copy number variant caller comprising a copy number variant model comprises: parameterizing the copy number variant model based on actual sequence read numbers mapped to segments within a region of interest from a test sample to determine one or more copy number variant model parameters; generating a plurality of synthetic copy number variants, each synthetic copy number variant comprising synthetic copy numbers of one or more of the segments, each synthetic copy number represented by a synthetic sequence read number based on the actual sequence read number of the corresponding segment from the test sample; calling copy numbers of one or more segments of the synthetic copy number variants using the copy number variant model and the one or more determined copy number variant model parameters; determining sample-specific performance statistics for the copy number variant caller based on the difference between the called copy numbers and the synthetic copy numbers of the synthetic copy number variants; and evaluating the sample-specific performance of the copy number variant caller based on the sample-specific performance statistics.
[0009] In some embodiments of the methods for evaluating sample-specific performance of copy number variant callers, synthetic sequence read numbers for one or more segments are generated by increasing, decreasing, or maintaining the actual sequence read numbers for the corresponding segments from the test sample in proportion to a predetermined copy number for the one or more segments. In some embodiments, the predetermined copy number is an integer copy number. In some embodiments, the predetermined copy number is a non-integer copy number.
[0010] In some embodiments of methods for evaluating sample-specific performance of copy number variant callers, synthetic sequence read numbers are generated by sampling a binomial distribution with probability of success equal to m / x and number of trials equal to the actual sequence read number for the corresponding segment from the test sample, where m is the synthetic copy number of the segment in the synthetic copy number variant and x is the hypothesized copy number of the corresponding segment from the test sample.
[0011] In some embodiments of a method for evaluating sample-specific performance of a copy number variant caller, the synthetic sequence read numbers are generated by sampling sequence read numbers as a negative binomial distribution with a probability of success equal to m / x and a number of successes equal to the actual sequence read number for the corresponding segment from the test sample, where m is the synthetic copy number of a segment in the synthetic copy number variant and x is the hypothesized copy number of the corresponding segment from the test sample, and adding the sampled sequence read numbers to the actual sequence read number for the corresponding segment from the test sample. In some embodiments, the synthetic sequence read numbers are generated by sampling sequence read numbers as the expected value of the negative binomial distribution.
[0012] In some embodiments of the methods for evaluating sample-specific performance of copy number variant callers, the copy number variant model is a Hidden Markov Model. In some embodiments, the Hidden Markov Model is a Hidden Markov Model that: (i) estimates the sample-specific performance of the queried segment or the queried segment; The method includes (i) one or more hidden states including copy numbers corresponding to multiple subsegments within the queried segment; (ii) an observed state including an actual or synthetic sequence read number for the queried segment; and (iii) a copy number likelihood model based on the expected actual or synthetic sequence read number for the queried segment. In some embodiments, the method includes determining the copy number likelihood model. In some embodiments, parameterizing the hidden Markov model includes adjusting the copy number likelihood model to fit the actual sequence read numbers mapped to the queried segment from the test sample. In some embodiments, the copy number likelihood model includes a distribution of two or more copy number states. In some embodiments, the copy number likelihood model includes a negative binomial distribution, and the negative binomial distribution is not a Poisson distribution. In some embodiments, the expected actual or composite sequence read numbers are based on the representative mapped sequence read numbers at the segments corresponding to the interrogated segment across multiple samples and the representative mapped sequence read numbers at the segments in the test sample, where the representative mapped sequence read numbers at the segments corresponding to the interrogated segment across multiple samples or the representative mapped sequence read numbers at the segments in the test sample are normalized representative values. In some embodiments, the copy number likelihood model is adjusted to account for the presence of GC content bias. In some embodiments, the hidden Markov model includes transition probabilities for copy numbers of the interrogated segment for a given copy number of spatially adjacent segments. In some embodiments, the hidden Markov model includes multiple transition probabilities for copy numbers of subsegments for a given copy number of spatially adjacent subsegments within the interrogated segment. In some embodiments, the transition probabilities consider the length of the representative copy number variant. In some embodiments, the transition probabilities consider the prior probability of copy number variants in the interrogated segment or spatially adjacent segments.In some embodiments, the representative length of the copy number variant, or the probability of the copy number variant at the interrogated segment, is determined based on observations in human populations.
[0013] In some embodiments of the method for evaluating sample-specific performance of a copy number variant caller, parameterizing the copy number variant model includes considering one or more spurious capture probes. In some embodiments, considering one or more spurious capture probes includes weighting one or more observation states of the plurality of observation states with a spurious capture probe indicator. In some embodiments, the spurious capture probe indicator is determined using a Bernoulli process. In some embodiments, considering one or more capture probes as spurious includes using expectation maximization. In some embodiments, if a capture probe is determined to be spurious, sequence reads originating from that capture probe are discarded in the copy number variant model.
[0014] In some embodiments of the method for evaluating sample-specific performance of a copy number variant model, parameterizing the copy number variant model includes accounting for noise in the number of mapped sequence reads.
[0015] In some embodiments of the methods for assessing sample-specific performance of copy number variant callers, the copy number variant model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more copy number variant model parameters.
[0016] In some embodiments of the method for assessing sample-specific performance of a copy number variant caller, the copy number variant model is parameterized by solving a trust-region Newton conjugate gradient algorithm.
[0017] In some embodiments of the method for assessing sample-specific performance of copy number variant callers, the copy number variant model is iteratively parameterized using expectation maximization.
[0018] In some embodiments of methods for evaluating sample-specific performance of copy number variant callers, the method includes mapping actual sequence reads from the test sample to segments within the region of interest and determining the number of actual sequence reads mapped to the segments.
[0019] In some embodiments of methods for assessing sample-specific performance of copy number variant callers, the test sample is enriched using one or more direct target sequence capture probes.
[0020] In some embodiments of methods for evaluating sample-specific performance of copy number variant callers, the methods include calling copy numbers of one or more segments for a test sample.
[0021] In some embodiments of the method for assessing sample-specific performance of a copy number variant caller, the segments comprise spatially adjacent segments.
[0022] In some embodiments of the method for assessing sample-specific performance of copy number variant callers, the sample-specific performance statistic is limit of detection, sensitivity, specificity, precision, recall, accuracy, positive predictive value, or negative predictive value. In some embodiments of the methods for assessing sample-specific performance of copy number variant callers, the sample-specific performance statistic is sensitivity or precision.
[0023] In some embodiments of the methods for evaluating sample-specific performance of copy number variant models, the methods include rejecting the test sample if the sample-specific performance of the copy number variant model is below a desired performance threshold.
[0024] Also described herein is a method for determining the copy number of an interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to the interrogated segment, wherein the test sequence library is enriched using one or more direct target sequence capture probes; (b) determining the number of sequence reads mapped to the interrogated segment; (c) determining a copy number likelihood model based on the expected number of sequence reads mapped to the interrogated segment; and (d) generating a hidden Markov model that includes (i) copy numbers corresponding to the interrogated segment or a plurality of subsegments within the interrogated segment. (ii) an observed state including the determined sequence read numbers mapped to the queried segment; and (iii) a copy number likelihood model; (e) parameterizing the hidden Markov model by adjusting the copy number likelihood model to fit the determined sequence read numbers mapped to the queried segment, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; and (f) determining the most probable copy number of the queried segment based on the parameterized hidden Markov model.
[0025] Further provided herein are methods for determining the copy number of an interrogated segment within a region of interest. A method for mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, the plurality of spatially adjacent segments including a queried segment, the test sequence library being enriched using a plurality of spatially adjacent direct target sequence capture probes, (b) determining the number of sequence reads mapped to each spatially adjacent segment, (c) determining a copy number likelihood model for each spatially adjacent segment based on the expected number of mapped sequence reads in the spatially adjacent segments, and (d) using a hidden Markov model to (i) estimate the copy number of each of the spatially adjacent segments or a plurality of subsegments within each of the spatially adjacent segments. (ii) constructing a hidden Markov model including: (i) a plurality of hidden states including a plurality of sequence reads mapped to each spatially adjacent segment; (ii) a plurality of observed states including sequence read numbers mapped to each spatially adjacent segment; and (iii) a copy number likelihood model for each spatially adjacent segment; (e) parameterizing the hidden Markov model including adjusting each copy number likelihood model to fit the determined sequence read numbers mapped to each spatially adjacent segment, wherein the hidden Markov model is parameterized using an analytical first derivative gradient and second derivative Hessian of one or more parameters in the copy number likelihood model; and (f) determining the most probable copy number of the queried segment based on the parameterized hidden Markov model.
[0026] Also described herein is a method for determining copy number variant aberrations in a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to an interrogated segment within the region of interest, wherein the test sequence library is enriched using one or more direct target sequence capture probes; (b) determining a number of sequence reads mapped to the interrogated segment; (c) determining a copy number likelihood model based on an expected number of sequence reads mapped to the interrogated segment; and (d) a hidden Markov model comprising: (i) one or more hidden states including copy numbers corresponding to the interrogated segment or a plurality of subsegments within the interrogated segment; and (ii) a copy number likelihood model based on the expected number of sequence reads mapped to the interrogated segment. (ii) constructing a hidden Markov model including an observation state including the number of sequence reads mapped to the interrogated segment; and (iii) a copy number likelihood model; (e) parameterizing the hidden Markov model by adjusting the copy number likelihood model to fit the determined number of sequence reads mapped to the interrogated segment, wherein the hidden Markov model is parameterized using an analytical first derivative gradient and second derivative Hessian of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the interrogated segment based on the parameterized hidden Markov model; and (g) determining a copy number variant aberration based on the most probable copy number of the interrogated segment.
[0027] Further described herein is a method for determining copy number variant aberrations in a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, wherein the plurality of spatially adjacent segments includes a queried segment, and wherein the test sequence library is enriched using a plurality of spatially adjacent direct target sequence capture probes; (b) determining a number of sequence reads mapped to each spatially adjacent segment; (c) determining a copy number likelihood model for each spatially adjacent segment based on expected numbers of mapped sequence reads in the spatially adjacent segments; and (d) a hidden Markov model comprising: (i) a plurality of hidden states including the copy number of each of the spatially adjacent segments, or a plurality of subsegments within each of the spatially adjacent segments; and (ii) the number of sequence reads mapped to each spatially adjacent segment. (iii) constructing a hidden Markov model including a plurality of observation states; and (iv) a copy number likelihood model for each spatially adjacent segment; (e) parameterizing the hidden Markov model, including adjusting each copy number likelihood model to fit the determined sequence read numbers mapped to each spatially adjacent segment, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the queried segment based on the parameterized hidden Markov model; and (g) determining a copy number variant aberration based on the most probable copy number of the queried segment.
[0028] Also described herein is a method for determining the copy number of an interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to the interrogated segment, wherein the test sequence library is enriched using one or more capture probes; (b) determining a number of sequence reads mapped to the interrogated segment; (c) determining a copy number likelihood model based on the expected number of sequence reads mapped to the interrogated segment; and (d) a hidden Markov model comprising: (i) one or more hidden states including copy numbers corresponding to the interrogated segment or a plurality of subsegments within the interrogated segment; and (ii) a copy number likelihood model based on the expected number of sequence reads mapped to the interrogated segment. (ii) constructing a hidden Markov model including: (a) an observation state including the number of sequence reads mapped to the queried segment; and (ii) a copy number likelihood model; (e) parameterizing the hidden Markov model by adjusting the copy number likelihood model to fit the determined number of sequence reads mapped to the queried segment and taking into account one or more spurious capture probes, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; and (f) determining the most probable copy number of the queried segment based on the parameterized hidden Markov model.
[0029] Further described herein is a method for determining the copy number of an interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, wherein the plurality of spatially adjacent segments includes the interrogated segment, and wherein the test sequence library is enriched using a plurality of spatially adjacent direct target sequence capture probes; (b) determining the number of sequence reads mapped to each spatially adjacent segment; (c) determining a copy number likelihood model for each spatially adjacent segment based on the expected number of mapped sequence reads in the spatially adjacent segments; and (d) using a hidden Markov model to (i) estimate the copy number of each of the spatially adjacent segments, or the number of sequences within each of the spatially adjacent segments. (ii) constructing a hidden Markov model including a plurality of hidden states including subsegments; (ii) a plurality of observed states including sequence read numbers mapped to each spatially adjacent segment; and (iii) a copy number likelihood model for each spatially adjacent segment; (e) parameterizing the hidden Markov model including adjusting each copy number likelihood model to fit the determined sequence read numbers mapped to each spatially adjacent segment and taking into account one or more spurious capture probes, wherein the hidden Markov model is parameterized using an analytical first derivative gradient and second derivative Hessian of one or more parameters in the copy number likelihood model; and (f) determining the most probable copy number of the queried segment based on the parameterized hidden Markov model.
[0030] In some embodiments of the above-described methods, one or more parameters of the copy number likelihood model are determined by the variance of the number of mapped sequence reads of a segment (d i ), the number of mapped sequence reads for the representative segment (μ i ), the variance of the number of mapped sequence reads of the segments in the test sequence library (d j), or the number of mapped sequence reads (μ j ) is included.
[0031] In some embodiments of the above-described method, the method further comprises determining the most probable copy number of a section within the region of interest, the section comprising a plurality of spatially adjacent segments including the queried segment.
[0032] In some embodiments of the methods described above, the copy number likelihood model comprises a distribution of two or more copy number states.
[0033] In some embodiments of the above-described methods, the copy number likelihood model comprises a negative binomial distribution, and the negative binomial distribution is not a Poisson distribution.
[0034] In some embodiments of the above-described method, the expected number of sequence reads is based on the number of representative mapped sequence reads at corresponding segments across multiple sequence libraries and the number of representative mapped sequence reads at corresponding segments across multiple segments of interest in the test sequence library, and the number of representative mapped sequence reads at corresponding segments across multiple sequence libraries or the number of representative mapped sequence reads at corresponding segments across multiple segments of interest in the test sequence library are normalized representative values.
[0035] In some embodiments of the above-described methods, the copy number likelihood model is adjusted to account for the presence of GC content bias. In some embodiments, the adjustment depends on the GC content of the capture probe corresponding to the interrogated segment or the GC content of the interrogated segment.
[0036] In some embodiments of the above-described method, the hidden Markov model includes transition probabilities for copy numbers of the interrogated segment given copy numbers of spatially adjacent segments. In some embodiments, the transition probabilities consider the length of a representative value of copy number variants. In some embodiments, the transition probabilities consider the prior probability of copy number variants in the interrogated segment or spatially adjacent segments. In some embodiments, the length of a representative value of copy number variants or the probability of copy number variants in the interrogated segment is determined based on observations in a human population.
[0037] In some embodiments of the above-described method, the hidden Markov model includes multiple transition probabilities for subsegment copy numbers at multiple subsegments within the interrogated segment for a given copy number of spatially adjacent subsegments. In some embodiments, the transition probabilities consider a representative length of copy number variants. In some embodiments, the transition probabilities consider a prior probability of copy number variants at the interrogated segment or spatially adjacent segments. In some embodiments, the representative length of copy number variants or the probability of copy number variants at the interrogated segment is determined based on observations in a human population.
[0038] In some embodiments of the above-described method, parameterizing the hidden Markov model includes considering one or more spurious capture probes. In some embodiments, considering the one or more spurious capture probes includes weighting one or more observation states of the plurality of observation states with a spurious capture probe indicator. In some embodiments, the spurious capture probe indicator is determined using a Bernoulli process. In some embodiments, considering one or more capture probes to be spurious includes using expectation maximization. In some embodiments, if a capture probe is determined to be spurious, likelihood information from that capture probe is discarded in the copy number likelihood model.
[0039] In some embodiments of the above-described methods, parameterizing the hidden Markov model includes accounting for noise in the number of mapped sequence reads.
[0040] In some embodiments of the above-described method, accounting for noise in the mapped sequence read counts comprises adjusting a copy number likelihood model. In some embodiments, adjusting the copy number likelihood model to account for noise comprises an expectation maximization step. In some embodiments, the expectation maximization step comprises weighting the level of noise in the mapped sequence read counts from the test sequence library. In some embodiments, if the noise in the mapped sequence read counts is above a predetermined threshold, the most probable copy number of the queried segment is not called.
[0041] In some embodiments of the above-described methods, sequence reads from overlapping capture probes are merged.
[0042] Some embodiments of the above-described methods use a Viterbi algorithm, a quasi-Newton solver, or Markov Chain Monte Carlo to determine the most probable copy number of the queried segment.
[0043] Some embodiments of the above-described methods further comprise determining a confidence level for the most probable copy number of the segment.
[0044] In some embodiments of the above-described methods, one or more parameters of the copy number likelihood model are the variance of the number of mapped sequence reads of a segment (d i ), the number of mapped sequence reads for the representative segment (μ i ), the variance of the number of mapped sequence reads of the segments in the test sequence library (d j ), or the number of mapped sequence reads (μ j ) is included.
[0045] In some embodiments of the above-described methods, the analytical first derivative gradient and second derivative analytical Hessian of one or more parameters in the copy number likelihood model are solved using a trust-region Newton conjugate gradient algorithm.
[0046] Also described herein is a computer system that includes a computer-readable medium containing instructions for performing any one of the above-described methods. [Brief explanation of the drawings]
[0047] [Figure 1] 1 shows a flowchart of one embodiment of a method for determining the copy number of a segment. [Figure 2A] Median sequence read counts (i.e., sequencing depth) across approximately 2500 segments (approximately 2500 unique capture probes) in 48 test sequence libraries are shown. [Figure 2B] The median normalized sequencing depth (i.e., the sequencing depth of a single segment normalized to the median of the same segment across all test sequence libraries) of the 48 different test sequence libraries shown in Figure 2A is shown. [Figure 3A] This figure shows a plot of the average number of sequence reads ("mean depth") of approximately 2500 capture probes used to enrich sequence libraries of regions of interest for multiple different samples versus sequence depth variance. The data was fitted using a negative binomial distribution, which is not a Poisson distribution. For comparison, a Poisson distribution, which assumes a linear relationship between variance and mean depth, is also illustrated. As can be seen in the graph, the variance for the depth distribution across the probes follows a negative binomial distribution, not a simple Poisson distribution. [Figure 3B]The figure shows a copy number likelihood model that includes a negative binomial distribution, which is either a non-Poisson distribution or a Poisson distribution where the copy number of the segment is 1, 2, or 3 copies. The distribution is the probability mass function (pmf) as a function of the number of sequence reads from the capture probe corresponding to the segment. "CN" = copy number. [Figure 4A] In the exemplary hidden Markov model, c1, c2, c3, and c4 represent hidden states (i.e., the most probable copy numbers of four different segments), and k1, k2, k3, and k4 represent observed states (i.e., the number of mapped sequence reads for each corresponding segment). The probabilities between the observed and hidden states for each corresponding segment are denoted by p(c1|k1), p(c2|k2), p(c2|k2), and p(c2|k2)), and the transition probabilities between hidden states are denoted by p(c2|c1), p(c3|c2), and p(c4|c3). A copy number likelihood model is used to parameterize the probabilities between the observed and hidden states. Both sets of probabilities are optimized using expectation maximization (EM). [Figure 4B] Illustrates a hidden Markov model of two segments subdivided into subsegments. A subsegment contains hidden states but not observed states. Transition probabilities for a subsegment are based on the copy number states of adjacent subsegments. This can be performed with base-by-base (or subsegment-by-subsegment) segmentation. [Figure 5A] 1 illustrates a hidden Markov model in which spurious capture probe indicators are placed in the observed states. [Figure 5B] Illustrate prior distributions that can be adjusted to determine whether a given capture probe is a spurious capture probe that is used to determine the prior b for the observed state k. A Bernoulli process can be used to determine the spurious capture probe probability for each test sequence library and how this probability may affect the spuriousness of probes for that test sequence library. [Figure 6A]For the less noisy test sequence library, the number of sequence reads normalized by the consumed sequencing depth for multiple segments across 22 genes is shown. [Figure 6B] Figure 1 shows the determined sequence read counts normalized by the consumed sequencing depth for multiple segments across 22 genes for the noisier test sequence library. Two different test sequence libraries enriched with the same capture probe display different levels of noise. [Figure 7] Copy number calls across multiple segments within the same region of interest (y-axis) for several test sequence libraries (x-axis) relying solely on the copy number likelihood model are shown. Darker areas indicate deviations from a copy number state of 2. Boxed areas indicate true copy number variants spanning multiple segments, whereas deviations from a copy number state of 2 observed only within a segment are likely false positives rather than true copy number variants. [Figure 8] Copy number calls across multiple segments within the same region of interest (y-axis) for several test sequence libraries (x-axis) are shown after using a hidden Markov model to determine the most probable copy number. Darker areas indicate deviations from a copy number state of 2. Boxed areas show how true copy number variants span multiple segments, minimizing false positives. The HMM takes into account the influence of the copy number state of adjacent segments on subsequent segments. This allows the model to call true copy number variants as opposed to variations observed within a single segment. [Figure 9]
[0014] A schematic diagram is provided for evaluating a copy number variant model by parameterizing the copy number variant model using actual sequence read numbers from a test sample, generating synthetic copy number variants based on the actual sequence read numbers, and calling the copy numbers of segments within the synthetic copy number variants based on the parameterized copy number variant model. [Figure 10]1 illustrates binomial sampling of actual sequence reads from a test sample with two copies of a segment to generate a synthetic copy number variant with one copy of the segment. [Figure 11] 1 depicts an exemplary computing system configured to perform any one of the processes described above, including various exemplary methods for calling copy numbers of queried segments or evaluating the performance of copy number variant models. [Figure 12] Figure 1 shows the sensitivity results of two hidden Markov model copy number variant collaborators plotted against increasing percentage of saliva samples. Saliva samples generally have noisy sequencing depth. The reference collaborator does not account for sequence library noise or spurious capture probes, while the test collaborator takes both of these factors into account. DETAILED DESCRIPTION OF THE INVENTION
[0048] The methods described herein allow for accurate determination of the copy number of an interrogated segment of the genome, such as a gene or gene segment. In some embodiments, the quality of the copy number variant calls is controlled by generating quality control metrics (such as sensitivity) for each sample. Accurate copy number calls allow for improved diagnosis of specific genetic abnormalities and aid in making important medical decisions.
[0049] Copy number variant callers can be used to screen test sequence libraries for copy number variants at one or more segments within a region of interest. These callers operate by constructing a copy number variant model, such as a hidden Markov model (HMM), which is parameterized to generate one or more copy number variant (CNV) model parameters for the test sample. The CNV model parameters may vary depending on sequencing depth, sample noise, capture probe efficiency, and / or other artifacts that arise during sequencing of the test sequence library.
[0050] Synthetic copy number variants can be generated to evaluate the performance of a copy number variant collaborator (or a copy number variant model used by the collaborator). The collaborator can be used to call copy numbers at one or more segments within the synthetic copy number variants, and performance statistics can be determined to provide an evaluation of the collaborator. Parameterizing a copy number variant model is computationally intensive. Therefore, sample-specific performance evaluation by parameterizing a CNV model for each synthetic copy number variant is impractical, especially when using a collaborator to screen a large number of samples. However, as described herein, the copy number variant model can be parameterized using sequence reads from a test sample to determine sample-specific CNV model parameters. Because the synthetic copy number variants can be generated based on sequence reads from the test sample, the CNV model parameters are specific to the test sample, and the synthetic copy number variants are generated based on the test sample, the determined sample-specific CNV model parameters can be used by the collaborator to call the copy numbers of segments of the synthetic copy number variants without reparameterizing the model. Thus, the methods described herein conserve substantial computing power while generating reliable performance statistics for the evaluation of CNV models.
[0051] Copy number variant models, such as hidden Markov models (HMMs), use analytical first-derivative gradients and second-derivative Hessians of one or more parameters of the copy number likelihood model. In some embodiments, the first derivative gradient and second derivative Hessian are solved using a trust-region Newton conjugate gradient algorithm. An expectation-maximization (EM) step can be used to determine the copy number variant model parameters, which can include multiple optimization loops. In some embodiments, EM parameterizes the CNV model to maximize the log-likelihood weighted by the expected copy number call.
[0052] Certain methods involve using a hidden Markov model (HMM) to determine the most probable copy number of a queried segment in a test sequence library. In some embodiments, the test sequence library is enriched using a direct target sequencing (DTS) method. DTS methods provide high-resolution targeting of the queried sequence, and the HMM callers described herein substantially benefit from the large amount of data collected for copy number calling. To further improve the accuracy of the HMM callers, sequencing depth artifacts that may arise from direct target sequencing methods can be taken into account. Such sequencing depth artifacts may include, for example, GC bias correction and spurious probe determination. In addition, the methods described herein provide accurate copy number calling when sequence reads are generated from a noisy sequence library.
[0053] A sequence library derived from a patient sample can be sequenced to obtain several sequence reads. The copy number of a segment is related to the sequence depth (i.e., the number of sequence reads or the normalized number of sequence reads) at that segment. The present disclosure describes a method for determining the existence of a copy number state at a segment using the sequence depth at that segment. The sequence depth can be obtained by determining the sequence reads mapped to the segment. The sequence depth can be obtained by determining the sequence reads mapped to the capture probe corresponding to the segment. This method takes into account several factors associated with the sequence technology to optimize the call for more accuracy.
[0054] Determining the number of mapped sequence reads for a segment depends, at least in part, on the segment's actual copy number state. While the majority of mammalian gene regions are diploid, and thus generally expected to have two copies of a gene segment, this may not always be the case. For example, some regions of the genome are not diploid due to their location (e.g., located on the Y chromosome). Other regions of the genome lose their diploid status as a result of functional specialization of some cells, such as immune cells, resulting in genome rearrangements. However, regardless of deviations from these standards, the copy number state of most genomic regions is expected to be 2, and deviations from a copy number state of 2 are expected to be reflected in the number of mapped sequence reads.
[0055] Before mapping sequence reads to segments, one or more upstream steps can be performed, such as sample preparation, including fragmentation, forming a sequence library (e.g., by ligating sequence adapters to nucleic acid molecules in the sequence library), and sequencing the sequence library. Noise in the sequence depth during any of these upstream steps can introduce noise into the sequence read counts. Furthermore, various capture probes within a capture probe library may not perform identically. For example, a particular segment within a region of interest may not allow for ideal capture probe design, which can lead to spurious capture probes. Therefore, using a determined number of mapped sequence reads to determine the copy number state of a segment is less straightforward than simply recognizing the existing dependency between the copy number state of the segment and the determined number of mapped sequence reads at the segment. The method of the present invention uses a hidden Markov model that is parameterized and optimized to account for the dependency between the number of mapped sequence reads and the copy number state of the segment to identify queried segments within the region of interest. This method allows for copy number calls of interrogated segments or subsegments within a region of interest to be made. Hidden Markov models can also account for various sources and levels of confounding factors. This method allows for a particularly effective and efficient process for determining the copy number of interrogated segments or subsegments within a region of interest and for determining copy number variant abnormalities within the region of interest.
[0056] In some embodiments of the present invention, a sequence library is enriched for a region of interest using direct target sequences. The direct target sequences use a capture probe library containing multiple capture probes that hybridize to nucleic acid molecules in the sequence library. The capture probes are designed to hybridize to segments within the region of interest, and each capture probe has a corresponding segment. Thus, the region of interest is determined by the capture probes used to enrich the sequence library. The capture probes are extended using the nucleic acid molecule hybridized to the capture probe as a template. The extended capture probes can then be sequenced to obtain the sequence of a portion of the nucleic acid molecule (i.e., the portion corresponding to the segment from the region of interest). Because the sequence of the capture probe itself is determined, the segment corresponding to the capture probe begins following the end of the capture probe. In some embodiments, the extended capture probes are amplified to obtain additional copies. Amplification of the extended capture probes can also introduce artifacts into the sequencing depth, which can be normalized as described herein.U.S. Patent No. 9,309,556, entitled "Direct Capture, Amplification, and Sequencing of Target DNA using Immobilized Primers," U.S. Patent No. 9,092,401, entitled "System and Method for Detecting Genetic Variation," U.S. Patent Application No. 2014 / 0024541, entitled "Methods and Compositions for High-Throughput Screening," Myllykangas et al., "Efficient targeted resequencing of human germline and cancer genomes by oligonucleotide-selective sequencing." Nat Biotechnol. 29(11):1024-7 (2011), and Hopmans et al., "A. Programmable method for massively parallel targeted sequencing." Nucleic Acids Res. 42(10):e88 (2014), describe embodiments of direct target sequencing. Direct target sequencing need not be performed using surface-based methods, but can also be performed in solution.
[0057] In some embodiments, the sequence library is enriched for the region of interest using a method other than direct target sequencing. For example, the sequence library can be enriched using hybrid capture technology, which involves combining the sequence library with a capture probe library and hybridizing the capture probe with the nucleic acid molecules in the sequence library. The hybridized nucleic acid molecules can then be isolated from the rest of the sequence library (e.g., by using biotinylated capture probes and streptavidin beads to separate the hybridized molecules). The nucleic acid molecules in the enriched sequence library can then be sequenced. Because nucleic acid molecules from the sequence library are directly sequenced (as opposed to direct target sequencing), the capture probe does not necessarily correspond to a specific segment within the region of interest. Instead, the sequencing depth at any given base within the region of interest can be determined by the number of sequence reads at that base.
[0058] Provided herein are definitions, explanations, examples and illustrations that will enable those skilled in the art to understand the scope of the methods provided and to practice the invention. It is understood that one, some, or all of the features of the various described embodiments may be combined to form other embodiments of the present invention. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0059] definition As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.
[0060] Reference herein to "about" or "approximately" a value or parameter includes (and describes) a variation that is directed to the value or parameter itself. For example, a description that refers to "about X" includes a description of "X."
[0061] The term "typical value" as used herein refers to the mean or median value, or any value used to estimate the mean or median value, unless the context clearly dictates otherwise.
[0062] "Capture probe" refers to a DNA or RNA molecule that hybridizes to a nucleic acid molecule present in a sequence library that has a complementary sequence, or a segment with a sequence sufficiently complementary to allow hybridization under normal hybridization conditions.
[0063] "Copy number likelihood" refers to the likelihood of the copy number in a segment or subsegment of interest.
[0064] "Copy number likelihood model" refers to a statistical model used to determine the copy number likelihood given the number of mapped sequence reads at that segment. The copy number likelihood model includes a statistical distribution for each copy number state covered by the model, with each distribution reflecting the probability that the copy number state is correct for a given number of mapped sequence reads.
[0065] "Copy number variant" or "CNV" refers to a deviation of copy number state from wild type. As used herein, "wild type" refers to a predetermined copy number state of a particular segment that is considered normal. Determining what is "wild type" can be based on population data from humans, mammals, or other animals. Determining what is "wild type" can also be based on reference runs, internal experiments, and data generated from such experiments.
[0066] A "direct target sequence capture probe" is a capture probe used to enrich sequences from a sequence library using a direct target sequence.
[0067] "Interrogated segment" refers to a segment within a region of interest for which a copy number variant model is used to determine copy number state. The interrogated segment can be divided into subsegments that can be as small as one base pair but no longer than the length of the interrogated segment.
[0068] A "noisy sequence library" or "noise" from a sequence library refers to a sequence library that produces poor data across one or more capture probes.
[0069] "Number of sequence reads" as used herein refers to the absolute number of sequence reads or the normalized number of sequence reads.
[0070] "Actual sample" refers to a nucleic acid sequence or sequence read originating from a physical sample subjected to genetic sequencing without alteration of the sequence, sequence read, or number of sequence reads. "Actual reference sample" refers to an actual sample that is compared to a synthetic sample (e.g., a synthetic copy number variant) by a genetic variant caller.
[0071] "Actual sequence reads" refer to sequence reads originating from actual samples without sequence alterations. "Actual sequence read count" refers to the absolute number of actual sequence reads or the normalized number of sequence reads, but does not refer to sequence read counts altered to reflect increased copy number of any segment or region of interest.
[0072] A "segment" refers to a chain of nucleotides containing two or more bases. A segment can be subdivided into one or more "subsegments." A "subsegment" can be as small as one nucleotide, but not longer than the segment in which it is located. A region of interest can be divided into one or more segments. Segments can be, but need not be, contiguous. Thus, a region of interest can optionally contain non-contiguous subregions. Segments can be the same length or different lengths. Two or more segments within a region of interest can be grouped to create a section within the region of interest. The segments that make up a section within a region of interest can be, but need not be, contiguous.
[0073] "Spurious capture probe" refers to a capture probe that generates artifacts in sequence read counts that are not related to copy number. Artifacts can result from substandard sequence reads, inconsistent sequence reads, sequence reads of lengths below a predetermined level, sequence read counts below a predetermined level, or poor display quality compared to other capture probes.
[0074] "Spatially adjacent segments" refers to a set of consecutive segments that are located within the same chromosome but do not need to be contiguous. That is, two spatially adjacent segments may be separated by some intervening nucleotides, but cannot be separated by intervening segments outside the set of spatially adjacent segments. When two spatially adjacent segments are not contiguous, the copy number of the intervening nucleotide can be estimated by a hidden Markov model. "Spatially adjacent capture probes," including "spatially adjacent direct target sequence capture probes," refer to capture probes that correspond to spatially adjacent segments.
[0075] The term "synthetic copy number variant" refers to an artificial sample generated using actual sequence reads, or actual sequence read numbers from an actual sample, with an increase or decrease in copy number of one or more segments within a region of interest relative to the actual sample.
[0076] "Synthetic copy number" refers to the copy number of a segment within a region of interest of a synthetic copy number variant, and the copy number may be increased, decreased, or the same as that of an actual sample. Since a synthetic copy number variant does not necessarily have to change the copy number of each segment and may include the wild-type copy number of one or more segments, the synthetic copy number of one or more segments of a synthetic copy number variant may be the same as the actual copy number of the segment.
[0077] "Composite sequence read number" refers to the sequence read number used to represent the composite copy number of a segment within a region of interest. The composite sequence read number of a segment can be increased, decreased, or maintained compared to the actual sequence read number of the corresponding segment.
[0078] It will be understood that embodiments and variations of the present invention described herein include those "consisting of" and / or "consisting essentially of".
[0079] When a range of values is provided, it is understood that each intervening value between the upper and lower limits of that range, and any other stated or intervening value in that stated range, is encompassed within the scope of the disclosure. When a stated range includes an upper or lower limit, ranges excluding either of those included limits are also included in the disclosure.
[0080] Methods for determining copy number The present disclosure provides a method for determining the copy number of an interrogated segment of a region of interest (or a subsegment of the interrogated segment) or copy number variant abnormalities within the region of interest based on the determined number of mapped sequence reads for the segment. The method includes determining a copy number likelihood model based on the expected number of mapped sequence reads for one or more copy number states. The first- and second-derivative Hessians of one or more parameters of the copy number likelihood model can be used in conjunction with expectation maximization (EM) to enable latent parameter estimation and optimization of the model. The first- and second-derivative Hessians can be solved, for example, using a trust-region Newton conjugate gradient algorithm. Additional steps and adjustments to the model can be used to account for other factors that affect the relationship between copy number and number of mapped sequence reads. This information can be used to parameterize a hidden Markov model, which can then be used to determine the most probable copy number state in the interrogated segment. Methods for constructing copy number likelihood models, expectation maximization implementations, tuning models to account for multiple factors, parameterization of hidden Markov models, and methods for solving various steps and the entire model are generally provided below.
[0081] Briefly, a method for determining the copy number of a segment or subsegment can include: (1) determining the number of sequence reads mapped to the interrogated segment; (2) constructing and parameterizing a hidden Markov model by determining a copy number likelihood model; and (3) using the parameterized hidden Markov model to determine the most probable copy number of the interrogated segment (or a subsegment of the interrogated segment). The hidden Markov model is parameterized using the first derivative gradient and second derivative Hessian of one or more parameters of the copy number likelihood model, along with expectation maximization (EM), which can be solved using a trust-region Newton conjugate gradient algorithm. In some embodiments, the methods provided herein also include a step for improving the model by considering confounding effects that may occur during the process.
[0082] In some embodiments of the methods described herein, a hidden Markov model is used to determine the most probable copy number state of a segment. The hidden Markov model can include a hidden layer containing the copy number state of the segment of interest, an observation layer containing the number of mapped sequence reads, transition probabilities between the copy number state of the hidden layer and the number of mapped sequence reads (probability inter-layer), and transition probabilities of the copy number state of the segment given the copy number states of the preceding adjacent segments (probability intra-hidden layer). Figure 1 illustrates one embodiment of a method for determining the copy number of a queried segment within a region of interest. In step 110, sequence reads generated for a test sequence library are mapped to one or more segments within one or more regions of interest. In step 120, the number of mapped sequence reads at the segment(s) within the region(s) of interest is determined. In step 130, a copy number likelihood model is determined, which is used to set the transition probabilities of the copy number state given the number of observed mapped sequence reads. In step 140, the hidden layer, the observation layer, and the transition probabilities are calculated. A Markov model is constructed. In step 150, a hidden Markov model is parameterized using the first derivative gradient and second derivative Hessian of one or more parameters of the copy number likelihood model, which can be solved using a trust-region Newton conjugate gradient algorithm. In its simplest form, the hidden Markov model includes at least two unknown parameters: the copy number state and the transition probability between the copy number state and the observed sequence read number, as determined by the copy number likelihood model. The first derivative gradient and second derivative Hessian of one or more parameters of the copy number likelihood model are used in conjunction with expectation maximization to determine these parameters (i.e., parameterize the model) based on a best fit of the data to determine the most probable copy number. The model desirably maximizes the probability of the copy number state given the observed sequence read number to determine the most probable copy number of the segment. In step 160, the most probable copy number state of the segment is determined. This process may take into account other variables that affect the observed state, such as GC content bias, spuriousness of capture probes associated with the segment, or a noisy test sequence library that affects transition probabilities. Additional variables are treated as potentials and determined by EM given the available data. The transition probabilities are then adjusted to account for these other variables. The EM process can be cumulative (adjusting all variables at once) or can accommodate variables in separate EM iterations before solving the HMM to determine the most probable copy number state.
[0083] Determining the number of mapped sequence reads In some embodiments, the methods described herein include mapping a plurality of sequence reads generated from a test sequence library to one or more segments, such as an interrogated segment. In some embodiments, the methods described herein may include mapping a plurality of sequence reads generated from the test sequence library to a plurality of segments (which may be spatially adjacent), where the plurality of segments includes the interrogated segment. The sequence library is enriched for a region of interest, such as by direct target sequences. The mapped sequence reads can be counted to determine the number of sequence reads that map to the interrogated segment or spatially adjacent segments.
[0084] In some embodiments, the segments are located within the same chromosome. In some embodiments, the segments are located within the same chromosomal region. In some embodiments, the segments are located within the same gene. In some embodiments, the segments are located within the same region of interest. In some embodiments, the segments are located within the same portion of the region of interest.
[0085] The sequence library can be sequenced to generate multiple sequence reads that can be mapped to the region of interest. The sequence library contains multiple nucleic acid fragments that can be isolated from bodily fluids such as blood, plasma, saliva, and urine, or from tissues or cultured cells. The nucleic acid fragments can be from an animal. The nucleic acid fragments can be from a mammal, such as a human. In a preferred embodiment, the test sequence library contains multiple nucleic acid fragments isolated from a patient. The nucleic acid molecules in the sequence library can be ligated to sequence adaptors that can assist in alignment in a specific sequence method. For example, the adaptors can be indexed, and indexing can be used to assist in sequence alignment. The sequence library can be enriched for the region of interest (such as by direct target sequences) either before or after ligating the nucleic acid molecules to the sequence adaptors.
[0086] The nucleic acid fragments in the test sequence library can be RNA or DNA nucleic acid fragments. The nucleic acid fragments can be cell-free DNA. In some embodiments, the cell-free DNA includes fetal cell-free DNA. In some embodiments, the cell-free DNA includes circulating tumor cell-free DNA. include.
[0087] The nucleic acid fragments in the sequence library comprise a region of interest. The region of interest can be the entire genome or any portion of a genome. In some embodiments, the region of interest comprises one or more chromosomes. In some embodiments, the region of interest comprises one or more genes of interest (e.g., two or more, three or more, four or more, five or more, about 10 or more, about 15 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 75 or more, about 100 or more, about 150 or more, about 200 or more, about 250 or more genes, about 300 or more, about 350 or more, about 400 or more, about 450 or more, about 500 or more, about 550 or more, about 600 or more, about 650 or more, about 700 or more, about 750 or more, about 800 or more, about 850 or more, about 900 or more, about 950 or more, or about 1000 or more, etc.). The one or more genes of interest can be any genes associated with a disease. The one or more genes of interest can include any genes associated with a genetic disease. The one or more genes of interest may include genes associated with a form of cancer, such as hereditary cancer. In some embodiments, the region of interest is one or more exons (e.g., 2 or more, 3 or more, 4 or more, 5 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, 50 or more, 75 or more, 100 or more, 150 or more, 200 or more, 250 or more, 500 or more, 1000 or more, or 2000 or more exons). In some embodiments, the region of interest is selected from the group consisting of APC, ATM, BARD1, BMPR1A, BRCA1, BRCA2, BRIPI, CDH1, CDK4, CDKN2A, CHEK2, EPCAM, GREM1, MEN1, MLH1, MRE11A, MSH2, MSH6, MUTYH, NBN, PALB2, PMS2, POLD1, POLE, PTEN, RAD50, RAD51C, RAD51D, RET, SDHA, SDHB, SDH C, SMAD4, STK11, TP53, VHL, PEX10, MTHFR, ALPL, HMGCL, DHDDS, PPT1, MPL, MMACHC, POMGNT1, CPT2, ALG6, RPE65, ACAD M, DPYD, AGL, SLC35A3, DBT, PHGDH, CTSK, NTRK1, NPHS2, LAMC2, LAMB3, USH2A, PHYH, ERCC6, PCDH15, LIPA, HOGA1, OAT,<h2 style=";text-align:left;direction:ltr">TH, HBB, SMPD1, TPP1, KCNJ11, ABCC8, USH1C, RAG2, RAPSN, TMEM216, PYGM, BBS1, PC, TCIRG1, CPT1A, DHCR7, MYO7A, MED17, PTS, SLC37A4, HY LS1、PFKM、BBS10、GNPTAB、PAH、MMAB、ACADS、PUS1、GJB2、GJB6、SGCG、 SACS、ATP7B、CLN5、PCCA、TGM1、ZFYVE26、VSX2、NPC2、GALC、SERPINA1、 VRK1、TECPR2、SLC12A6、IVD、CAPN3、CLN6、NR2E3、HEXA、MPI、FAH、MES P2、BLM、GNPTG、MEFV、PMM2、CLN3、BBS2、TAT、CYBA、FANCA、VPS53、ASPA 、CTNS、ACADVL、ALDH3A2、PEX12、NAGLU、G6PC、SGCA、MKS1、DNAI2、GAL K1、GAA、SGSH、NPC1、LAMA3、LOXHD1、MCOLN1、MAN2B1、GCDH、NPHS1、BCK DHA, OPA3, FKRP, HADHA, LRPPRC, FAM161A, ATP6V1B1, DYSF, ALMS1, NEB, CERKL, CPS1, BCS1L, CYP27A1, COL4A4, COL4A3, AGXT, NDUFAF5, ADA 、RTEL1、HLCS、CBS、AIRE、TRMU、MLC1、TYMP、ARSA、SUMF1、XPC、BTD、GL B1、AMT、GBE1、HGD、PCCB、HPS3、CLRN1、BCHE、IDUA、EVC2、EVC、SEPSECS 、SGCB、MTTP、BBS12、MMAA、AGA、F11、NDUFS6、DNAH5、NDUFS4、ERCC8、H EXB、HSD17B4、SLC22A5、SLC26A2、SGCD、PROP1、ADAMTS2、PEX6、MUT、PK HD1, EYS, SLC17A5, BCKDHB, RARS2, LAMA2, ARG1, PEX7, ASL, PEX1, SAMD9, ASNS, SLC26A4, DLD, CFTR, CLN8, STAR, HGSNAT, TTPA, PEX2, CNGB3,<h2 style=";text-align:left;direction:ltr"> The gene or portion of a gene, exon or portion of an exon is selected from the group consisting of VPS13B, CYP11B1, CYP11B2, GLDC, DNAI1, GALT, RMRP, GNE, GRHPR, VPS13A, FANCC, XPA, ALDOB, FKTN, IKBKAP, ASS1, RS1, NR0B1, DMD, OTC, IL2RG, ATP7A, CHM, GLA, COL4A5, IDS, MTM1, ABCD1, or a combination thereof.
[0088] A region of interest can be divided into multiple segments. Each segment can be further divided into subsegments. A subsegment can be one or more nucleotides in length. Segments within a region of interest can be, but need not be, contiguous. For example, in some embodiments, a region of interest can include one or more non-contiguous segments, two or more non-contiguous segments, three or more non-contiguous segments, four or more non-contiguous segments, five or more non-contiguous segments, ten or more non-contiguous segments, twenty-five or more non-contiguous segments, fifty or more non-contiguous segments, one hundred or more non-contiguous segments, one hundred or more non-contiguous segments, one hundred or more non-contiguous segments, two hundred or more non-contiguous segments, two hundred or more non-contiguous segments, two hundred or more non-contiguous segments, three ... The non-contiguous segments may include 400 or more non-contiguous segments, 450 or more non-contiguous segments, 500 or more non-contiguous segments, 550 or more non-contiguous segments, 600 or more non-contiguous segments, 650 or more non-contiguous segments, 700 or more non-contiguous segments, 750 or more non-contiguous segments, 800 or more non-contiguous segments, 850 or more non-contiguous segments, 900 or more non-contiguous segments, 950 or more non-contiguous segments, or 1000 non-contiguous segments. In some embodiments, each of the non-contiguous segments includes one or more contiguous bases, two or more contiguous bases, three or more contiguous bases, four or more contiguous bases, or five or more contiguous bases. For example, in some embodiments, each of the non-contiguous segments includes 1 to about 20 contiguous bases (e.g., 1 to about 10 contiguous bases, or about 1 to about 5 contiguous bases).In some embodiments, the region of interest comprises 1 or more contiguous segments, 2 or more contiguous segments, 3 or more contiguous segments, 4 or more contiguous segments, 5 or more contiguous segments, 10 or more contiguous segments, 25 or more contiguous segments, 50 or more contiguous segments, 100 or more contiguous segments, 150 or more contiguous segments, 200 or more contiguous segments, 250 or more contiguous segments, 300 or more contiguous segments, 350 or more contiguous segments, The non-contiguous segments may comprise 400 or more contiguous segments, 450 or more contiguous segments, 500 or more contiguous segments, 550 or more contiguous segments, 600 or more contiguous segments, 650 or more contiguous segments, 700 or more contiguous segments, 750 or more contiguous segments, 800 or more contiguous segments, 850 or more contiguous segments, 900 or more contiguous segments, 950 or more contiguous segments, or 1000 contiguous segments. In some embodiments, each of the non-contiguous segments comprises one or more contiguous bases, two or more contiguous bases, three or more contiguous bases, four or more contiguous bases, or five or more contiguous bases. For example, in some embodiments, each of the non-contiguous segments comprises 1 to about 20 contiguous bases (e.g., 1 to about 10 contiguous bases, or about 1 to about 5 contiguous bases). In some embodiments, the region of interest includes a combination of non-adjacent and adjacent segments. In some embodiments, the region of interest includes only one segment. In some embodiments, the region of interest includes at least one segment. In some embodiments, the region of interest includes at least two segments. In some embodiments, the region of interest includes at least two adjacent segments. In some embodiments, one segment in a first region of interest may be adjacent to a segment in a second region of interest that is adjacent to the first region of interest.
[0089] A region of interest can be enriched with one or more capture probes. The reference locations of the capture probes relative to the region of interest are known. For example, the capture probes contain reference sequences corresponding to predetermined probe coordinates. In some embodiments, the region of interest is divided into segments based on the locations of the capture probes (i.e., the capture probes correspond to the segments). The capture probes contain reference sequences corresponding to the probe coordinates. For example, the first nucleotide of a segment can match the first nucleotide of the sequence that hybridizes to the 3' end of the capture probe. In some embodiments, the first nucleotide of a segment matches the first nucleotide of the sequence that hybridizes to the 5' end of the capture probe. In some embodiments, the region of interest includes two spatially adjacent segments. A segment within a region of interest can be divided into subsegments. The subsegments can be as small as one nucleotide in length as the segment. The subsegments can overlap. For example, a first subsegment can be the first nucleotide of the segment plus one downstream nucleotide. A second subsegment can include the first subsegment plus an additional downstream nucleotide. In some embodiments, a segment that is n nucleotides long comprises n-1 subsegments, where each subsequent subsegment is one nucleotide longer than the previous one. In some embodiments, a segment that is n nucleotides long comprises n subsegments, where each subsegment is one nucleotide in length.
[0090] The region of interest includes at least one interrogated segment. The interrogated segment is a segment for which it is desirable to know the copy number. The copy number state of the interrogated segment is unknown, and the most probable copy number of the interrogated segment is determined by solving a hidden Markov model. Like other segments, the interrogated segment can be divided into subsegments. In some embodiments, the first nucleotide of the interrogated segment matches the first nucleotide of the sequence that hybridizes to the 5' end of a capture probe. In some embodiments, the first nucleotide of the interrogated segment matches the first nucleotide of the sequence that hybridizes to the 3' end of a capture probe. In some embodiments, the interrogated segment comprises a sequence that spans two spatially adjacent capture probes. In a preferred embodiment, the interrogated segment comprises a nucleotide sequence between two adjacent capture probes, where the first nucleotide of the sequence is the first nucleotide that hybridizes to the 5' or 3' end of a capture probe, and the last nucleotide of the segment is adjacent to the first nucleotide that hybridizes to the 5' or 3' end of the spatially adjacent probe.
[0091] The test sequence library can be sequenced using next-generation sequencing to generate sequence reads.Next-generation sequencing technology is well known in the art.The test sequence library can be sequenced using high-throughput sequencers such as Illumina HiSeq2500, Illumina HiSeq3000, Illumina HiSeq4000, Illumina HiSeqX, Roche 454, PacBio Sequel System PacBio RS II, Life Technologies Ion Proton sequence system, etc.Other sequencing methods are known in the art.
[0092] In some embodiments, the sequence library is enriched with one or more capture probes by direct target sequences. In direct target sequences, the capture probe hybridizes a specific target region of a nucleic acid molecule from within the sequence library. This method allows for enrichment of the target region, allowing subsequent sequencing work to focus on the genomic region or transcript of interest. By enriching the target region with a capture probe for the region of interest, the target region of interest can be enriched. This allows for more efficient, high-throughput sequencing of regions. This efficiency preserves the overall cost of the sequencing library while maintaining or improving the sensitivity and specificity of diagnostic tests or screens. Capture probes can be selected based on the region of interest so that those nucleic acid molecules in the sequence library that contain a portion of the region of interest can hybridize to the capture probe and be enriched, while those nucleic acid molecules in the sequence library that do not contain a portion of the region of interest do not hybridize to the capture probe and are not enriched.
[0093] In direct target sequencing, a capture probe that hybridizes to a target sequence adjacent to a corresponding segment in a region of interest is combined with a sequence library, thereby hybridizing the capture probe to a nucleic acid molecule, including hybridizing the capture probe to a target sequence.In direct target sequencing, the capture probe is extended using a nucleic acid molecule as a template, and the extended capture probe is sequenced.Since the extended capture probe (or an amplified copy of the extended capture probe) is itself sequenced, the sequence of the capture probe can be used to assist sequence alignment, but is not interpreted as a sequence derived from a test sequence library.
[0094] Other methods for enriching sequence libraries using capture probes are commonly known in the art and can include hybrid capture techniques (e.g., using biotinylated capture probes) and PCR amplification using capture probes as PCR primers.
[0095] In some embodiments, hybrid capture technology is used to enrich regions of interest by combining a capture probe that is substantially complementary to a portion of the region of interest with a sequence library, thereby allowing the capture probe to hybridize to nucleic acid molecules that contain this portion of the region of interest. Nucleic acid molecules that hybridize to the capture probe can be isolated from unhybridized nucleic acid molecules (e.g., by pull-down). The hybridized complex can be denatured, and the enriched nucleic acid molecules from the sequence library can be sequenced. In some embodiments, the enriched nucleic acid molecules are re-enriched by a second (or more) round of hybridization to the capture probe, isolation, and denaturation before being sequenced. Optionally, the nucleic acid molecules in the sequence library can be amplified (e.g., by PCR) either before or after enrichment.
[0096] In some embodiments, one or more of the capture probes are attached to an additional oligonucleotide (such as a primer binding site or other specialized nucleic acid segment). In some embodiments, the capture probes in the capture probe library are DNA oligonucleotides, RNA oligonucleotides, or a mixture of DNA and RNA oligonucleotides. In some embodiments, the capture probes are about 10-100 bases in length. In some embodiments, the capture probes are about 20-60 bases in length. In some embodiments, the capture probes are about 30-50 bases in length. In some embodiments, the capture probes are 40 bases in length.
[0097] Generally, the larger the region of interest, the more capture probes are required for adequate coverage, so the number of capture probes in a capture probe library can depend on the size of the region of interest. In some embodiments, the capture probe library contains about 10 or more unique capture probes (such as about 50 or more, about 100 or more, about 250 or more, about 500 or more, about 1000 or more, about 2500 or more, about 5000 or more, about 10,000 or more, about 25,000 or more, about 50,000 or more, about 100,000 or more, or about 200,000 or more).
[0098] Sequencing enriched sequence library generates multiple sequence reads.To determine the sequencing depth of segment or sub-segment, determine the number of sequence reads that are mapped to this segment.Sequence reads can be mapped by, for example, aligning sequence reads (or a portion of sequence reads) with reference sequence, or by assigning sequence reads to a segment based on a portion of sequence reads.
[0099] In some embodiments, sequence reads are mapped by aligning the sequence reads (or a portion of the sequence reads) to a reference sequence. For example, a sequence read resulting from a direct target sequence can include a capture probe portion (i.e., a portion of the sequence read that can be attributed to the capture probe itself) and a segment portion (i.e., a portion of the sequence read that can be attributed to a segment targeted by and associated with the capture probe). In some embodiments, the segment portion is aligned with the reference sequence, the capture probe portion is aligned with the reference sequence, or the capture probe portion and the segment portion are aligned with the reference sequence. The reference sequence includes a region of interest that has been pre-divided into segments. Thus, a sequence read aligned to a reference sequence can be aligned to a corresponding segment, and the aligned sequence read is assigned or "mapped" to that segment.
[0100] In some embodiments, sequence reads are mapped by assigning the sequence read to a segment based on a portion of the sequence read. In such embodiments, it is not necessary to align the sequence read to a reference sequence. Because each capture probe corresponds to a segment, and the corresponding segment is known by the capture probe design, sequence reads that contain the sequence of a capture probe (or the complement of the capture probe) can be assigned (or "mapped") to the corresponding segment.
[0101] In some embodiments, the sequence depth may be obtained by determining the sequence reads mapped to the segment. In some embodiments, the sequence depth may be obtained by determining the sequence reads mapped to the capture probes corresponding to the segment.
[0102] In some embodiments, two or more capture probes overlap (i.e., the capture probes can hybridize to overlapping sequences within the region of interest). Two or more capture probes may overlap by about 0%-10%, about 10%-20%, about 20%-30%, about 30%-40%, about 40%-50%, about 50%-60%, about 60%-70%, about 70%-80%, about 80%-90%, or about 90%-99% of the probe length. In some embodiments, two or more capture probes overlap 100%. In some embodiments, the number of sequences attributable to two or more capture probes correlates with each other. Overlapping or correlated capture probes may be considered by merging (i.e., summing) the number of sequence reads attributable to the overlapping or correlated capture probes.
[0103] When multiple sequence reads are mapped to the queried segment or multiple spatially adjacent segments (including the queried segment), the number of sequence reads mapped to the queried segment or spatially adjacent segments (including the queried segment) can be determined by counting the number of sequence reads assigned to the segments.
[0104] Building, initializing, and maximizing copy number likelihood models The copy number likelihood model can be any statistical model that can be used to determine the likelihood of observing the number of sequence reads mapped at a segment, given the copy number state of the segment. The initial copy number likelihood model is a model in which the parameters of the model are defined. "copy number likelihood model" refers to a model that is based on a copy number state, but before the model is optimized. In preferred embodiments, the copy number likelihood model includes one or more likelihood distributions of the expected number of mapped sequence reads given a copy number state. That is, each likelihood distribution corresponds to a copy number state. For example, the copy number likelihood model may include a likelihood distribution of the expected number of sequence reads given a copy number state of 1, a likelihood distribution of the expected number of sequence reads given a copy number state of 2, a likelihood distribution of the expected number of sequence reads given a copy number state of 3, and a likelihood distribution of the expected number of sequence reads given a copy number state of 4. The copy number likelihood model need not include a likelihood distribution for each possible copy number state, but will include at least one likelihood distribution. Similarly, the copy number likelihood model may include distributions for copy number states greater than 4, such as a copy number state of 5, 6, 7, or 8. In some embodiments, the distribution included in the copy number likelihood model is a Poisson distribution. In some embodiments, the distribution included in the copy number likelihood model is a binomial distribution. In some embodiments, the copy number likelihood model includes a negative binomial distribution. For example, in some embodiments, the copy number likelihood model is based on the copy number state c i,j
[0046] includes one or more negative binomial distributions (or one or more negative binomial distributions where the negative binomial distributions are not Poisson distributions) for the expected mapped sequence reads of queried segment i in test sequence library j.
[0105] The likelihood distribution of the copy number likelihood model can be further characterized by its mean (μ) and variance (d). The mean and variance of the likelihood distribution are optimized by using the expected number of sequence reads determined by sequencing test sequence library j at multiple segments (i.e., using the same capture probe) at segment i, and by setting the copy number state at segment i in sequence library j. The expected number of sequence reads is based on at least three factors: the number of mapped sequence reads of a representative value of the segment across multiple sequence libraries, the number of mapped sequence reads of a representative value of the test sequence library across multiple segments, and the local copy number state of the segment. The mean of the distribution is given by μ = c i,j μ i μ j and can be set as μ i is N s is the number of representative mapped sequence reads for segment i across sequence libraries, and μ j is N p is the number of representative mapped sequence reads for test sequence library j across segments, and c i,j is the copy number state at segment i of test sequence library j, and k i,j is the determined number of sequence reads in segment i of test sequence library j, and μ i and / or μ j is normalized.
[0106] Formally,
number
[0107] The copy number likelihood model is established by determining the distribution of expected sequence read counts for different copy number states, and then determining the most probable c given the actual mapped sequence read counts for a segment. i,j is maximized for
[0108] For the majority of genes, the expected copy number (i.e., "wild type") is assumed to be 2 (i.e., diploid). This is not always the case. For example, For genes on the Y chromosome, it may be assumed that the expected copy number (i.e., "wild type") is 1. Given this relationship, in some embodiments, the copy number likelihood distribution for any given copy number state is centered around the representative value,
number
[0109] The copy number likelihood distribution also includes a variance (d), which is estimated for segment i as follows:
number
[0110] The copy number likelihood distribution can be a Poisson distribution, a binomial distribution, a negative binomial distribution (such as a generalized Poisson negative binomial distribution or a non-Poisson negative binomial distribution), or any other suitable distribution. A negative binomial distribution, which is not a Poisson distribution, has been found to be particularly useful for determining copy number likelihood distributions. Figure 3A shows a plot of the average number of sequence reads ("mean depth") of approximately 2,500 capture probes used to enrich sequence libraries for regions of interest in multiple different test sequence libraries versus the sequencing depth variance. The data was fitted using a negative binomial distribution, which is not a Poisson distribution. For comparison, a Poisson distribution, which assumes a linear relationship between variance and mean depth, is also illustrated. As can be seen in Figure 3A, the data contradicts the Poisson assumption that the mean sequencing depth is equal to the sequencing depth variance, as plotting the data shows that the variance is greater than the mean. Thus, the data fit significantly better to the negative binomial distribution than to the Poisson distribution.
[0111] Figure 3B illustrates a copy number likelihood model including copy number likelihood distributions for copy numbers of 1 (CN=1), 2 (CN=2), and 3 (CN=3). Figure 3B illustrates a Poisson distribution and a negative binomial distribution, where the negative binomial distribution is not a Poisson distribution for each copy number. The distribution is a probability mass function (pmf) as a function of the number of sequence reads from capture probes corresponding to the segment.
[0112] Building a Hidden Markov Model A hidden Markov model allows for the determination of the most probable copy number (hidden state) from the number of mapped sequence reads (observed states). Generally, a hidden Markov model has four main parameters: one or more hidden states, one or more observed states, one or more release probabilities from the hidden states to the observed states, and transition probabilities between the hidden states. Provided herein are methods for constructing a hidden Markov model and parameterizing the hidden Markov model. Also provided herein are methods for training a hidden Markov model using an incomplete data set. Also provided herein are methods for optimizing the hidden Markov model parameters by considering variables that affect the release probability between the hidden state and the observed state. Specifically, provided below are methods and descriptions for hidden Markov model layers, Markov model transition probabilities, copy number likelihood models, parameterizing a hidden Markov model using expectation maximization, adjusting a hidden Markov model to consider the number of latent variables, and solving a hidden Markov model.
[0113] An exemplary hidden Markov model that can be used in the disclosed method is illustrated in Figure 4A. In Figure 4A, c1, c2, c3, and c4 represent hidden states (i.e., the most probable copy numbers of four different segments, although it is understood that the model can include n segments), and k1, k2, k3, and k4 represent observed states (i.e., the mapped sequence read numbers of each corresponding segment). The transition probabilities are the probability of transitioning from the copy number of one segment to the copy number of an adjacent segment, and are represented by p(c2|c1), p(c3|c2), and p(c4|c3). Finally, the probabilities of the hidden states (i.e., the copy numbers of the segments) given the observed states (the mapped sequence read numbers of the segments) are represented by p(c1|k1), p(c2|k2), p(c2|k2), and p(c2|k2). The latter are the posterior probabilities to be solved. To determine the posterior probability, p(k n |c n) copy number likelihood model is used.
[0114] In some embodiments, the hidden Markov model includes only one hidden state and a corresponding observed state. In some embodiments, the hidden state corresponds to the copy number state of a segment, and the observed state corresponds to the number of mapped sequence reads in the segment. In some embodiments, the hidden Markov model includes multiple hidden states and multiple observed states. In some embodiments, the multiple hidden states correspond to the copy number states in multiple segments, and the multiple observed states correspond to the number of mapped sequence reads in the multiple segments. In some embodiments, each segment in the region of interest corresponds to a capture probe in the region of interest. In some embodiments, two adjacent hidden states correspond to two spatially adjacent segments in the region of interest.
[0115] Segments may be divided into subsegments, as previously described herein. In some embodiments, the hidden state corresponds to the copy number of the subsegment. A subsegment does not contain a mapped sequence read number that is independent of the mapped sequence read number of its parent segment (i.e., the segment of which the subsegment is a member). In some embodiments, a mapped sequence read number for a segment is attributed to each subsegment within the segment. In some embodiments, a subsegment contains hidden state (i.e., copy number), but a mapped sequence read number is attributed only to the first subsegment of the segment. This is exemplified in Figure 4B, which includes two segments identified by dashed lines: segment A and segment B. Segment A includes subsegment 1, subsegment 2, and subsegment 3, while segment B includes subsegment 4, subsegment 5, and subsegment 6. The mapped sequence read number for segment A is attributed to subsegment 1, the first subsegment of that segment. The mapped sequence read number for segment B is attributed to the first subsegment of that segment. The first subsegment in the sequence belongs to subsegment 4, the first subsegment in the sequence. C1, C2, C3, C4, C5, and C6 represent the hidden states (copy numbers) of each subsegment, and k1 and k4 represent the observed states (number of sequence reads) of subsegment 1 and subsegment 4, respectively. The transition probabilities between the hidden states of the subsegments are identified by p(c2|c1), p(c3|c2), p(c4|c3), p(c5|c4), and p(c6|c5). Because only subsegments 1 and 4 contain observed states, only two probabilities of the number of mapped sequence reads given the copy numbers of the subsegments are included: p(k1|c1) and p(k4|c4).
[0116] The copy number state of a segment is related to the number of sequence reads that have been mapped to that location. The number of mapped sequence reads (k i,j ) is given, the segment or subsegment (c i,j Determining the copy number state of a segment (which can be denoted as k) allows calling the copy number of that segment or subsegment. The probability that a given copy number state is the correct copy number depends at least on the number of mapped sequence reads. In Bayesian statistics, k i,j (i.e., p(c i,j |k i,j )) given c i,j The posterior probability of can be determined using a copy number likelihood distribution. The posterior probability is the probability of a parameter given some data, while the likelihood model is the probability of data given the parameters. In this case, the posterior probability is the probability of a copy number state of a segment or subsegment given the number of sequence reads mapped at the segment or subsegment (i.e., p(c i,j |k i,j )), whereas the copy number likelihood model measures the likelihood of observing the number of sequence reads mapped at a segment given the copy number state of the segment (i.e., p(ki,j |c i,j )). p(c i,j |k i,j ) cannot be determined directly, so we use the copy number likelihood model p(k i,j |c i,j ) can be used to parameterize the hidden Markov model, which can be used to calculate the posterior probability p(c i,j |k i,j ) can be solved. Below, we consider the copy number likelihood model as a negative binomial distribution, but it is understood that similar aspects apply to other distribution forms. In some embodiments, the copy number likelihood model is p(k i,j |c i,j )=NegBinom(k i,j |μ c,i,j =c i,j μ i μ j ;d=d i ) can be defined as k i,j is the number of mapped sequence reads in segment i of test sequence library j.
[0117] The negative binomial distribution is parameterized to best fit the data. In its simplest form, the copy number likelihood model is a negative binomial model. However, depending on the data generated, a different type of distribution may fit the data better and may be more suitable. The general aspects of the present invention may apply to models that include different statistical distributions.
[0118] The copy number transition probability of a segment or subsegment depends, in part, on the copy number state of spatially adjacent segments or subsegments. The length and frequency of copy number variants can also affect the transition probability.
[0119] In some embodiments, the transition probabilities can be predetermined or fixed. In preferred embodiments, the transition probabilities are variable. For example, assuming hidden copy number states limited to 0, 1, 2, 3, or 4 copies (assuming a wild-type copy number of 2), the transition probabilities can be formally represented by the following probabilistic transition matrix:
number
[0120] Copy number variants have a representative length, and copy numbers longer or shorter than this length tend to be lower than the copy number of the representative length. In some embodiments, the transition probability (or multiple transition probabilities) takes into account the representative length of the copy number variant. The representative length of the copy number variant can be based on observations from a historical population (e.g., a historical human population). The historical population is the historical population of the sequence library from which the copy number variants were called. A larger historical population can result in more accurate representative copy number variant lengths. In some embodiments, the historical population includes about 1,000 or more sequence libraries (e.g., about 5,000 or more, about 10,000 or more, about 25,000 or more, about 50,000 or more, about 100,000 or more, about 250,000 or more, or about 500,000 or more sequence libraries, etc.). The representative length of the copy number variant is predetermined. In some embodiments, the length of the representative copy number variant is about 3,000 to about 1,000 bases (e.g., about 4,000 to about 8,000 bases, about 5,000 to about 7,000 bases, about 5,500 to about 6,500 bases, or about 6,200 bases). Taking the length of the representative copy number variant into account, the transitions (or subsegment transition probabilities) of the probabilistic transition matrix used to calculate the transition probability for each base are
number
number
[0121] Transition probabilities can also consider the probability of a copy number variant at the queried segment given the copy number states at spatially adjacent segments. Certain portions of the genome may contain "hotspots" of genetic variants, including copy number variants. Hotspots refer to regions in the genome that exhibit a high propensity for any type of mutation. This may be due to the structural organization of the region or functional aspects of the region, which make the region prone to mutation. The probability of a copy number variant at any given segment (such as the queried segment or a spatially adjacent segment) may be based on observations from a historical population (e.g., a historical human population). The historical population is the historical population of the sequence library from which the copy number variants were called. The larger the historical population, the more accurate the copy number variants. In some embodiments, the historical population comprises about 1000 or more sequence libraries (e.g., about 5000 or more, about 10,000 or more, about 25,000 or more, about 50,000 or more, about 100,000 or more, about 250,000 or more, or about 500,000 or more sequence libraries, etc.). To consider the probability of copy number variants in the queried segment or spatially adjacent segments, the transitions in the probabilistic transition matrix are
number
number
[0122] In some embodiments, the hidden Markov model includes one transition probability of the copy number state of the segment or subsegment. In some embodiments, the hidden Markov model includes multiple transition probabilities of the copy number state of the segment or subsegment. In some embodiments, the transition probability of the copy number state given the copy number state of the adjacent preceding segment depends on the length of the copy number variant. In some embodiments, the length of the copy number variant is specific to that particular region of the genome. In some embodiments, the length of the copy number variant is the length of the representative value of the copy number variant across the genome.
[0123] In some embodiments, the copy number state transition probability given the copy number states of the adjacent preceding segments depends on the probability of observing a copy number variant. In some embodiments, the probability of observing a copy number variant is specific to that particular region of the genome. In some embodiments, the probability of observing a copy number variant is the representative probability of observing a copy number variant across the genome.
[0124] Parameterizing hidden Markov models and determining the most probable copy number As described above, a hidden Markov model includes (i) one or more hidden states including copy numbers corresponding to one or more segments or subsegments (including at least the queried segment or a subsegment of the queried segment), (ii) one or more observed states including the number of sequence reads mapped to the one or more segments, and (iii) a copy number likelihood model. The copy number likelihood model defines the probability of observing an observed state for a given hidden state (i.e., p(k i,j |c i,j )) The Hidden Markov Model also includes hidden state transition probabilities, which can be fixed or variable as described above.
[0125] The hidden Markov model is initiated using a copy number likelihood model. The hidden Markov model can also be initiated by assuming that the copy number state (i.e., hidden state) has a wild-type copy number (e.g., 2 copies), which can be used to back-calculate the transition (r) to determine the transition probability. The copy number likelihood model is based on the expected number of sequence reads mapped to the segment, as explained above, but the copy number likelihood model can also be calculated using, for example, the mean μ of each copy number likelihood distribution in the copy number likelihood model. c,i,j and variance d i can be adjusted to fit the determined number of sequence reads mapped to the segment (i.e., the observed states) by allowing them to float when parameterizing the hidden Markov model. The transition probabilities, if variable, can also be adjusted during parameterization of the hidden Markov model.
[0126] Parameterizing a hidden Markov model involves adjusting a copy number likelihood model to fit a determined number of sequence reads mapped to a segment (e.g., a queried segment or a spatially adjacent segment). In some embodiments, the copy number likelihood model is optimized to fit a determined number of sequence reads mapped to a segment (e.g., a queried segment or a spatially adjacent segment). The copy number likelihood model is "optimized" after multiple rounds of adjustment to best fit the observed state. In some embodiments, parameterizing a hidden Markov model involves adjusting (or optimizing) transition probabilities. A hidden Markov model can be parameterized by optimizing the copy number likelihood model using analytical first-derivative gradients and second-derivative Hessians of one or more parameters of the copy number likelihood model, which can be solved, for example, using a trust-region Newton conjugate gradient algorithm. Typical parameters for the optimized copy number likelihood model include the variance (d) of the mapped sequence read numbers of the segment. i ), the number of mapped sequence reads for the representative segment (μi ), the variance of the number of mapped sequence reads of the segments in the test sequence library (d j ), or the number of mapped sequence reads (μ j ) The expectation-maximization (EM) algorithm can then be used to optimize the parameters in one or more iterations of applying the Hidden Markov Model to determine the most probable number of copies of the segment (e.g., using a Viterbi algorithm, a Quasi-Newton solver, or Markov Chain Monte Carlo with a Baum-Welch algorithm) and reparameterization of the Hidden Markov Model.
[0127] For example, expectation maximization (EM) can be used to adjust (or optimize) a copy number likelihood model (based on the expected number of sequence reads) and / or one or more additional model parameters to maximize expected sequence reads mapped to a segment (i.e., adjusted μ c , i,j ) and the adjusted variance of that segment (i.e., adjusted d i ), i.e., such that the probability of the expected number of sequence reads in the queried segment is maximized for a given copy number state in that segment.
[0128] In general, expectation maximization (EM) can be used to estimate latent or unknown parameters despite incomplete data. The EM algorithm consists of an expectation "E" step that selects the most likely copy number distribution from a copy number likelihood model given the determined number of sequence reads mapped to a segment (so that the most probable copy number can be determined), and ...). c,i,j and d iThe EM process can be iteratively switched between a maximization "M" step, which re-estimates the probability of any c. The maximization step assumes a fixed probability model and number of sequence reads, and finds the copy number state that, when applied to the model, results in the highest probability of the actual mapped sequence read number from all other possible copy numbers. The EM process can be applied to different parameters of the HMM; for example, the EM process can use the expected values generated in the "E" step to consider the transitions (r) between hidden states, if applicable. Briefly, EM is used to maximize the model to find any c i,j Formally, the Viterbi algorithm can determine the maximum likelihood of a copy number likelihood model as follows:
number
[0129] In some embodiments, the BaumWelch algorithm determines the copy number code of a segment. The Baum-Welch algorithm is used in the expectation step of the EM process to determine the expected probability of a copy number state at segment i for a given number of mapped sequence reads at segment i. i |k [0,i] ) and likelihood β(k [i,I] |c i ) and The Baum-Welch algorithm can be solved using methods known to those skilled in the art.
[0130] A parameterized hidden Markov model can be used to determine the most probable copy number of the queried segment or a subsegment of the queried segment during the maximization step. The most probable copy number of the queried segment can be determined using any useful algorithm known in the art, such as the Viterbi algorithm, a quasi-Newton solver, or Markov Chain Monte Carlo.
[0131] GC content bias correction The GC content of a segment of interest or a capture probe corresponding to the segment can affect the number of sequence reads mapped to the segment, for example, due to differences in the hybridization efficiency of the capture probe. Thus, depending on the GC content, the capture probe can have a strong influence on the number of sequence reads mapped to the segment, regardless of the copy number status of the segment. This GC content bias is well known and has been described in the art. In some embodiments of the methods described herein, GC content bias is taken into account when determining the copy number of a segment. GC content bias correction can be useful for any method of determining copy number variants and does not need to be used only with direct target sequences. For example, in some embodiments, GC content bias is corrected when determining the copy number of a segment within the region of interest, and the sequence library is enriched using hybrid capture technology. In addition, the method for correcting GC content bias does not need to be limited to methods of determining copy number using hidden Markov models, but GC content bias can be corrected for any method, including the use of copy number likelihood models.
[0132] In some embodiments, the sequence read number of any given segment (such as the expected sequence read number used to determine copy number likelihood model) is corrected for GC content by multiplying the sequence read number by GC bias correction coefficient.The GC bias correction coefficient is specific to a given segment and to a test sequence library.That is, the GC bias correction coefficient is uniquely determined for a segment and a test sequence library, and the GC bias correction coefficient must be re-determined for different segments and for each different test sequence library.
[0133] The number of sequence reads mapped to a given segment (which may include the queried segment) can be normalized by dividing the number of mapped sequences in that segment by the number of mapped sequence reads of a representative value of multiple segments enriched from the test sequence library. The normalized number of sequence reads for each segment within the multiple segments can be plotted against the GC content of that segment. The data points are then subjected to a quadratic correction, g i,j =a+b(GC)+c(GC) 2 can be fitted using g i,j is the GC bias correction factor specific to segment i of a multi-segment test sequence library j, (GC) is the GC content, and a, b, and c are constants determined by quadratic fitting.
[0134] Thus, a GC bias correction factor can be determined by fitting a quadratic function to a plurality of data points, each data point comprising a normalized number of sequence reads mapped to a segment and the GC content of that segment, the plurality of data points representing a plurality of segments enriched by capture probes in a test sequence library, and defining the GC bias correction factor to be the normalized number of sequence reads determined by a quadratic function of the GC content of the segment.
[0135] The copy number likelihood model can be adjusted in a similar manner to account for the presence of GC content bias. That is, the expected sequence read count used as the basis for the copy number likelihood model can be adjusted to account for the presence of GC content. For example, the representative value of the copy number likelihood distribution in the model can be adjusted as follows: μ c,i,j =c i,j μ i μ j g i,j Furthermore, the copy number likelihood model p(k i,j |c i,j )=NegBinom(k i、j |μ c,i,j =c i,j μ i μ j g i,j ,d) can be formulated as follows, and k i,j refers to the number of sequence reads at segment i in test library j, and d is the number of sequence reads at segment i in test library j. i , d j , or d i,j is.
[0136] In some embodiments, there is a method for determining the copy number of an interrogated segment or a subsegment of the interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a segment within the region of interest, wherein the test sequence library is enriched using capture probes; (b) determining the number of sequence reads mapped to the segment; (c) determining a copy number likelihood model for the segment based on the expected number of mapped sequence reads at the segment, wherein the expected number of mapped sequence reads is corrected for the GC content of the segment; and (d) determining the most probable copy number of the interrogated segment based on the copy number likelihood model. The most probable copy number of the interrogated segment can be determined based on the copy number likelihood model using the hidden Markov model described herein, or by any other method known in the art. For example, the most probable copy number can be determined based on the maximum copy number probability of each region based on the capture probes for that region. In another example, the most probable copy number can be determined using a brute force segmentation approach.
[0137] In some embodiments, there is a method for determining the copy number of an interrogated segment or a subsegment of the interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to the interrogated segment, wherein the test sequence library is enriched using one or more capture probes; (b) determining a number of sequence reads mapped to the interrogated segment; and (c) determining a copy number likelihood model based on an expected number of sequence reads mapped to the interrogated segment, wherein the expected number of mapped sequence reads is proportional to the GC content of the interrogated segment. (d) constructing a hidden Markov model including: (i) one or more hidden states including copy numbers corresponding to the queried segment or a plurality of subsegments within the queried segment; (ii) observed states including sequence read numbers mapped to the queried segment; and (iii) a copy number likelihood model; (e) parameterizing the hidden Markov model by adjusting the copy number likelihood model to fit the determined sequence read numbers mapped to the queried segment; and (f) calculating the copy number likelihood of the queried segment or the subsegments within the queried segment based on the parameterized hidden Markov model. and determining the most probable copy number of a subsegment of the segment selected.
[0138] In some embodiments, there is a method for determining the copy number of an interrogated segment or a subsegment of the interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, wherein the plurality of spatially adjacent segments comprises the interrogated segment, and wherein the test sequence library is enriched using a plurality of spatially adjacent capture probes; (b) determining a number of sequence reads mapped to each spatially adjacent segment; and (c) determining a copy number likelihood model for each spatially adjacent segment based on expected numbers of mapped sequence reads in the spatially adjacent segments, wherein the expected numbers of mapped sequence reads are determined based on the GC content of the spatially adjacent segments. (d) constructing a hidden Markov model including (i) a plurality of hidden states including copy numbers of each of the spatially adjacent segments or a plurality of subsegments within each of the spatially adjacent segments, (ii) a plurality of observed states including sequence read numbers mapped to each of the spatially adjacent segments, and (iii) a copy number likelihood model for each of the spatially adjacent segments; (e) parameterizing the hidden Markov model including adjusting each copy number likelihood model to fit the determined sequence read numbers mapped to each of the spatially adjacent segments; and (f) determining the most probable copy number of the queried segment or a subsegment of the queried segment based on the parameterized hidden Markov model.
[0139] Spurious Capture Probes Certain capture probes used to enrich segments within a region of interest can generate spurious results. For example, the number of sequence reads generated by a spurious capture probe may not match the copy number of the corresponding segment, either due to under- or over-enrichment of the segment. These spurious results can arise, for example, due to the design of the capture probe or sequence variants (e.g., SNPs) within the sequence to which the capture probe is designed to hybridize. Spurious capture probes can affect the number of mapped sequence reads and artificially confound copy number likelihood models and parameters. Therefore, it is desirable to consider spurious capture probes. Spurious capture probes do not necessarily have to be direct target sequence capture probes; similar methods can be applied to capture probes used to enrich test sequence libraries (e.g., hybrid capture techniques). Determining whether a capture probe is a spurious capture probe can be performed using EM. For example, the determination of whether a capture probe is spurious can be made during the expectation step; if the probability of the capture probe being spurious changes during the EM iterations, the maximization step will also change, and the maximization step will determine the most likely copy number state of the segment, taking into account the new spuriousness of the capture probe. If the capture probe is determined to be a spurious capture probe, the probability of the mapped sequence read number of the copy number state segment is set to 1 during the expectation maximization process. Setting the probability to a constant allows the model to effectively discard spurious capture probes, since they do not provide additional information and therefore are not considered when the model is parameterized. The determination of the spuriousness of a capture probe can be iterated, for example, by determining whether the capture probe is spurious after several EM cycles.
[0140] In some embodiments, a Bernoulli process is used to determine whether a given capture probe is spurious. The Bernoulli process can be applied to some or all of the capture probes; that is, for each capture probe, its spuriousness is determined independently. For capture probe i, the indicator variable b i is introduced, where 1 means that the capture probe t is spurious and 0 means that the capture probe is not spurious.
number
[0141] By using this index, it is possible to take spurious capture probes into account by adjusting the copy number likelihood model. If a capture probe is determined to be spurious, the probability of the number of mapped sequence reads of the corresponding segment at any given copy number is set to 1. If the capture probe is not spurious, the copy number likelihood distribution of the copy number likelihood model is not changed. Formally,
number
[0142] The spuriousness of a capture probe may depend on the test sequence library. That is, some test sequence libraries may be more prone to spurious capture probes than other test sequence libraries. In some embodiments, whether a test sequence library is prone to spurious capture probes is determined based on a prior distribution of the test sequence library. In some embodiments, determining whether a test sequence library is prone to a particular probe being spurious depends on a general prior distribution.
[0143] Figure 5B illustrates an example prior distribution that can be adjusted to determine whether a given capture probe is a spurious capture probe. i,jis the observation status (number of mapped sequence reads) of segment i, k i The indicator variable b is the Bernoulli prior distribution for i,j can be specific to segment i and to test sequence library j. The test sequence library prior distribution π j is the indicator variable b t is set to be the same across all segments in the region of interest of the test sequence library. The general prior distribution π is set to be the same across all segments in the region of interest of the test sequence library. j The prior distribution π is set for π and is the same for all similarly enriched sequence libraries. A general prior distribution π can be predetermined and validated to reduce false calls without losing sensitivity. The tuning step (such as the maximization step of the EM algorithm) can be set by assuming that the capture probes follow a Bernoulli distribution with spurious probabilities. The prior distribution π j The probability of capture probe i in test sequence library j being spurious given σ can be expressed as:
number
[0144] Given the determined number of sequence reads mapped to spatially adjacent segments (or spatially adjacent subsegments) 0 to I, the probability that capture probe i is spurious can be derived as follows:
number
[0145] Indicator expected value b i Given a test sequence library prior distribution π jcan be determined as follows:
number
[0146] In some embodiments, the most probable copy number of the interrogated segment, or one or more subsegments of the interrogated segment, is not called if the capture probe associated with the interrogated segment is determined to be spurious. In some embodiments, the most probable copy number of the interrogated segment, or one or more subsegments of the interrogated segment, is called based on the probability that capture probe i is spurious (i.e., p(b i |k [0,I] )) is above a predetermined threshold (such as about 0.1 or more, about 0.2 or more, about 0.3 or more, about 0.4 or more, or about 0.5 or more), no call is made.
[0147] In some embodiments, there is a method for determining the copy number of an interrogated segment or a subsegment of the interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to the interrogated segment, wherein the test sequence library is enriched using one or more capture probes; (b) determining a number of sequence reads mapped to the interrogated segment; (c) determining a copy number likelihood model based on the expected number of sequence reads mapped to the interrogated segment; and (d) a hidden Markov model that estimates the copy number of the interrogated segment or a subsegment of the interrogated segment. (i) constructing a hidden Markov model including one or more hidden states including copy numbers corresponding to a plurality of subsegments within the queried segment; (ii) an observed state including sequence read numbers mapped to the queried segment; and (iii) a copy number likelihood model; (e) parameterizing the hidden Markov model by adjusting the copy number likelihood model to fit the determined sequence read numbers mapped to the queried segment and taking into account one or more spurious capture probes; and (f) determining the most probable copy number of the queried segment or a subsegment of the queried segment based on the parameterized hidden Markov model.
[0148] In some embodiments, there is a method for determining the copy number of an interrogated segment or a subsegment of the interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, wherein the plurality of spatially adjacent segments includes the interrogated segment, and wherein the test sequence library is enriched using a plurality of spatially adjacent direct target sequence capture probes; (b) determining a number of sequence reads mapped to each spatially adjacent segment; (c) determining a copy number likelihood model for each spatially adjacent segment based on expected numbers of mapped sequence reads in the spatially adjacent segments; and (d) a hidden Markov model having: (i) a plurality of hidden states including copy numbers of each of the spatially adjacent segments or a plurality of subsegments within each of the spatially adjacent segments; (ii) a plurality of observed states including numbers of sequence reads mapped to each of the spatially adjacent segments; and (iii) an empty state. (e) parameterizing the hidden Markov model, including adjusting each copy number likelihood model to fit the determined sequence read numbers mapped to each spatially adjacent segment and taking into account one or more spurious capture probes; and (f) determining the most probable copy number of the queried segment or a subsegment of the queried segment based on the parameterized hidden Markov model.
[0149] Noisy test array library During the preparation of a test sequence library, some steps can cause the nucleic acids in the test sequence library to be prone to "noise" across a large number of capture probes. This leads to inconsistent data and a large number of false positives. Figure 6A shows an example of a test sequence library with low noise, and Figure 6B shows an example of a test sequence library with more noise, even when two sequence libraries are enriched using the same capture probe library. Noise can be introduced, for example, during the preparation or sequence of the test sequence library, isolation of nucleic acids from test samples, storage of the sequence library, or fragmentation of nucleic acids isolated from test samples, which can compromise the integrity of oligonucleotides, which can in turn affect the oligonucleotide method.
[0150] In some embodiments, parameterizing the hidden Markov model includes accounting for noise in the number of mapped sequence reads. In some embodiments, accounting for noise in the number of mapped sequence reads includes adjusting a copy number likelihood model. For example, parameterizing the hidden Markov model may include an expectation maximization step, and accounting for noise may occur during the expectation maximization step.
[0151] The variance d of the copy number likelihood distribution in the copy number likelihood model was discussed above. If only the variance due to the capture probes (i.e., in the segments) is considered, then d=d i The variance of the copy number likelihood distribution can also be used to account for noise across segments in the test sequence library j. The variance due to accounting for noise across segments in the test sequence library j and the noise due to capture probes can be determined by arithmetic combination, for example, by multiplying or adding the variance due to sequence library noise and the variance due to capture probe noise. For example, in some embodiments, the variance of the copy number likelihood distribution can be formally viewed as: d=d i *d j The variance due to sequence library noise and the variance due to capture probe noise can be combined in some embodiments, for example, by adding:
number
[0152] The parameterization of the hidden Markov model adjusts the copy number likelihood model, including the variance of the copy number likelihood distribution due to the model. Therefore, both components of the variance d (i.e., d i and d j ) can be adjusted during parameterization of the hidden Markov model, for example, using the expectation-maximization algorithm. In some embodiments, to account for noise, a test sequence library (d j ) are used. In some embodiments, a quasi-Newton method can be used to account for noise during the maximization step. In particular, the expectation step seeks to maximize
number
[0153] During the ceremony,
number
number
number
number
number
[0154] Once the parameters of the distribution are set, a parameterized hidden Markov model can be used to determine the most probable copy number state of the segment.
[0155] Performance of copy number variant screens In certain embodiments, the methods described herein are used to evaluate the sample-specific performance of a copy number variant screen or copy number variant model. Synthetic copy number variants are generated in silico using actual sequence reads from a test sample. Thus, the synthetic copy number variants are sample-specific. The copy number variant model is parameterized using actual sequence read numbers mapped to segments within the region of interest of the test sample to determine copy number variant model parameters. Because the synthetic copy number variants are based on the test sample and the determined copy number variant model parameters are sample-specific, the determined sample-specific copy number variant model parameters are based on the components of the segments within the synthetic copy number variants. Used by copy number variant callers to call copy numbers.
[0156] The synthetic copy number variants include synthetic copy numbers of one or more segments having a region of interest, where the synthetic copy number is represented by the synthetic sequence read number from one or more segments within the region of interest. In some embodiments, the synthetic sequence read number is obtained by adjusting the sequence read number of one or more segments within the region of interest from a test sample. The adjustment is made in proportion to the synthetic copy number. In some embodiments, the synthetic sequence read number is obtained by directly manipulating a database containing sequence reads of one or more segments within the region of interest from an actual sample, for example, by random deletion or duplication of sequence reads in the database. In some embodiments, the synthetic sequence read number is generated by sampling a distribution (such as a binomial distribution or a negative binomial distribution). Multiple synthetic copy number variants can be generated, for example, based on multiple test or reference samples.
[0157] The synthetic copy numbers of one or more segments within the region of interest present in the synthetic copy number variant are called using a copy number variant caller. In some embodiments, the caller compares the synthetic sequence read counts from one or more segments within the synthetic copy number variant with the sequence read counts from one or more segments in an actual reference sample having known copy numbers of the segments. The caller can determine the copy numbers of the segments within the synthetic copy number variant using, for example, a hidden Markov model (HMM) as described herein. The actual reference sample is preferably a different actual sample than the actual sample used as the basis for generating the synthetic copy number variant.
[0158] The copy number variant caller uses synthetic copy number variants and determined copy number variant model parameters, as shown in FIG. 9 . The actual sequence read counts from the test sample mapped to the segment within the region of interest are used to initialize a copy number variant model to determine initial copy number variant model parameters, such as a copy number variant model in a hidden Markov model. The copy number variant model can be parameterized to determine the copy number variant model parameters, for example, using analytical first-derivative gradients and second-derivative Hessians. Using the initial CNV model parameters, a CNV model is applied, for example, using the Viterbi algorithm and the Baum-Welch algorithm. An expectation-maximization step can be iteratively performed to optimize the CNV model parameters to fit the actual sequence read counts, thereby determining one or more copy number variant model parameters optimized for the test sample (i.e., sample-specific copy number variant model parameters). The copy number variant model can call copy number variants in the test sample using the sample-specific copy number variant model parameters and the actual sequence read counts of the segment. The actual sequence read numbers from the test sample are also used to generate synthetic sequence read numbers, which are used to represent the synthetic copy numbers of segments within the region of interest of the synthetic copy number variant. Multiple synthetic copy number variants, for example, about 10 to about 10,000 synthetic copy number variants, can be generated in this manner. Copy number variant models and sample-specific copy number variant model parameters can use the synthetic sequence read numbers to call the copy numbers of one or more segments of the synthetic copy number variant.
[0159] Performance statistics for copy number variant screens can be determined to assess sample-specific performance of the copy number variant screen based on the difference between the called copy number and the synthetic copy number of the synthetic copy number variants. Because multiple synthetic copy number variants are generated and called by the caller, performance statistics can be used to evaluate the sample-specific performance of the copy number variant screen in the context of the synthetic variants. Therefore, a greater diversity of synthetic copy number variants (which can be based on multiple real samples) provides more accurate performance statistics to characterize the performance of a copy number variant model.
[0160] In some embodiments, a method for evaluating sample-specific performance of a copy number variant model includes: parameterizing the copy number variant model based on actual sequence read numbers mapped to segments within a region of interest from a test sample to determine one or more copy number variant model parameters; generating a plurality of synthetic copy number variants, each synthetic copy number variant comprising synthetic copy numbers of one or more of the segments, wherein each synthetic copy number is represented by a synthetic sequence read number based on the actual sequence read number of the corresponding segment from the test sample; calling copy numbers of one or more segments of the synthetic copy number variants using the copy number variant model and the one or more determined copy number variant model parameters; determining sample-specific performance statistics for the copy number variant model based on the difference between the called copy numbers and the synthetic copy numbers of the synthetic copy number variants; and evaluating the sample-specific performance of the copy number variant model based on the sample-specific performance statistics. In some embodiments, the sample-specific performance statistic is a limit of detection, sensitivity, specificity, precision, recall, accuracy, positive predictive value, or negative predictive value.
[0161] In some embodiments, the copy number variant caller uses a hidden Markov model to call the copy number of a synthetic copy number variant. Because copy number variants of a given segment are relatively rare, it can be assumed that the test sample does not have a copy number variant in the given segment. Even if the test sample does not have a copy number variant, the test sample contains enough non-variant (i.e., wild-type) segments so that the evaluation method is reliable.
[0162] For the purpose of generating synthetic copy number variants, the test sample is assumed to be wild-type with respect to the copy number of the segment containing the region of interest, and the sequence read numbers can be assumed to form a negative binomial distribution with a representative value (mean or median) and variance.The variance of the distribution can be caused, for example, by segment enrichment or noise in the sequence.The distribution of sequence reads from the population of synthetic copy number variants is preferably similar to the expected negative binomial distribution of sequence reads from a theoretical population of actual copy number variants that are treated equivalently and therefore have the same copy number variant model parameters.
[0163] In some embodiments, a plurality of synthetic copy number variants are generated, including synthetic copy numbers of one or more segments represented by synthetic sequence read numbers from one or more segments within the region of interest. The synthetic sequence read numbers for each of the one or more segments can be generated by increasing, decreasing, or maintaining the actual sequence read numbers from one or more segments within the region of interest from the test sample. For example, if a first actual sequence read number corresponds to a first segment within the region of interest, and a second actual sequence read number corresponds to a second segment within the region of interest, and the test sample is assumed or expected to have two copies of the region of interest, a synthetic copy number variant having three copies of the region of interest can be generated by generating a first synthetic sequence read number corresponding to the first segment by increasing the first actual sequence read number to reflect three copies of the first segment, and generating a second synthetic sequence read number corresponding to the second segment by increasing the second actual sequence read number to reflect three copies of the second segment. The number of synthetic sequence reads corresponding to the first and second segments is increased to reflect three copies, so that the synthetic copy number variants are the first and second segments of interest. In some embodiments, synthetic sequence read numbers are generated by multiplying the actual sequence read numbers by a factor (e.g., 1.5 to increase the copy number from 2 to 3, or 0.5 to decrease the copy number from 2 to 1). In some embodiments, synthetic sequence read numbers are generated by adding (or subtracting) a sequence read number (e.g., 50% of the representative actual sequence read number corresponding to all segments within the region of interest) to the actual sequence read number. In some embodiments, the sequence read numbers are normalized (e.g., as described below) so that a single copy of the region of interest is represented by a normalized sequence read number (e.g., 0.5) and two copies of the region of interest are represented by a normalized sequence read number (e.g., 1). Thus, in some embodiments, a normalized sequence read number (e.g., 0.5) is added to the normalized sequence read number to increase the copy number of the synthetic copy number variant, and a normalized sequence read number (e.g., 0.5) is subtracted from the normalized sequence read number to decrease the copy number of the synthetic copy number variant. Preferably, the actual sequence read numbers are increased or decreased to generate synthetic sequence copy numbers to represent synthetic copy number variants having a predetermined number (which may be an integer or non-integer) of copies of the segment (such as 1 or more, 2 or more, 3 or more, 4 or more, or 5 or more copies of the segment).
[0164] In some embodiments, synthetic sequence read numbers are generated by adding or subtracting sequence read numbers from sequence read numbers from a test sample to generate synthetic copy number variants. Synthetic copy number variants containing duplications are generated by adding sequence read numbers, and synthetic copy number variants containing deletion events are generated by deleting sequence read numbers. The sequence read numbers added or subtracted from the sequence read numbers from the test sample are based in part on the number of duplication or deletion events simulated in the synthetic copy number variant. In some embodiments, the synthetic sequence read numbers of synthetic copy number variants containing n copies of a region of interest (or segment of region of interest) that are more (or less) than the assumed (e.g., wild-type) copy number x in the test sample are calculated by subtracting a representative (e.g., average or median) sequence read number from multiple test samples for that region of interest (or segment of region of interest) from the sequence read number of that region of interest (or segment of region of interest) to (or from) the sequence read number of that region of interest (or segment of region of interest).
number
number
number
number
number
[0165] In some embodiments, the number of synthetic sequence reads of a synthetic copy number variant comprising m copies of a region of interest (or segment of a region of interest) is
number
number
number
[0166] In some embodiments, the number of synthetic sequence reads for a synthetic copy number variant comprising m copies of the segment is equal to the number of sequence reads from the test sample.
number
number
number
number
number
number
[0167] In some embodiments, the number of synthetic sequence reads for a synthetic copy number variant is determined by sampling a binomial or negative binomial distribution of actual sequence reads from a test sample. For example, for a synthetic copy number deletion variant with m copies of a region of interest (or a segment of a region of interest), the number of synthetic sequence reads is
number
number
number
[0168] In some embodiments, synthetic sequence reads for synthetic copy number duplication variants having m copies of a region of interest (or segment of a region of interest) are generated by sampling from a negative binomial distribution, where the number of successes is equal to the number of actual sequence reads from a test sample having a hypothesized x copy number of the region of interest (or segment of a region of interest), and the probability of success is
number
number
number
number
number
number
number
[0169] The copy number variant caller can call the copy number of one or more segments within the region of interest of each synthetic copy number variant of the multiple synthetic copy number variants. The copy numbers of the segments within the synthetic copy number variants are known because they are represented by synthetic sequence read counts generated by adjusting the actual sequence read counts from the test sample to the desired number of copies of one or more segments. The called copy numbers can be compared to the copy numbers of each synthetic copy number variant of the multiple synthetic copy number variants to determine performance statistics of the copy number variant model. Performance statistics can be, for example, sensitivity, specificity, precision, recall, accuracy, positive predictive value, negative predictive value, or any other metric of agreement.
[0170] Performance statistics indicate the performance of copy number variant screens or models. For example, it is desirable for copy number variant models to have a high number of true positives and a low number of false negatives. Therefore, performance statistics can be used to evaluate the performance of copy number variant models. In some embodiments, a predetermined threshold value for the performance statistics can be selected. In some embodiments, if the performance statistics is below the predetermined threshold value, the test sample can be reanalyzed and / or a new set of sequence reads can be generated for the test sample.
[0171] Computer Systems In some embodiments, the methods described herein are implemented by a program running on a computer system. FIG. 11 depicts an exemplary computing system 1100 configured to perform any one of the processes described above, including various exemplary methods for calling copy numbers of queried segments or evaluating the performance of copy number variant models. The computing system 1100 may include, for example, a processor, memory, storage, and input / output devices (e.g., a monitor, keyboard, disk drives, an Internet connection, etc.). The computing system 1100 may include circuitry or other specialized hardware for performing some or all aspects of the processes. For example, in some embodiments, the computing system includes a sequencer (such as a massively parallel sequencer). In some operational settings, the computing system 1100 may be configured as a system including one or more units, each configured to perform some aspects of the processes in either software, hardware, or some combination thereof. do.
[0172] 11 depicts a computing system 1100 having several components that may be used to execute the processes described above. The main system 1102 includes a motherboard 1104 having an input / output ("I / O") section 1106, one or more central processing units ("CPUs") 1108 (e.g., processors), and a memory section 1110, which may have associated flash memory cards 1112. The I / O section 1106 is connected to a display 1114, a keyboard 1116, a disk storage unit 1118, and a media drive unit 1120. The media drive unit 1120 can read / write computer-readable media 1122, which may contain programs 1124 and / or data.
[0173] At least some values based on the results of the above-described processes can be saved for subsequent use. Additionally, a non-transitory computer-readable medium can be used to store one or more computer programs for performing any one of the above-described processes by a computer (e.g., one or more central processing units (“CPUs”) 1108 can execute one or more stored computer programs (or instructions) to perform the above-described processes). The computer program can be written, for example, in a general-purpose programming language (e.g., Pascal, C, C++, Java, Python, JSON, R, etc.) or some dedicated application-specific language.
[0174] In some embodiments, summary statistics are reported (e.g., to the patient, a physician, a caregiver, or a regulatory agency). In some embodiments, summary statistics are displayed, for example, on a monitor.
[0175] Various exemplary embodiments are described herein. Reference is made to these examples in a non-limiting sense. They are provided to illustrate more broadly applicable aspects of the disclosed technology. Various changes may be made, and equivalents may be substituted, without departing from the actual spirit and scope of the various embodiments. In addition, many modifications may be made to adapt a particular situation, material, composition of matter, process, process act(s), or step(s) to the purpose(s), spirit, or scope of the various embodiments. Moreover, as will be understood by those skilled in the art, each of the individual variations described and illustrated herein has individual components and features that can be readily separated or combined with the features of any of the other several embodiments without departing from the scope or spirit of the various embodiments. All such modifications are intended to be within the scope of this disclosure and the associated claims.
[0176] Illustrative Embodiments The following embodiments are illustrative and are not intended to limit the present invention.
[0177] Embodiment 1. A method for evaluating sample-specific performance of a copy number variant classifier comprising a copy number variant model, comprising: parameterizing the copy number variant model based on the number of actual sequence reads from the test sample that mapped to the segment within the region of interest to determine one or more copy number variant model parameters; generating a plurality of synthetic copy number variants, each synthetic copy number variant comprising a synthetic copy number of one or more of the segments, each synthetic copy number represented by a synthetic sequence read number based on the actual sequence read number of the corresponding segment from the test sample; calling copy numbers of one or more segments of synthetic copy number variants using the copy number variant model and the one or more determined copy number variant model parameters; determining sample-specific performance statistics for the copy number variant callers based on the difference between the called copy number and the synthetic copy number of the synthetic copy number variant; and evaluating sample-specific performance of the copy number variant callers based on sample-specific performance statistics.
[0178] Embodiment 2. The method of embodiment 1, wherein synthetic sequence read numbers for one or more segments are generated by increasing, decreasing, or maintaining the actual sequence read number for the corresponding segment from the test sample in proportion to the predetermined copy number of the one or more segments.
[0179] Embodiment 3. The method of embodiment 2, wherein the predetermined copy number is an integer copy number.
[0180] Embodiment 4. The method of embodiment 2, wherein the predetermined copy number is a non-integer copy number.
[0181] Embodiment 5. The method of any one of embodiments 1 to 4, wherein the synthetic sequence read numbers are generated by sampling a binomial distribution with a probability of success equal to m / x and a number of trials equal to the actual sequence read number at the corresponding segment from the test sample, where m is the synthetic copy number of the segment in the synthetic copy number variant and x is the hypothesized copy number of the corresponding segment from the test sample.
[0182] Embodiment 6. The number of synthetic sequence reads is sampling sequence read numbers as a negative binomial distribution with a probability of success equal to m / x and a number of successes equal to the actual number of sequence reads at the corresponding segment from the test sample, where m is the synthetic copy number of the segment in the synthetic copy number variant and x is the hypothesized copy number of the corresponding segment from the test sample; 6. The method of any one of embodiments 1 to 5, wherein the number of sampled sequence reads is generated by adding the number of actual sequence reads of the corresponding segment from the test sample.
[0183] Embodiment 7. The method of embodiment 6, wherein the synthetic sequence read numbers are generated by sampling the sequence read numbers as the expectation value of a negative binomial distribution.
[0184] Embodiment 8. The method of any one of embodiments 1 to 7, wherein the copy number variant model is a hidden Markov model.
[0185] Embodiment 9. The hidden Markov model is (i) one or more hidden states containing copy numbers corresponding to the queried segment or multiple subsegments within the queried segment; (ii) an observation state including the actual or synthetic sequence read counts for the queried segment; (iii) a copy number likelihood model based on expected actual or synthetic sequence read numbers for the queried segment.
[0186] Embodiment 10. The method of embodiment 9, comprising determining a copy number likelihood model.
[0187] Embodiment 11. Parameterizing a Hidden Markov Model to Parameterize a Copy Number Likelihood Model 11. The method of embodiment 9 or 10, comprising adjusting the number of sequence reads from the test sample to fit the actual number of sequence reads mapped to the queried segment.
[0188] Embodiment 12. The method of any one of embodiments 9 to 11, wherein the copy number likelihood model comprises a distribution of two or more copy number states.
[0189] Embodiment 13. The method of any one of embodiments 9 to 12, wherein the copy number likelihood model comprises a negative binomial distribution, and the negative binomial distribution is not a Poisson distribution.
[0190] Embodiment 14. The method of any one of embodiments 9 to 13, wherein the expected actual or synthetic number of sequence reads is based on the number of representative mapped sequence reads at the segments corresponding to the queried segments across multiple samples and the number of representative mapped sequence reads at the segments in the test sample, and the number of representative mapped sequence reads at the segments corresponding to the queried segments across multiple samples or the number of representative mapped sequence reads at the segments in the test sample are normalized representative values.
[0191] Embodiment 15. The method of any one of embodiments 9 to 14, wherein the copy number likelihood model is adjusted to account for the presence of GC content bias.
[0192] Embodiment 16. The method of any one of embodiments 9 to 15, wherein the hidden Markov model comprises transition probabilities of the copy number of the queried segment given the copy number of the spatially adjacent segment.
[0193] Embodiment 17. The method of any one of embodiments 9 to 15, wherein the hidden Markov model includes multiple transition probabilities of subsegment copy numbers at multiple subsegments within the queried segment for a given copy number of spatially adjacent subsegments.
[0194] Embodiment 18. The method of embodiment 16 or 17, wherein the transition probabilities take into account the length of the representative value of the copy number variant.
[0195] Embodiment 19. The method of any one of embodiments 16 to 18, wherein the transition probabilities take into account the prior probabilities of copy number variants in the queried segment or spatially adjacent segments.
[0196] Embodiment 20. The method of embodiment 18 or 19, wherein the representative length of the copy number variant or the probability of the copy number variant at the interrogated segment is determined based on observations in a human population.
[0197] Embodiment 21 The method of any one of embodiments 1 to 20, wherein parameterizing the copy number variant model comprises considering one or more spurious capture probes.
[0198] Embodiment 22. The method of embodiment 21, wherein considering one or more spurious capture probes includes weighting one or more observation states of the plurality of observation states with a spurious capture probe indicator.
[0199] Embodiment 23. The method of embodiment 22, wherein the spurious capture probe indicators are determined using a Bernoulli process.
[0200] Embodiment 24. The method of embodiment 22 or 23, wherein considering one or more of the capture probes to be spurious comprises using expectation maximization.
[0201] Embodiment 25. The method of any one of embodiments 21 to 24, wherein if a capture probe is determined to be spurious, sequence reads derived from that capture probe are discarded in the copy number variant model.
[0202] Embodiment 26 The method of any one of embodiments 1 to 25, wherein parameterizing the copy number variant model comprises taking into account noise in the number of mapped sequence reads.
[0203] Embodiment 27. The method of any one of embodiments 1 to 26, wherein the copy number variant model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more copy number variant model parameters.
[0204] Embodiment 28. The method of any one of embodiments 1 to 27, wherein the copy number variant model is parameterized by solving a trust-region Newton conjugate gradient algorithm.
[0205] Embodiment 29. The method of any one of embodiments 1 to 28, wherein the copy number variant model is iteratively parameterized using expectation maximization.
[0206] Embodiment 30. The method of any one of embodiments 1 to 29, comprising mapping actual sequence reads from the test sample to segments within the region of interest and determining the number of actual sequence reads mapped to the segments.
[0207] Embodiment 31. The method of any one of embodiments 1 to 30, wherein the test sample is enriched using one or more direct target sequence capture probes.
[0208] Embodiment 32. The method of any one of embodiments 1 to 31, comprising calling the copy number of one or more segments for the test sample.
[0209] Embodiment 33. The method of any one of embodiments 1 to 32, wherein the segments comprise spatially adjacent segments.
[0210] Embodiment 34. The method of any one of embodiments 1 to 33, wherein the sample-specific performance statistic is a limit of detection, sensitivity, specificity, precision, recall, accuracy, positive predictive value, or negative predictive value.
[0211] Embodiment 35. The method of any one of embodiments 1 to 34, wherein the sample-specific performance statistic is sensitivity or precision.
[0212] Embodiment 36 The method of any one of embodiments 1 to 35, comprising rejecting the test sample if the sample-specific performance of the copy number variant model is below a desired performance threshold.
[0213] Embodiment 37. A method for determining the copy number of an interrogated segment within a region of interest, comprising: (a) Multiple sequence reads generated from a test sequence library are analyzed using the queried sequence library. mapping the test sequence library to a target segment, wherein the test sequence library is enriched using one or more direct target sequence capture probes; (b) determining the number of sequence reads that map to the queried segment; (c) determining a copy number likelihood model based on the expected number of sequence reads mapped to the queried segment; and (d) a hidden Markov model, (i) one or more hidden states containing copy numbers corresponding to the queried segment or multiple subsegments within the queried segment; (ii) an observation state including the number of sequence reads that mapped to the queried segment; (iii) constructing a hidden Markov model including a copy number likelihood model; (e) parameterizing a hidden Markov model by fitting a copy number likelihood model to the determined number of sequence reads mapped to the queried segment, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the queried segment based on a parameterized hidden Markov model.
[0214] Embodiment 38. A method for determining the copy number of an interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, wherein the plurality of spatially adjacent segments includes an interrogated segment, and the test sequence library is enriched using a plurality of spatially adjacent direct target sequence capture probes; (b) determining the number of sequence reads that map to each spatially adjacent segment; (c) determining a copy number likelihood model for each spatially adjacent segment based on the expected number of mapped sequence reads in the spatially adjacent segment; (d) a hidden Markov model, (i) a plurality of hidden states including copy numbers of each of the spatially adjacent segments or a plurality of subsegments within each of the spatially adjacent segments; (ii) a plurality of observation states including the number of sequence reads mapped to each spatially adjacent segment; (iii) constructing a hidden Markov model including a copy number likelihood model for each spatially adjacent segment; (e) parameterizing a hidden Markov model, including adjusting each copy number likelihood model to fit the determined number of sequence reads mapped to each spatially adjacent segment, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the queried segment based on a parameterized hidden Markov model.
[0215] Embodiment 39. One or more parameters of the copy number likelihood model are the variance of the mapped sequence read counts of a segment (d i ), the average number of mapped sequence reads in the segment (μ i ), the variance of the number of mapped sequence reads of the segments in the test sequence library (d j ), or the average number of mapped sequence reads for a segment in the test sequence library (μ j 39. The method of embodiment 37 or 38, comprising:
[0216] Embodiment 40. The method of any one of embodiments 37 to 39, further comprising determining the most probable copy number of a section within the region of interest, wherein the section comprises a plurality of spatially adjacent segments including the queried segment.
[0217] Embodiment 41 The method of any one of embodiments 37 to 40, wherein the copy number likelihood model comprises a distribution of two or more copy number states.
[0218] Embodiment 42 The method of any one of embodiments 37 to 41, wherein the copy number likelihood model comprises a negative binomial distribution, and the negative binomial distribution is not a Poisson distribution.
[0219] Embodiment 43. A method according to any one of embodiments 37 to 42, wherein the expected number of sequence reads is based on the average number of mapped sequence reads at corresponding segments across multiple sequence libraries and the average number of mapped sequence reads at multiple segments of interest in the test sequence library, and the average number of mapped sequence reads at corresponding segments across multiple sequence libraries or the average number of mapped sequence reads at multiple segments of interest in the test sequence library are normalized representative values.
[0220] Embodiment 44 The method of any one of embodiments 37 to 43, wherein the copy number likelihood model is adjusted to account for the presence of GC content bias.
[0221] Embodiment 45. The method of embodiment 44, wherein the adjustment depends on the GC content of the capture probe corresponding to the interrogated segment or the GC content of the interrogated segment.
[0222] Embodiment 46. The method of any one of embodiments 37 to 45, wherein the hidden Markov model comprises transition probabilities of the copy number of the queried segment given the copy number of the spatially adjacent segment.
[0223] Embodiment 47. The method of any one of embodiments 37 to 45, wherein the hidden Markov model includes multiple transition probabilities of subsegment copy numbers at multiple subsegments within the queried segment for a given copy number of spatially adjacent subsegments.
[0224] Embodiment 48. The method of embodiment 46 or 47, wherein the transition probabilities take into account the length of the representative value of the copy number variant.
[0225] Embodiment 49. The method of any one of embodiments 46 to 48, wherein the transition probabilities take into account the prior probabilities of copy number variants in the queried segment or spatially adjacent segments.
[0226] Embodiment 50. The method of embodiment 48 or 49, wherein the representative length of the copy number variant or the probability of the copy number variant at the interrogated segment is determined based on observations in a human population.
[0227] Embodiment 51. The method of any one of embodiments 37 to 50, wherein parameterizing the hidden Markov model includes taking into account one or more spurious capture probes.
[0228] Embodiment 52. Considering one or more spurious capture probes includes weighting one or more observation states of a plurality of observation states with a spurious capture probe indicator. 52. The method of embodiment 51.
[0229] Embodiment 53. The method of embodiment 52, wherein the spurious capture probe indicators are determined using a Bernoulli process.
[0230] Embodiment 54. The method of embodiment 52 or 53, wherein considering one or more of the capture probes to be spurious comprises using expectation maximization.
[0231] Embodiment 55. The method of any one of embodiments 52 to 54, wherein if a capture probe is determined to be spurious, the likelihood information from that capture probe is discarded in the copy number likelihood model.
[0232] Embodiment 56. The method of any one of embodiments 37 to 55, wherein parameterizing the hidden Markov model includes taking into account noise in the number of mapped sequence reads.
[0233] Embodiment 57. The method of any one of embodiments 37 to 56, wherein accounting for noise in the number of mapped sequencing reads comprises adjusting a copy number likelihood model.
[0234] Embodiment 58. The method of embodiment 57, wherein adjusting the copy number likelihood model to account for noise comprises an expectation maximization step.
[0235] Embodiment 59. The method of embodiment 58, wherein the expectation maximization step includes weighting the level of noise in the number of mapped sequence reads from the test sequence library.
[0236] Embodiment 60. The method of any one of embodiments 56 to 59, wherein the most probable copy number of the queried segment is not called if the noise in the number of mapped sequence reads is above a predetermined threshold.
[0237] Embodiment 61 The method of any one of embodiments 37 to 60, wherein sequence reads from overlapping capture probes are merged.
[0238] Embodiment 62. The method of any one of embodiments 37 to 61, wherein the Viterbi algorithm, a quasi-Newton solver, or Markov Chain Monte Carlo is used to determine the most probable copy number of the queried segment.
[0239] Embodiment 63 The method of any one of embodiments 37 to 62, further comprising determining the confidence of the most probable copy number of the segment.
[0240] Embodiment 64. A method for determining copy number variant abnormalities in a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to interrogated segments within a region of interest, wherein the test sequence library is enriched using one or more direct target sequence capture probes; (b) determining the number of sequence reads that map to the queried segment; (c) determining a copy number likelihood model based on the expected number of sequence reads mapped to the queried segment; and (d) a hidden Markov model, (i) the queried segment or multiple segments within the queried segment one or more hidden states containing copy numbers corresponding to the subsegments; (ii) an observation state including the number of sequence reads that mapped to the queried segment; (iii) constructing a hidden Markov model including a copy number likelihood model; (e) parameterizing a hidden Markov model by fitting a copy number likelihood model to the determined number of sequence reads mapped to the queried segment, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the queried segment based on a parameterized hidden Markov model; and (g) determining copy number variant abnormalities based on the most probable copy number of the queried segment.
[0241] Embodiment 65. A method for determining copy number variant abnormalities in a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, wherein the plurality of spatially adjacent segments includes an interrogated segment, and the test sequence library is enriched using a plurality of spatially adjacent direct target sequence capture probes; (b) determining the number of sequence reads that map to each spatially adjacent segment; (c) determining a copy number likelihood model for each spatially adjacent segment based on the expected number of mapped sequence reads in the spatially adjacent segment; (d) a hidden Markov model, (i) a plurality of hidden states including copy numbers of each of the spatially adjacent segments or a plurality of subsegments within each of the spatially adjacent segments; (ii) a plurality of observation states including the number of sequence reads mapped to each spatially adjacent segment; (iii) constructing a hidden Markov model including a copy number likelihood model for each spatially adjacent segment; (e) parameterizing a hidden Markov model, including adjusting each copy number likelihood model to fit the determined number of sequence reads mapped to each spatially adjacent segment, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the queried segment based on a parameterized hidden Markov model; and (g) determining copy number variant abnormalities based on the most probable copy number of the queried segment.
[0242] Embodiment 66. A method for determining the copy number of an interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to interrogated segments, wherein the test sequence library is enriched using one or more capture probes; (b) determining the number of sequence reads that map to the queried segment; (c) determining a copy number likelihood model based on the expected number of sequence reads mapped to the queried segment; and (d) a hidden Markov model, (i) one or more hidden states containing copy numbers corresponding to the queried segment or multiple subsegments within the queried segment; (ii) an observation state including the number of sequence reads that mapped to the queried segment; (iii) constructing a hidden Markov model including a copy number likelihood model; (e) parameterizing a hidden Markov model by adjusting the copy number likelihood model to fit the determined number of sequence reads mapped to the queried segment and accounting for one or more spurious capture probes, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the queried segment based on a parameterized hidden Markov model.
[0243] Embodiment 67. A method for determining the copy number of an interrogated segment within a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, wherein the plurality of spatially adjacent segments includes an interrogated segment, and the test sequence library is enriched using a plurality of spatially adjacent direct target sequence capture probes; (b) determining the number of sequence reads that map to each spatially adjacent segment; (c) determining a copy number likelihood model for each spatially adjacent segment based on the expected number of mapped sequence reads in the spatially adjacent segment; (d) a hidden Markov model, (i) a plurality of hidden states including copy numbers of each of the spatially adjacent segments or a plurality of subsegments within each of the spatially adjacent segments; (ii) a plurality of observation states including the number of sequence reads mapped to each spatially adjacent segment; (iii) constructing a hidden Markov model including a copy number likelihood model for each spatially adjacent segment; (e) parameterizing a hidden Markov model, including adjusting each copy number likelihood model to fit the determined number of sequence reads mapped to each spatially adjacent segment and accounting for one or more spurious capture probes, wherein the hidden Markov model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the queried segment based on a parameterized hidden Markov model.
[0244] Embodiment 68. One or more parameters of the copy number likelihood model are determined by the variance of the number of mapped sequence reads of a segment (d i ), the average number of mapped sequence reads in the segment (μ i ), the variance of the number of mapped sequence reads of the segments in the test sequence library (d j ), or the average number of mapped sequence reads for a segment in the test sequence library (μ j 68. The method of any one of embodiments 64 to 67, comprising:
[0245] Embodiment 69. The method of any one of embodiments 37 to 68, wherein the analytical first derivative gradient and second derivative analytical Hessian of one or more parameters in the copy number likelihood model are solved using a trust-region Newton conjugate gradient algorithm.
[0246] Embodiment 70. A computer system including a computer-readable medium containing instructions for carrying out the method according to any one of embodiments 1 to 68. [Example]
[0247] Biological samples from blood or saliva were sequenced using the Illumina HiSeq2500 platform after direct target sequence enrichment across a panel of 178 genes. Batches of 46 samples were analyzed, with each batch containing different ratios of saliva and blood samples. Saliva samples generally produce noisier sequence results, which can affect the sensitivity of other samples within the same flow cell batch. The number of sequence reads from segments within each sample was used to parameterize a segment hidden Markov model, generating 400 synthetic copy number variants. The hidden Markov model call was used to call the copy number of the segment within each sample synthetic copy number variant. The hidden Markov model included (i) a hidden state for the copy number of a given segment, (ii) an observed state with the number of synthetic sequence reads for a given segment, and (iii) a copy number likelihood model based on the number of synthetic reads for a given segment. The sensitivity of each test sample was determined using the called copy number of the segment in the synthetic variant and the actual copy number in the synthetic variant.
[0248] Copy number variant calling analysis was performed using two different hidden Markov model callers. In the reference hidden Markov model caller, sample noise (i.e., variance due to noise in the test sequence library) and spurious capture probe noise were ignored. In the test hidden Markov model, noise in the test sequence library was ignored.
number
[0249] The determined sensitivity for each sample is shown in Figure 12, plotted against the number of saliva samples (46, the remainder being blood samples). The sensitivity using the reference hidden Markov collaborator generally worsens when there are more saliva samples in the batch. However, even when the batch contains 44 saliva samples, the sensitivity of the test hidden Markov collaborator generally remains above 90%.
Claims
1. 1. A method for determining copy number variant abnormalities in a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to interrogated segments within the region of interest, the test sequence library being enriched using one or more direct target sequence capture probes; (b) determining the number of sequence reads that map to the queried segment; (c) determining a copy number likelihood model based on the expected number of sequence reads mapped to the queried segment; (d) A hidden Markov model, (i) one or more hidden states comprising copy numbers corresponding to the interrogated segment or multiple subsegments within the interrogated segment; (ii) an observation state including the number of sequence reads that mapped to the queried segment; and (iii) a transition probability from the copy number of the queried segment to the copy number of a segment spatially adjacent to the queried segment; and (iv) the copy number likelihood model; constructing a Hidden Markov Model, including: (e) parameterizing the Hidden Markov Model to adjust the copy number likelihood model to fit the determined number of sequence reads mapped to the queried segment, wherein the Hidden Markov Model is parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood model; (f) determining the most probable copy number of the queried segment based on the parameterized hidden Markov model; and (g) determining copy number variant abnormalities based on the most probable copy number of the queried segment; A method comprising:
2. 1. A method for determining copy number variant abnormalities in a region of interest, comprising: (a) mapping a plurality of sequence reads generated from a test sequence library to a plurality of spatially adjacent segments, the plurality of spatially adjacent segments including an interrogated segment, the test sequence library being enriched using a plurality of spatially adjacent direct target sequence capture probes; (b) determining the number of sequence reads that map to each spatially adjacent segment; (c) determining a copy number likelihood model for each spatially adjacent segment based on the expected number of sequence reads that map in each spatially adjacent segment; (d) A hidden Markov model, (i) a plurality of hidden states including copy numbers of each spatially adjacent segment or a plurality of subsegments within each spatially adjacent segment; (ii) a plurality of observation states including the number of sequence reads mapped to each of the spatially adjacent segments; (iii) a transition probability from the copy number of the queried segment to the copy number of a segment spatially adjacent to the queried segment; and (iv) the copy number likelihood model for each spatially adjacent segment; and constructing a Hidden Markov Model, including: (e) parameterizing the hidden Markov models, including adjusting each copy number likelihood model to fit the determined number of sequence reads mapped to each spatially adjacent segment, wherein the hidden Markov models are parameterized using analytical first derivative gradients and second derivative Hessians of one or more parameters in the copy number likelihood models; (f) determining the most probable copy number of the queried segment based on the parameterized hidden Markov model; and (g) determining copy number variant abnormalities based on the most probable copy number of the queried segment; A method comprising:
3. 3. The method of claim 1 or 2, wherein the one or more parameters of the copy number likelihood model comprise the variance (di) of the number of mapped sequence reads of the interrogated segment, the number of mapped sequence reads of a representative value (μi) of the interrogated segment, the variance (dj) of the number of mapped sequence reads of all segments in the test sequence library, or the number of mapped sequence reads of a representative value (μj) of all segments in the test sequence library.
4. 4. The method of claim 1, further comprising determining a most probable copy number of a section within the region of interest, the section comprising a plurality of spatially adjacent segments including the interrogated segment; the copy number likelihood model comprising a distribution of two or more copy number states; the copy number likelihood model comprising a negative binomial distribution, wherein the negative binomial distribution is not a Poisson distribution; the expected sequence read counts are based on a representative number of mapped sequence reads at corresponding segments across a plurality of sequence libraries and a representative number of mapped sequence reads across a plurality of segments of interest in the test sequence library; the representative number of mapped sequence reads at corresponding segments across a plurality of sequence libraries or the representative number of mapped sequence reads across a plurality of segments of interest in the test sequence library are normalized representative values; and the copy number likelihood model is adjusted based on the presence of a GC content bias, wherein the adjustment depends on the GC content of the direct target sequence capture probe corresponding to the interrogated segment or the GC content of the interrogated segment.
5. A method described in any of claims 1 to 4, wherein the transition probability takes into account the representative length of copy number variants or the prior probability of copy number variants in the queried segment or spatially adjacent segments, and the representative length of copy number variants is determined based on observations in a human population.
6. 6. The method of claim 1, wherein a Viterbi algorithm, a quasi-Newton solver, or Markov Chain Monte Carlo is used to determine the most probable copy number of the queried segment.
7. 7. The method of claim 1, further comprising determining the confidence level of the most probable copy number of the segment.
Citation Information
Patent Citations
Methods and systems for copy number variant detection
WO2016187051A1
Genetic variant-phenotype analysis system and methods of use
WO2017172958A1
Methods for assessing genetic variant screen performance
WO2018085779A1