Screening of cancer-related genetic variants using mutation characteristics
By screening cancer-related mutation signatures and creating tumor-informed assays, the problem of low detection sensitivity in existing technologies is addressed, enabling efficient and accurate detection of residual cancer cells.
Patent Information
- Application Number
- CN202480011480.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-18
- Filing Date
- 2024-01-17
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies have the problem of low sensitivity when detecting residual cancer cells, especially due to interference from genetic variants caused by artificial mutations, which affects the accuracy of detection.
By analyzing cancer-related mutational signatures, we screen for high-quality genetic variant sets, use mutational signature analysis methods to identify genetic variants associated with cancer types, create tumor-informed assays, exclude sources of human error and therapy-related variants, and improve detection sensitivity.
It improves the sensitivity of detecting residual cancer cells, reduces the interference of artificial mutations, enhances the accuracy of detection, and can more effectively monitor the presence of cancer in patients.
Smart Images

Figure CN120660139A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the use of mutational signatures to screen for cancer-associated genetic variants. Background Art
[0002] After cancer treatment, a small number of cancer cells may remain in patients who appear to be in remission. These residual cells are called "minimal residual disease" (MRD) and can be the cause of relapse. Assays for detecting MRD (e.g., circulating tumor DNA (ctDNA) assays) can employ a variety of methods, including sequencing a patient's tumor tissue to identify tumor-informed genetic variants that, when detected in a patient's cell-free DNA (cfDNA), can indicate MRD. Summary of the Invention
[0003] In some embodiments, data analysis techniques are provided for selecting sets of genetic variants for use in creating tumor-informed assays. Information about genetic variants in mutational signatures of different cancer types can be used, at least in part, to screen sets of genetic variants observed in a sample's mutational catalog to identify genetic variants that are more likely to be attributable to a sample-specific mutational process than to other processes, including artifacts resulting from sample manipulation. The identified genetic variants can then be used to create tumor-informed assays.
[0004] In some embodiments, a method for selecting genetic variants for use in a tumor-informed assay is provided. The method includes receiving a sample collected from a patient, the sample associated with a cancer type, generating a mutation catalog for the sample, the mutation catalog indicating proportions of genetic mutation types observed in the sample, selecting a feature set associated with the cancer type, the set comprising one or more features, each feature comprising a mutation profile, determining a set of genetic variants that are most likely true somatic variants associated with the sample based on the feature set associated with the cancer type and the mutation catalog, and outputting the set of genetic variants for use in creating a tumor-informed assay for the patient.
[0005] In one aspect, selecting a feature set associated with a cancer type includes accessing a database configured to store a plurality of features associated with the cancer type, and the plurality of features are population-level features determined from a plurality of cancer samples. In another aspect, the method further includes including in the feature set only features from the database having one or more genetic mutation types representing at least 5% of all genetic mutations in the mutation spectrum. In another aspect, the method further includes including in the feature set only features from the database having one or more genetic mutation types representing at least 10% of all genetic mutations in the mutation spectrum. In another aspect, the method further includes including in the feature set only features from the database having one or more genetic mutation types representing at least 25% of all genetic mutations in the mutation spectrum.
[0006] In another aspect, determining the set of genetic variants comprises fitting the mutation profiles in the feature set to the mutation catalog to determine the corresponding amount of each feature observed in the sample. In another aspect, determining the set of genetic variants further comprises determining whether the amount of each feature observed in the sample is greater than a threshold. In another aspect, determining the set of genetic variants further comprises, for each of the genetic variants in the sample, associating a context probability based on the frequency of the genetic variant in the feature set and the determined amount of each feature observed in the sample, and sampling the genetic variants weighted by their associated context probabilities to determine the genetic variants to include in the set of genetic variants. In another aspect, the context probabilities are based on trinucleotide contexts.
[0007] In another aspect, the method further includes selecting the set of genetic variants for use in a tumor-informed assay based at least in part on the suitability of each genetic variant in the set of genetic variants, and creating a tumor-informed assay based on the set of genetic variants from the selected set. In another aspect, the feature set associated with the cancer type includes at least two features. In another aspect, selecting the feature set associated with the cancer type further includes selecting a mutation profile associated with an error source introduced during the sequencing process, and the method further includes excluding from the set of genetic variants any genetic variants determined to be attributable to the error source. In another aspect, the error source is associated with a formalin-fixed paraffin-embedded (FFPE) process, an amplification process, or a sequencing process.
[0008] In another aspect, selecting the feature set associated with the cancer type further comprises selecting a mutation profile associated with a therapy, and the method further comprises excluding from the set of genetic variants any genetic variants determined to be attributable to the therapy. In another aspect, the therapy is chemotherapy.
[0009] In another aspect, generating a mutation catalog for a sample comprises performing whole exome sequencing on the sample. In another aspect, the signature set comprises at least one two-base substitution signature. In another aspect, at least some of the mutation profiles are associated with different exposure types. In another aspect, the mutation profiles indicate proportions of genetic mutation types observed in a population. In another aspect, the genetic mutation types comprise trinucleotide contexts. In another aspect, the signature represents exposure to a mutational process. In another aspect, the mutational process is associated with cancer.
[0010] In some embodiments, a method for selecting genetic variants is provided. The method includes receiving a sample collected from a patient, generating a mutation catalog from a genetic variant set (the mutation catalog indicating the proportions of genetic mutation types observed in the sample), selecting a feature set (the set including one or more features, each feature including a mutation profile), determining a genetic variant set, excluding from the genetic variant set any genetic variants determined to be attributable to the one or more features, and outputting the genetic variant set for use in creating an assay for the patient.
[0011] In one aspect, the feature set is associated with an error source. In another aspect, the error source is associated with a formalin-fixed paraffin-embedded (FFPE) process, an amplification process, or a sequencing process. In another aspect, the feature set is associated with a therapy. In another aspect, the therapy is chemotherapy. In another aspect, excluding further comprises excluding from the set of genetic variants any genetic variants determined to be subclonal genetic variants associated with the therapy. In another aspect, the feature set further comprises at least one mutation profile associated with a cancer type, wherein determining further comprises determining the set of genetic variants associated with the sample based on the at least one mutation profile associated with the cancer type and the mutation catalog.
[0012] In some embodiments, a system is provided. The system includes at least one hardware computer processor programmed to perform any of the methods described herein.
[0013] In some embodiments, a computer-readable medium is provided. The computer-readable medium is encoded with a plurality of instructions that, when executed by at least one hardware computer processor, perform any of the methods described herein.
[0014] In some embodiments, a tumor-informed assay for monitoring the presence of cancer in a patient is provided. The tumor-informed assay is configured to detect at least one genetic variant from a set of genetic variants output according to any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Various non-limiting embodiments of the present technology will be described with reference to the following drawings.It should be understood that the drawings are not necessarily drawn to scale.
[0016] Figure 1 is an example representation of a mutation catalog of a tumor sample that can be used to identify genetic variants for tumor-informed assays, according to some embodiments of the present disclosure.
[0017] Figure 2A is an example representation of a single base substitution mutation signature that can be used in some embodiments of the present disclosure to identify sets of genetic variants for tumor-informed assays.
[0018] Figure 2Bis an example representation of a dibase substitution mutation signature that can be used in some embodiments of the present disclosure to identify sets of genetic variants for tumor-informed assays.
[0019] Figure 3 is a flow chart of a process for obtaining a set of genetic variants for use in creating a tumor-informed assay, according to some embodiments of the present disclosure.
[0020] Figure 4 is a flow chart of a process for obtaining a set of genetic variants for use in creating a tumor-informed assay, according to some embodiments of the present disclosure.
[0021] Figure 5A is an example representation of a flat mutation signature that may be present in a database of mutation signatures according to some embodiments of the present disclosure.
[0022] Figure 5B is an example representation of prominent mutation features that may be selected for use in the fitting process according to some embodiments of the present disclosure.
[0023] Figure 6 is an example representation of selecting multiple single-base substitution mutation features for use in a fitting process according to some embodiments of the present disclosure.
[0024] Figure 7 is an example representation of selecting single-base substitution mutation signatures and double-base substitution mutation signatures for use in the fitting process according to some embodiments of the present disclosure.
[0025] Figure 8 is an example of a process for determining whether a genetic mutation type is similar to one or more known mutation signatures according to some embodiments of the present disclosure.
[0026] Figure 9 is an example representation of a mutational signature of a FFPE process that can be used to screen for artifactual variants according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0027] Aspects of the technology described herein relate to techniques for determining sets of genetic variants to be included in a patient's tumor-specific assay (e.g., an MRD assay). The inventors have recognized that at least some of the genetic variants identified by sequencing of patient tumor tissue may be artifacts caused by processes such as cytosine deamination, early incorporation errors, and polymerase bias from PCR enrichment, as well as incorporation errors from sequencing. Including such genetic variants in a patient's tumor-informed assay can reduce the sensitivity of the assay because they are not present in the patient's cancer. Some embodiments of the present disclosure relate to techniques for identifying sets of genetic variants for use in tumor-informed assays that improve the sensitivity of assays for detecting, for example, residual cancer cells.
[0028] Some embodiments of the present disclosure relate to a tumor-informed variant screening method that applies mutation signature analysis to select high-quality genetic variants that are most likely to be present in a tumor or cancer type. Genetic mutations occur in human DNA through various mutational processes, including but not limited to inherent errors in DNA replication, and exposure to various physical, chemical or biological mutagens. Different mutational processes can generate specific combinations of genetic mutation types, so that unique mutational signatures can be associated with the mutational processes. The existence of such mutational signatures has recently been investigated for various mutational processes associated with cancer and published to publicly accessible databases (e.g., the COSMIC database, available at https: / / cancer.sanger.ac.uk / cosmic).
[0029] The inventors have recognized and appreciated that mutational signatures associated with different forms of cancer can be used to screen for genetic variants present in samples from patients. For example, a screened set of genetic variants can be used for tumor-informed assays for patients. As described in more detail below, some embodiments attempt to fit genetic variants in a mutational catalog determined for a tumor sample from a patient to multiple mutational signatures associated with the patient's cancer. Genetic variants associated with features related to the tumor sample (e.g., variants that fit the expected cancer type) can be prioritized for inclusion in tumor-informed assays to improve the sensitivity of the assay in detecting cancer in future samples collected from the patient. In addition, by prioritizing genetic variants based on association with known cancer mutational signatures, additional normal or healthy patient samples may not be required for comparison with tumor tissue to identify tumor-associated variants. The technology described herein may also have applications in quality control, for example, by rejecting samples that appear to be from a different type of cancer (indicating potential processing issues) or excluding genetic variants that appear to be associated with an ex vivo mutational process that are unlikely to be associated with the patient's cancer, as described in more detail below.
[0030] A mutation catalog is a mutation spectrum that includes a set of genetic mutation types observed in a single biological sample (such as a tumor sample). In some embodiments, the mutation catalog can include a context, such as a trinucleotide context, which identifies different combinations of nucleotides located directly at the 5' (e.g., preceding) and 3' (e.g., following) positions of a given mutation in a DNA sequence. Figure 1An example of a mutation catalog generated for a tumor sample according to some embodiments of the present disclosure is illustrated. As shown, when considering six mutation types (C>A, C>G, C>T, T>A, T>C, T>G) and 16 trinucleotide contexts for each mutation type, the mutation catalog provides information for 96 "channels". In some embodiments, the mutation catalog of a tumor sample can be determined by identifying multiple genetic mutations in the tumor sample relative to a reference sequence and grouping the mutations according to their context. Figure 1 As shown in , a mutation catalog can be represented as a histogram, where the height of the bar in the histogram can correspond to the mutation count (e.g. Figure 1 ) or proportions (e.g., percentage of single-base substitutions). Thus, a tumor sample's mutation catalog can indicate whether certain mutation types are overrepresented or underrepresented in the sample, providing insight into the most prevalent genetic mutation types in that tumor sample.
[0031] As described above, mutational signatures of various mutational processes (e.g., mutational processes associated with cancer) have been studied and are publicly accessible in various databases (e.g., via the COSMIC database, which is available at https: / / cancer.sanger.ac.uk / cosmic). Mutational signatures can be generated from large sample groups from different individuals. For example, a mutational signature can represent a recurring pattern of sequence context-dependent genetic variants observed in patients with similar etiologies. Some unique mutational signatures associated with different cancer types and causes (e.g., exposure to different mutagens) have been identified. Mutational signatures associated with different cancer types can be generated based on whole genome sequencing (WGS) of cancer tissue from thousands of samples, where the results are clustered to produce different mutational signatures for each cancer type. Each mutational signature can represent the landscape of both passenger and driver genetic mutations associated with the cancer. Mutational signatures can be associated with a disease (e.g., lung cancer, colon cancer, or skin cancer) and / or a potential cause of the disease (e.g., oxidation, microsatellite instability, exposure to ultraviolet light).
[0032] Similar to the mutation catalog of a sample (e.g. Figure 1 The mutation signature determined based on population data can be represented as a histogram of genetic mutations. The genetic mutation histogram of the example mutation signature (labeled as SBS4) is shown in Figure 2A The SBS4 mutational signature is associated with lung cancer derived from smoking-induced oxidation and is characterized by an excess of C>A and G>T mutations (as well as CC>AA and GG>TT double substitutions and C / G deletions, not shown). Figure 2AAs shown in , the genetic mutations in the mutation signature can be further characterized by context, such as trinucleotide context (i.e., different combinations of 5' and 3' nucleotides adjacent to any given genetic mutation). When characterized by trinucleotide context, the characterization of the genetic mutations produces a 96-category spectrum for any given sample.
[0033] although Figure 2A The example mutation signature shown in includes genetic mutations of single base substitutions, but it should be understood that mutation signatures can also include genetic mutation types other than single base substitutions, including but not limited to double base substitutions, insertions and deletions, copy number variations, and microsatellites. The genetic mutation histogram of the example mutation signature including double base substitutions (labeled as DBS1) is shown in Figure 2B As shown in . As shown, the DBS1 mutation signature associated with melanoma is mainly characterized by CC>NN mutations, especially CC>NN mutations with TT context (i.e., TCCT>TNNT). Other contexts of genetic variation can also be used. For example, additional nucleotides located at 5' and / or 3' of the variation can be used to provide contexts such as dinucleotides, trinucleotides, tetranucleotides, pentanucleotides, hexanucleotides, etc. Insertions and deletions can be further characterized by homopolymer length, number of repeat units, and microhomology length. Copy number variation can be further characterized by length (e.g., 0-100kb, 100kb-1Mb, >1Mb, 100kb-1Mb, 1Mb-10Mb, 10Mb-40Mb, >40Mb), hyperdiploid state, loss of heterozygosity, or heterozygosity. Furthermore, while classifying genetic variants by context may provide additional resolution, in some embodiments, no additional nucleotide context may be used and variants may be binned simply by their base changes (e.g., C->A, C->G, T->C, etc.).
[0034] As described above, mutation signatures can be generated from a large sample population from different individuals. For example, a mutation signature can be determined by identifying recurring patterns in multiple mutation catalogs from individual patient samples. For example, consider a dataset of a large number (e.g., 1000) of mutation catalogs generated from biopsies of lung cancer, skin cancer, and breast cancer samples. Figure 1As shown in , each mutation catalog represents the count or frequency of different genetic mutation types (e.g., single base substitution, double base substitution, insertion, deletion, etc.) associated with the trinucleotide context of, for example, genetic mutations, thereby providing information of 96 different channels (six mutation types across 16 different trinucleotide contexts) for the sample. By examining the recurrence pattern across the mutation catalog in the data set (e.g., whether a C->A base change occurs in a specific context across multiple samples), it is possible to identify mutation signatures associated with different processes and / or cancer types, wherein each mutation signature indicates the frequency of the recurrence mutation type observed in the sample set. Any suitable pattern recognition technology can be used to identify the recurrence pattern of genetic mutations in the mutation catalog data set. Such technology includes, but is not limited to, machine learning techniques, such as non-negative matrix decomposition (NMF), principal component analysis (PCA), and vector quantization (VQ). Available conventional software packages (e.g., open source MutationalPatterns R package, which is available at https: / / bioconductor.org / packages / release / bioc / html / MutationalPatterns.html) extract mutation signatures from mutation catalog large data sets.
[0035] Although users can define parameters for extracting mutation signatures from mutation catalog datasets, it is most common to extract mutation signatures for individual cancer types (e.g., lung, breast) and / or associated exposures (e.g., smoking, aging). Therefore, each mutation signature extracted in this way can represent the mutation process present in the tumor sample set. Although mutation signatures and mutation catalogs appear to convey similar information, there are significant differences. For example, each mutation signature represents the proportion of genetic mutations in the entire mutation catalog dataset from which the feature is extracted, but only includes the recurrence patterns observed in the dataset, rather than the counts or proportions of all genetic mutations in the sample, as represented in the mutation catalog. Mutation signatures can be extracted by, for example, clustering mutation catalogs from many samples and selecting only samples with recurring, common spectra. An example of such mutation signature extraction is described in Degasperi et al., Substitution mutational signatures in whole-genome-sequenced cancers in the UK population, Science 376:6591 (2022).
[0036] The inventors have recognized and appreciated that genetic mutation information in a mutation signature can be used to select sets of genetic variants for use in tumor-informed assays. Accordingly, some embodiments of the present disclosure align or "fit" a mutation catalog determined for a tumor sample to a mutation signature set, which enables determination of which mutational processes (as represented in the signature) are most closely associated with the sample.
[0037] Figure 3 Illustrated is a process 300 for determining a set of genetic variants for use in tumor-informed assays according to some embodiments of the present disclosure. In act 310, a mutation catalog for a sample from a patient is generated. For example, as described in more detail below, the sample can be processed and the DNA sequenced to generate a text-based file that describes the genetic mutations in the sample relative to a reference genetic profile. The occurrence of the genetic mutations and (optionally) their context can be formulated as a mutation catalog for the sample. The process 300 then proceeds to act 312, where a set of mutation signatures associated with a specific cancer associated with the sample is selected. For example, the COSMIC database v.3.3 from June 2022 includes approximately 130 different mutation signatures extracted from mutation catalogs of a large group of tumor samples from different cancers. In act 312, a set of mutation signatures that are most likely to be associated with the specific cancer associated with the sample analyzed in act 310 can be selected.
[0038] Then, process 300 proceeds to act 314, where the mutation catalog generated in act 310 and the mutation signature set from the set selected in act 312 are used to determine a set of genetic variants that are most likely to be true somatic variations associated with the sample. For example, as described in more detail below, the mutation catalog of the sample can be fitted to the mutation signature set to determine a weighted sum of mutation signatures representing the most likely exposures that led to the mutation catalog. As an example, the mutation catalog from a lung cancer patient may include mutations associated with aging (e.g., C->T changes due to deamination), smoking, and other mutational processes. Due to different amounts of exposure to these mutational processes, different genetic mutations associated with the signature may be present in some lung cancer samples but not in other lung cancer samples. As described herein, fitting the mutation catalog to the mutation signature set can deconvolute the mutational processes active in a particular patient's tumor. By fitting the mutation catalog to the mutation signature set, each genetic variant in the mutation catalog can be associated with a specific mutational signature, thereby gaining insight into the mutational processes associated with the evolution of the tumor sample.
[0039] After determining the set of genetic variants that characterize the sample in act 314, process 300 proceeds to act 316, where the set of genetic variants is output for use in creating a tumor-informed assay for the patient. By detecting which genetic variants are likely to be associated with cancer and not artifactual, and selecting variants that characterize the underlying mutational process associated with the sample, an assay sensitive to these genetic variants can be generated.
[0040] Figure 4 A process 400 for creating a tumor-informed assay based on a set of genetic variants selected according to the techniques described herein is illustrated. In act 410, a sample (e.g., a blood sample, a tumor biopsy, or other tissue sample) can be processed (e.g., using whole exome sequencing or another next-generation sequencing technology) to determine genetic sequence variations in the sample relative to a reference genome, wherein a text-based variant call format (VCF) file is output. Process 400 can then proceed to act 412, wherein the genetic variants identified in the VCF file can optionally be subject to one or more screening criteria to identify low-quality genetic variants, such as those with low coverage, low base or mapping quality, redundant surrounding sequences, presence in online databases (e.g., a frequency in dbSNP of >0.05 indicating that it is likely a germline variant), etc. Process 400 then proceeds to act 414, wherein a mutation catalog for the sample is generated using the genetic variants identified in the VCF file that were not screened out in act 412. For example, a matrix including the genetic variants and their counts or proportions in the mutation spectrum of a sample can be generated as a mutation catalog of the sample. An example mutation catalog of a sample is described above. Figure 1 Although screening for low-quality variants may be optional in some embodiments of the present disclosure, it may be useful to remove these variants because they may affect the feature fitting process described herein.
[0041] Then, process 400 proceeds to action 420, wherein the mutation catalog of sample is fitted to the set of one or more mutation signatures. As described above, the mutation catalog is fitted to one or more mutation signatures in an attempt to reveal the underlying mutation process experienced by the sample by mapping the variant in the sample mutation catalog to the variant in the known genetic spectrum for different mutation processes. For example, consider a mutation catalog containing 1000 mutations. After fitting in action 420, 500 mutations can be attributed to SBS4 features (smoking), 200 mutations can be attributed to SBS3 (defective homologous recombination DNA damage repair), and the remaining mutations can be associated with other features or not associated with any feature. Each attribution to a specific feature can be regarded as " exposure " to the process associated with the feature. In this way, according to some embodiments of the present disclosure, the mutation catalog of the sample from the patient is fitted to a feature set based on a population and can show that the patient has been exposed to certain mutation processes as defined by certain cancer features.
[0042] Fitting techniques attempt to determine the contribution of each mutation process to the genetic variants expressed in the mutation catalog by quantifying the presence and prevalence of each mutation signature in the mutation catalog. In some embodiments, fitting can be performed by creating two matrices--sample matrix M and feature matrix P. Sample matrix M can include 96 rows and n columns, 96 rows for each mutation / trinucleotide context combination, and n columns for each mutation catalog. Feature matrix P can include k rows and 96 columns, k rows for each mutation signature, and 96 columns for each trinucleotide context. Given these two matrices, it is possible to determine a weight matrix E comprising the k columns for each feature and the n rows for each sample so that the reconstructed tumor sample matrix R (calculated as M-(P*E)) minimizes a given error threshold e. In other words, it is possible to determine a weight matrix E so that matrix E optimally recreates M by finding the optimal exposure combination that minimizes the difference between (P*E) and M. For a given sample n, the highest weight of the corresponding row in E can be selected to understand which mutation process is most likely to cause each mutation type in the mutation catalog. It should be understood that in act 420 , other techniques including but not limited to other minimization techniques, golden search, minimal quadratic, etc. may alternatively be used to fit the mutation catalog to the mutation signature set.
[0043] The inventors have recognized and appreciated that the accuracy of the genetic variants output from the fitting process in act 420 depends largely on which mutation features are provided as input to the fitting process. Therefore, some embodiments of the present disclosure include one or more feature curation actions to select a feature set from a feature database for use in the fitting act 420. Figure 4As shown in , the feature curation action can include an action 416 of selecting features that are most commonly or otherwise expected to be observed in a particular cancer type associated with the sample. For example, the inventors have recognized that, due to the sensitivity of the fitting algorithm, fitting a mutation catalog to a mutational signature associated with a mutational process that is unrelated to, for example, the particular cancer type associated with the sample can implicate mutational processes in the fitting results that are unlikely to have occurred in a particular patient. In some cases, the fitting algorithm can force certain variants to be fitted to a feature even if the sample has not been exposed to the mutational process. Thus, by preselecting certain features that are expected to be observed in such samples, the accuracy of the genetic variants output from the fitting process in action 420 can be improved.
[0044] In one example, a variety of tumor types were found Figure 2A In the embodiment of the present invention, " smoking signature " SBS4 shown in , even when tobacco carcinogens are unlikely to arrive at this site (such as in prostate cancer) is also like this. This can be because SBS4 feature is dominated by the C>A and G>T mutations found in many cancers. Therefore, if SBS4 feature is included in the mutation signature for analyzing prostate cancer sample, then due to this strong phenotype, the fitting process in action 420 can preferentially be fitted to SBS4 rather than other features, even if the patient may never smoke. Similarly, the result produced in the fitting process with the mutation signature associated with the common mutation process in the sample across all types can be less meaningful than the result produced with the mutation signature associated with the uncommon more specific process in the sample across all or many types. For example, the SBS1 mutation signature (relevant to aging, and therefore not unique to somatic cell changes) and SBS3, SBS5 and SBS8 mutation signatures comprising the common mutation across all cancer types can not be the preferred candidates used in the fitting process of action 420.
[0045] The inventors have also recognized that fitting techniques using minimization often prefer "flat" features where the mutation types are evenly distributed. An example of a flat feature (labeled SBS3) is shown in Figure 5A As shown in Figure 5AAs can be observed in , each mutation type in the SBS3 signature contributes similarly to the overall mutation spectrum (e.g., no single mutation type has a frequency exceeding 1%, 2%, 2.5%, 3%, 3.5%, 4%, or 4.5%). If included in the feature set provided as input to the fitting process in act 420, many or all variants in many mutation catalogs would be associated with this flat feature, resulting in overfitting of the feature. To address the issue of such flat features being included, in some embodiments of the present disclosure, only "highlighted" features may be included in the feature set provided as input to the fitting process in act 420, as shown in act 418 of process 400. An example of a mutation feature that may be considered a highlight (labeled SBS7a) is shown in Figure 5B As shown, the SBS7a signature shows a clear preference for a specific mutation type (in this case, C>T substitutions) relative to other mutation types. Other non-limiting examples of mutational signatures that were highlighted include SBS2 (activity of the APOBEC family of cytidine deaminases) and SBS10d (defective POLD1 proofreading).
[0046] In some embodiments, flat features can be identified visually, empirically (e.g., by fitting multiple mutation catalogs to a feature and noting that the feature tends to be selected for the majority of variants across multiple mutation catalogs), or by setting a threshold for the proportion of contribution of a given mutation type required to be included. For example, in some embodiments, to be considered a prominent feature (as opposed to a flat feature), a feature can have a given mutation type that represents at least 5% of all genetic mutations in the feature's mutation spectrum. In some embodiments, the threshold can be at least 10%, at least 15%, at least 20%, at least 25%, at least 50%, or at least 75% of all genetic mutations in the feature's mutation spectrum.
[0047] Select prominent mutational features for Figure 4 An example of the fitting process in action 420 is shown in Figure 6 In the picture. Figure 6 In the example of , the tumor sample is a urothelial carcinoma sample. In act 416 of process 400, it is determined that four mutational signatures are frequently observed for this cancer type: SBS1, SBS13, SBS2, and SBS5. Although each of these signatures is associated with urothelial carcinoma, SBS5 can be characterized as a flat (e.g., not prominent) feature and can be removed from the feature set used in the fitting process in act 418 of process 400 because it is not prominent, and SBS1 can be removed in act 416 of process 400 because it is observed in many cancer types and therefore may not be informative. Therefore, as Figure 6As shown in , the resulting mutation signature set used for fitting only includes SBS2 and SBS13 signatures. Visual comparison of the SBS2 and SBS13 mutation signatures with the mutation catalog of the tumor sample confirmed that the mutation spectrum of the SBS2 and SBS13 signatures includes relevant genetic mutations that characterize the tumor sample and also has a strong mutation phenotype. Specifically, Figure 6 It was shown that the tumor samples were essentially a combination of the C>T mutation spectrum of SBS2 and the C>G mutation spectrum of SBS13, albeit in different proportions. This could be due to the fact that the mutational inventory represents the exposure that the sample received from each of the mutational processes associated with SBS2 and SBS13. Figure 6 In the example, nearly 80% of the sample's mutation catalog can be explained by the SBS2 and SBS13 signatures, of which 60% is the result of exposure with SBS2 and 40% is the result of exposure with SBS13.
[0048] Select prominent mutational features for Figure 4 Another example of the fitting process of action 420 in Figure 7 In the picture. Figure 7 In the example of , the tumor sample is a melanoma sample. Figure 7 As shown in , the mutation catalog of melanoma samples includes both single base substitutions (top) and double base substitutions (bottom). It can be determined that for melanoma cancer types, a prominent single base substitution cancer mutation signature (SBS7a) and a prominent double base substitution cancer mutation signature (DBS1) are frequently observed, and this mutation signature set can be used Figure 4 The fitting in act 420 of process 400 is shown in
[0066] Visual comparison of the SBS7a and DBS1 mutation signatures with the mutation inventory of the melanoma tumor sample confirmed that the mutation profiles of the SBS7a and DBS1 signatures included relevant genetic mutations that characterized the tumor sample.
[0049] In some embodiments, the number of features used in the fitting process is between 2 and 5. In some cases, using additional features can affect the results of the fitting process, such as by overfitting.
[0050] Return to Figure 4, after completing the fitting in act 420, process 400 proceeds to act 422, where it can be determined whether the exposure of the sample to a particular feature determined during the fitting process is above a threshold. If it is determined in act 422 that the exposure of the sample to at least one feature is above a threshold, process 400 proceeds to act 424, where context (e.g., trinucleotide context) is added to each variant in the variant set. This context can be used to assign a context probability (act 430) from the subset of related features (act 426) identified in act 422, such that each variant is annotated with a likelihood of being due to exposure to one of the cancer mutation features used for the fitting process. For example, in Figure 6 In the example SBS13 signature of a urothelial carcinoma sample, C>A (TCC context) has a lower proportion than C>G (TCA context). Therefore, C>G (TCA) changes in the mutation catalog can be assumed to have a higher probability of being associated with urothelial carcinoma than C>A (TCC) changes. In some embodiments, the context probability can be determined as the sum of the associated exposures. For example, if it is determined that 60% of the C>T (TCA) variants are the result of exposure to SBS2 and 40% are the result of exposure to SBS13, the context probability of C>T (TCA) can be calculated as the sum of 0.6*C>T (TCA) frequency in SBS2 and 0.4*C>T (TCA) frequency in SBS13, thereby adjusting for the amount of exposure determined to each cancer mutation signature.
[0051] After assigning contextual probabilities to variants in act 430, process 400 proceeds to act 432, where variants are sampled based on the assigned contextual probabilities to generate a set of genetic variants most likely associated with the tumor sample in act 442. In particular, the sampling can generate genetic variations corresponding to mutational signature peaks. Alternatively, the genetic variants can be ranked based on the contextual probabilities, and the top few (e.g., top 6, 12, 16, 20, 48, 96) can be selected for inclusion in the set of variants determined in act 442. By considering mutational processes that may be associated with tumor development and probabilistically associating variants with these processes, the selected variants are more likely to be true somatic alterations rather than variants introduced by artificial processes.
[0052] If it is determined in act 422 that the fitting process did not identify any exposures greater than a threshold, process 400 can proceed to act 440, where the variants included in the VCF file are ranked based on one or more quality criteria. For example, variants can be ranked based on coverage depth in normal or healthy samples, number of reads between tumor and normal samples, allele frequency, probability that a variant is a germline variant, etc. Other examples of such quality criteria can be found in International Patent Publication WO2022029688A1, the contents of which are hereby incorporated by reference. Then, in act 442, the ranked variants can be used to determine a set of genetic variants that can be used to create a tumor-informed assay.
[0053] In some embodiments, variants can also be weighted based on the likelihood that they are driver mutations (i.e., mutations that may be associated with causing cancer) or passenger mutations. The latter can be particularly useful for identifying minimal residual disease because they are less likely to be eliminated by targeted therapy and are therefore still useful for identifying resistant clones that have lost driver mutations. In some embodiments, variants can also be weighted or ranked according to other characteristics, such as clonality (i.e., clone or subclone), mappability, or likelihood of amplification, as further described in International Patent Publication WO2022029688A1.
[0054] After determining the set of genetic variants in act 442 (e.g., by ranking the variants in act 440 or sampling the variants based on contextual probabilities in act 432), process 400 proceeds to act 444, where the set of variants can be used to create a tumor-informed assay. The tumor-informed assay can be designed to preferentially amplify the set of variants in subsequent plasma samples to detect residual disease. Examples of tumor-informed assays can be found, for example, in International Patent Publication No. WO2022-029688A1 and U.S. Patent Publication No. 2020 / 0157604, the contents of which are hereby incorporated by reference.
[0055] Although the techniques described herein can prioritize cancer-associated variants based on context (e.g., by sampling variants based on contextual probabilities in act 432 of process 400), some artifactual variants may still be selected for inclusion in the assay (although generally they may be prioritized lower than cancer-based variants). In this case, it may be advantageous to specifically exclude any variants that are not associated with any previously discovered cancer signature, since the likelihood that a genetic variant observed in a mutation catalog has not been previously seen in any known cancer mutation signature is low. Screening out possible artifactual variants in a mutation catalog can be performed by comparing each variant in the mutation catalog to various known cancer signatures via, for example, cosine similarity. In some embodiments, all available mutation signatures (e.g., not limited to a specific cancer type) can be used for artifactual variant screening, since screening can determine whether a given mutation spectrum is similar to any mutational process (e.g., signature) that has been previously observed. Any variant that is not similar to a previously observed signature may be the result of errors introduced by some other process, which may be due to, for example, sample preparation or amplification. Return to Figure 6 In the example of a urothelial carcinoma sample in , the artifactual variant filtering process might not exclude C->T and C->G variants because these mutational profiles are present in the SBS2 and SBS13 mutational signatures, respectively. However, other variants (e.g., T->C variants) can be excluded if they are not sufficiently similar to some other mutational processes observed in the mutational signature.
[0056] In some embodiments, rather than analyzing all variants in a mutation catalog, samples that exhibit an excess (e.g., >50%, >60%, >70%, >80%) of mutations of a particular mutation type (e.g., C>T, including all contexts) can be flagged for further analysis. The mutation spectrum of a given flagged mutation can be compared to all mutation signatures (or subsets thereof) that also exhibit a significant proportion of this mutation type using, for example, cosine similarity, as exemplified in Figure 8 If the mutation spectrum is too low in similarity to any feature, the mutation type may not be further considered for inclusion in the final genetic variant set. Figure 8 As shown in the example, samples with an excess of C>T mutations can be identified and then compared pairwise with multiple mutational signatures. Using cosine similarity, it can be determined that the mutational profile of C>T mutations is sufficiently similar to the corresponding mutational profiles from the SBS1, SBS2, and SBS7a mutational signatures (e.g., using cosine similarity and a threshold of 0.85). Therefore, this variant type can be retained for further consideration in the variant selection process.
[0057] In some embodiments of the present disclosure, artifactual variants in a mutation catalog can be identified and removed by comparing the variants to an "artifactual signature" generated from a mutation catalog corresponding to a sample containing many artifactual mutations. For example, variants generated by the formalin-fixed paraffin-embedded (FFPE) process can be identified and removed based on their similarity to the FFPE mutation signature. Figure 9 An example of an artifact signature based on FFPE mutations is shown. Such a signature can be used to identify and exclude certain variants from the mutation catalog that may be due to the FFPE process rather than to the mutagenic process of exposure to the specific type of cancer associated with the sample. Similarly, an artifact signature can be generated from samples with PCR amplification errors or other error types to provide a filter for excluding certain variants from the mutation catalog.
[0058] In some embodiments, the mutation signature set includes at least one mutation signature associated with exposure to a therapy, such as chemotherapy. Some mutation exposures occur earlier in the development of the cancer than other mutation exposures, and therefore some signatures are more likely to represent clonal variations, while other signatures are more likely to be subclonal variations (i.e., signatures associated with earlier exposures are more likely to be clonal, while signatures associated with some later events, such as chemotherapy, will be subclonal). Using this information allows selection of variants that are more likely to be clonal variants. This is valuable, for example, when a patient has already had some treatment (e.g., a breast cancer sample taken after neoadjuvant therapy). Variants from the tumor can be prioritized based on features by selecting the following variants: a) variants that are more likely to be clonal variants, and b) variants that are more likely to be true somatic alterations. Similarly, if a variant is more likely to be a clonal variant, it can be excluded.
[0059] Having thus described several aspects and embodiments of the technology set forth in this disclosure, it is to be understood that various alterations, modifications, and improvements will readily occur to those skilled in the art.
[0060] Such changes, modifications and improvements are intended to be within the spirit and scope of the technology described herein. For example, a person of ordinary skill in the art will readily envision various other means and / or structures for implementing the functions described herein and / or obtaining one or more of the results and / or advantages described herein, and each of these changes and / or modifications is considered to be within the scope of the embodiments described herein. Those skilled in the art will recognize or be able to use unconventional experiments to determine many equivalents of the specific embodiments described herein. Therefore, it should be understood that the above-mentioned embodiments are presented only by way of example, and within the scope of the appended claims and their equivalents, embodiments of the present invention may be practiced in a manner different from that specifically described. In addition, any combination of two or more features, systems, articles, materials, kits and / or methods described herein is included within the scope of the present disclosure, as long as such features, systems, articles, materials, kits and / or methods do not contradict each other.
[0061] The above embodiments can be implemented in any of many ways. One or more aspects and embodiments of the present disclosure related to the performance of a process or method can utilize program instructions that can be executed by a device (e.g., a computer, a processor or other device) to perform a process or method, or to control the performance of a process or method. In this regard, various inventive concepts can be embodied as a computer-readable storage medium (or multiple computer-readable storage media) (e.g., a circuit configuration in a computer memory, one or more hard disk drives, a flash memory, a field programmable gate array or other semiconductor devices, or other tangible computer storage media), which is encoded with one or more programs, and when executed on one or more computers or other processors, these programs implement one or more of the methods in the various embodiments described above. One or more computer-readable media can be removable so that one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects described above. In some embodiments, the computer-readable medium can be a non-transient medium.
[0062] The terms "program" or "software" are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be used to program a computer or other processor to implement the various aspects described above. In addition, it should be understood that, according to one aspect, one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but can be distributed in a modular manner among multiple different computers or processors to implement various aspects of the present disclosure.
[0063] Computer-executable instructions can take many forms, such as program modules, that are executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. Typically, in various embodiments, the functionality of the program modules can be combined or distributed as needed.
[0064] In addition, the data structure can be stored in a computer-readable medium in any suitable form. To simplify the description, the data structure can be shown as having fields that are related by position in the data structure. Such relationships can also be achieved by allocating storage space for the fields, and these storage spaces have locations in the computer-readable medium that convey the relationship between the fields. However, any suitable mechanism can be used to establish the relationship between the information in the fields of the data structure, including by using pointers, tags, or other mechanisms that establish relationships between data elements.
[0065] The above-described embodiments of the present technology can be implemented in any of many ways. For example, the embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or set of processors, whether provided in a single computer or distributed among multiple computers. It should be understood that any component or set of components that implement the functions described above can be generally regarded as a controller that controls the above-described functions. The controller can be implemented in many ways, such as with dedicated hardware, or with general-purpose hardware (e.g., one or more processors) that is programmed using microcode or software to implement the above-described functions, and when the controller corresponds to multiple components of the system, it can be implemented in a combination of ways.
[0066] Furthermore, it should be understood that a computer may be embodied in any of a variety of forms, such as, by way of non-limiting example, a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Furthermore, a computer may be embedded in a device not generally considered a computer but having suitable processing capabilities, including a personal digital assistant (PDA), a smart phone, or any other suitable portable or fixed electronic device.
[0067] In addition, a computer may have one or more input and output devices. Among other things, these devices may be used to present a user interface. Examples of output devices that may be used to provide a user interface include printers or display screens for visual presentation of output, and speakers or other sound generating devices for auditory presentation of output. Examples of input devices that may be used for a user interface include keyboards and pointing devices, such as mice, touchpads, and digitizers. As another example, a computer may receive input information through speech recognition or in other audible formats.
[0068] Such computers may be interconnected by one or more networks of any suitable form, including local or wide area networks, such as enterprise networks, and intelligent networks (IN) or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0069] Furthermore, as described, some aspects may be embodied as one or more methods. The actions performed as part of the method may be ordered in any suitable manner. Thus, embodiments may be constructed in which the actions are performed in an order different from that shown, which may include performing some actions simultaneously, even though the actions are shown as sequential in the illustrative embodiments.
[0070] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0071] Unless explicitly indicated to the contrary, the indefinite articles "a" and "an" as used in the specification and claims herein should be understood to mean "at least one".
[0072] As used in the specification and claims herein, the phrase "and / or" should be understood to mean "either one or both" of the elements so combined, i.e., elements that are present in some cases jointly and in other cases separately. Multiple elements listed with "and / or" should be interpreted in the same manner, i.e., "one or more" of the elements so combined. In addition to the elements specifically identified by the "and / or" clause, other elements may optionally be present, whether related or unrelated to those specifically identified elements. Thus, as a non-limiting example, when used in conjunction with open language such as "comprising", a reference to "A and / or B" may refer to only A (optionally including elements other than B) in one embodiment; to only B (optionally including elements other than A) in another embodiment; to both A and B (optionally including other elements) in yet another embodiment; and so on.
[0073] As used in the specification and claims herein, the phrase "at least one" when referring to a list of one or more elements should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but does not necessarily include at least one of each element specifically listed in the list of elements, and does not exclude any combination of elements in the list of elements. This definition also allows for the optional presence of elements other than the elements specifically identified in the list of elements to which the phrase "at least one" refers, whether related to or unrelated to those specifically identified elements. Thus, as a non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") can refer to at least one A in one embodiment, optionally including more than one A, without the presence of B (and optionally including elements other than B); in another embodiment, at least one B, optionally including more than one B, without the presence of A (and optionally including elements other than A); in yet another embodiment, at least one (optionally including more than one) A and at least one (optionally including more than one) B (and optionally including other elements); and so on.
[0074] In addition, the words and terms used herein are for descriptive purposes and should not be considered limiting. The words "include," "comprising," or "having," "containing," "involving," and variations thereof used herein are intended to encompass the items listed thereafter and their equivalents as well as additional items.
[0075] In the claims and the foregoing description, all transitional phrases such as "comprising," "including," "carrying," "having," "containing," "involving," "having," "consisting of," etc., shall be understood as open-ended, meaning including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively.
[0076] The use of ordinal terms such as "first," "second," "third," etc. in the claims to modify claim elements does not in itself imply any priority, precedence, or order of one claim element with respect to another, nor does it imply a temporal order in which method actions may be performed, but is merely used as a label to distinguish one claim element having a particular name from another element having the same name (except for the use of the ordinal term) to distinguish the claim elements.
Claims
1. A method for selecting genetic variants for tumor-informed assays, the method comprising: receiving a sample collected from a patient, the sample being associated with a cancer type; generating a mutation catalog for the sample, the mutation catalog indicating proportions of genetic mutation types observed in the sample; selecting a feature set associated with the cancer type, the set comprising one or more features, each feature comprising a mutation profile; determining, based on the set of features associated with the cancer type and the catalog of mutations, a set of genetic variants that are most likely to be true somatic variants associated with the sample; and The set of genetic variants is output for use in creating a tumor-informed assay for the patient.
2. The method of claim 1 , wherein selecting a feature set associated with the cancer type comprises: accessing a database configured to store a plurality of features associated with the cancer type, wherein the plurality of features are population-level features determined from a plurality of cancer samples.
3. The method of claim 2, further comprising: Only features from the database having one or more genetic mutation types that represent at least 5% of all genetic mutations in the mutation spectrum are included in the feature set.
4. The method of claim 3, further comprising including in the feature set only features from the database having one or more genetic mutation types that represent at least 10% of all genetic mutations in the mutation spectrum.
5. The method of claim 4, further comprising including in the feature set only features from the database having one or more genetic mutation types that represent at least 25% of all genetic mutations in the mutation spectrum.
6. The method of any one of claims 1 to 5, wherein determining the set of genetic variants comprises: The mutation spectra in the feature set are fitted to the mutation catalog to determine the corresponding amount of each feature observed in the sample.
7. The method of claim 6, wherein determining the set of genetic variants further comprises: A determination is made as to whether the amount of each feature observed in the sample is greater than a threshold value.
8. The method of claim 6, wherein determining the set of genetic variants further comprises: for each of the genetic variants in the sample, associating a contextual probability based on the frequency of the genetic variant in the feature set and the determined amount of each feature observed in the sample; and The genetic variants are sampled weighted by their associated contextual probabilities to determine the genetic variants to include in the set of genetic variants.
9. The method of claim 8, wherein the context probability is based on a trinucleotide context.
10. The method according to any one of claims 1 to 9, further comprising: screening the set of genetic variants for use in the tumor-informed assay based at least in part on the suitability of each of the genetic variants in the set of genetic variants; as well as The tumor-informed assay is created based on the set of genetic variants within the screened set.
11. The method of any one of claims 1 to 10, wherein the set of features associated with the cancer type comprises at least two features.
12. The method of any one of claims 1 to 11, wherein selecting the feature set associated with the cancer type further comprises selecting a mutation profile associated with an error source introduced during a sequencing process, the method further comprising: Any genetic variant determined to be attributable to the error source is excluded from the set of genetic variants.
13. The method of claim 12, wherein the error source is associated with a formalin-fixed paraffin-embedded (FFPE) process, an amplification process, or a sequencing process.
14. The method of any one of claims 1 to 13, wherein selecting the feature set associated with the cancer type further comprises selecting a mutation profile associated with a therapy, the method further comprising: Any genetic variant determined to be attributable to the therapy is excluded from the set of genetic variants.
15. The method of claim 14, wherein the therapy is chemotherapy.
16. The method of any one of claims 1 to 15, wherein generating a mutation catalog for the sample comprises performing whole exome sequencing on the sample.
17. The method of any one of claims 1 to 16, wherein the feature set comprises at least one double base substitution feature.
18. The method of any one of claims 1 to 17, wherein at least some of the mutation profiles are associated with different exposure types.
19. The method of any one of claims 1 to 18, wherein the mutation spectrum indicates the proportion of genetic mutation types observed in a population.
20. The method of claim 19, wherein the genetic mutation type comprises a trinucleotide context.
21. The method according to any one of claims 1 to 20, wherein the feature Represents exposure to a mutational process.
22. The method of claim 21, wherein the mutational process is associated with cancer.
23. A method for selecting a genetic variant, the method comprising: receiving samples collected from patients; generating a mutation catalog from the set of genetic variants, the mutation catalog indicating proportions of genetic mutation types observed in the sample; selecting a feature set, the set comprising one or more features, each feature comprising a mutation profile; Determine the set of genetic variants; excluding from the set of genetic variants any genetic variants determined to be attributable to the one or more characteristics; and The set of genetic variants is output for use in creating an assay for the patient.
24. The method of claim 23, wherein the feature set is associated with an error source.
25. The method of claim 25, wherein the error source is associated with a formalin-fixed paraffin-embedded (FFPE) process, an amplification process, or a sequencing process.
26. The method of any one of claims 23 to 25, wherein the feature set is associated with a therapy.
27. The method of claim 26, wherein the therapy is chemotherapy.
28. The method of claim 26, wherein the excluding further comprises excluding from the set of genetic variants any genetic variant determined to be a subclonal genetic variant associated with the therapy.
29. The method of any one of claims 23 to 28, wherein the feature set further comprises at least one mutation profile associated with a cancer type, wherein the determining further comprises: The set of genetic variants associated with the sample is determined based on the at least one mutation profile associated with the cancer type and the mutation catalog.
30. A system comprising: At least one hardware computer processor programmed to perform any of the methods described in claims 1 to 29.
31. A computer readable medium encoded with a plurality of instructions which, when executed by at least one hardware computer processor, perform the method of any one of claims 1 to 29.
32. A tumor-informed assay for monitoring the presence of cancer in a patient, wherein the tumor-informed assay is configured to detect at least one genetic variant from the output set of genetic variants according to any one of claims 1 to 29.
Citation Information
Patent Citations
Method for the Analysis of Minimal Residual Disease
US20200157604A1
Highly sensitive method for detecting cancer DNA in a sample
WO2022029688A1