Filtering cancer-associated genetic variants using mutational signatures

By filtering genetic variants using cancer-specific mutational signatures, the method addresses the issue of artifactual variants in MRD detection, improving assay sensitivity and accuracy in identifying residual cancer cells.

JP2026504795APending Publication Date: 2026-02-10イニバタ エルティーディー
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025536487
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-18
Filing Date
2024-01-17
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing cancer treatment methods fail to accurately detect minimal residual disease (MRD) due to the inclusion of artifactual genetic variants caused by sample processing and sequencing errors, reducing the sensitivity of assays.

Method used

Utilize mutational signatures associated with different cancer types to filter and select genetic variants that are likely to be true somatic variants, excluding artifactual variants by matching a mutation catalog to a database of population-level signatures, and creating a tumor-informed assay.

Benefits of technology

Improves the sensitivity of MRD assays by prioritizing cancer-associated genetic variants, reducing the need for additional normal samples and enhancing the accuracy of detecting residual cancer cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026504795000001_ABST
    Figure 2026504795000001_ABST
Patent Text Reader

Abstract

A method and apparatus for selecting genetic variants for a tumor-informed assay is provided, comprising receiving a sample from a patient, the sample being associated with a cancer type, generating a mutation catalog for the sample, the mutation catalog indicating a proportion of genetic variants observed in the sample, selecting a signature set associated with the cancer type, the signature set including one or more signatures, each signature including a mutation profile, determining a set of genetic variants that are most likely to be true somatic variants associated with the sample based on the signature set associated with the cancer type and the mutation catalog, and outputting the set of genetic variants for use in creating a tumor-informed assay for the patient.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to using mutational signatures to filter genetic variants associated with cancer. [Background technology]

[0002] After cancer treatment, a small number of cancer cells may remain in a patient's body even if the patient appears to be in remission. These remaining cells are called "minimal residual disease" (MRD) and may cause recurrence. Assays for detecting MRD (e.g., circulating tumor DNA (ctDNA) assays) can use various techniques. These techniques involve sequencing a patient's tumor tissue to identify tumor-informative genetic variants that, when detected in the patient's cell-free DNA (cfDNA), may indicate MRD. Summary of the Invention

[0003] In some embodiments, data analysis techniques are provided for selecting a set of genetic variants for use in creating tumor-informed assays. Information about genetic variants in mutational signatures for different cancer types is used, at least in part, to filter the set of genetic variants observed in the sample's mutation catalog to identify genetic variants that are more likely to result from mutational processes specific to the sample than from other processes, including artifactual processes resulting from sample processing. The identified genetic variants can then be used to create tumor-informed assays.

[0004] In some embodiments, a method for selecting genetic variants for a tumor-informed assay is provided, comprising: receiving a sample collected from a patient, the sample associated with a cancer type; generating a mutation catalog for the sample, the mutation catalog indicating a proportion of genetic variants observed in the sample; selecting a signature set associated with the cancer type, the signature set including one or more signatures, each signature including a mutation profile; determining a set of genetic variants that are most likely to be true somatic variants associated with the sample based on the signature set associated with the cancer type and the mutation catalog; and outputting the set of genetic variants for use in creating an assay for the patient.

[0005] In one embodiment, selecting a signature set associated with a cancer type comprises accessing a database configured to store a plurality of signatures associated with the cancer type, wherein the plurality of signatures are population-level signatures determined from a plurality of cancer samples. In another embodiment, the method further comprises including in the signature set only signatures from the database having one or more genetic variants that account for at least 5% of all genetic variants in the mutation profile. In another embodiment, the method further comprises including in the signature set only signatures from the database having one or more genetic variants that account for at least 10% of all genetic variants in the mutation profile. In another embodiment, the method further comprises including in the signature set only signatures from the database having one or more genetic variants that account for at least 25% of all genetic variants in the mutation profile.

[0006] In another aspect, determining the set of genetic variants comprises matching the mutation profile in the signature set to the mutation catalog to determine the corresponding amount at which each signature is observed in the sample. In another aspect, determining the set of genetic variants further comprises determining whether the amount at which each signature is observed in the sample exceeds a threshold. In another aspect, determining the set of genetic variants further comprises associating a context probability for each genetic variant in the sample based on the frequency of the genetic variant in the signature set and the determined amount at which each signature is observed in the sample, and sampling the genetic variants weighted by the associated context probability to determine which genetic variants to include in the set of genetic variants. In another aspect, the context probability is based on trinucleotide context.

[0007] In another embodiment, the method further comprises filtering the set of genetic variants based at least in part on the suitability of each genetic variant in the set of genetic variants for use in a tumor-informed assay, and creating a tumor-informed assay based on the set of genetic variants in the filtered set. In another embodiment, the cancer type-associated signature set comprises at least two signatures. In another embodiment, selecting the cancer type-associated signature set further comprises selecting a mutation profile associated with an error source introduced during the sequencing process, and the method further comprises excluding from the set of genetic variants all genetic variants determined to be attributable to the error source. In another embodiment, the error source is associated with the formalin-fixed, paraffin-embedded (FFPE) process, the amplification process, or the sequencing process.

[0008] In another aspect, selecting a signature set associated with a cancer type further comprises selecting a mutation profile associated with a treatment, and the method further comprises excluding from the set of genetic variants all genetic variants determined to be caused by the treatment. In another aspect, the treatment is chemotherapy.

[0009] In another embodiment, generating a mutation catalog for the sample comprises performing whole-exome sequencing on the sample. In another embodiment, the signature set comprises at least one double base substitution signature. In another embodiment, at least a portion of the mutation profile is associated with different exposure types. In another embodiment, the mutation profile represents the proportion of genetic variants observed in a population. In another embodiment, the genetic variants comprise trinucleotide contexts. In another embodiment, the signature represents exposure to a mutational process. In another embodiment, the mutational process is associated with cancer.

[0010] In some embodiments, a method for selecting genetic variants is provided, comprising receiving a sample from a patient, generating a mutation catalog from a set of genetic variants, the mutation catalog indicating a proportion of genetic variants observed in the sample, selecting a signature set, the signature set including one or more signatures, each signature including a mutation profile, determining a set of genetic variants, excluding from the set of genetic variants all genetic variants determined to be attributable to the one or more signatures, and outputting the set of genetic variants for use in generating an assay for the patient.

[0011] In one embodiment, the signature set is associated with an error source. In another embodiment, the error source is associated with a formalin-fixed, paraffin-embedded (FFPE) process, an amplification process, or a sequencing process. In another embodiment, the signature set is associated with a treatment. In another embodiment, the treatment is chemotherapy. In another embodiment, the excluding step further comprises excluding from the set of genetic variants all genetic variants determined to be subclonal genetic variants associated with the treatment. In another embodiment, the signature set further comprises at least one mutation profile associated with a cancer type, and the determining step further comprises determining a set of genetic variants associated with the sample based on the at least one mutation profile associated with the cancer type and the mutation catalogue.

[0012] In some embodiments, a system is provided, the system including at least one hardware computer processor programmed to perform any one of the methods described herein.

[0013] In some embodiments, a computer-readable medium is provided that is encoded with a plurality of instructions that, when executed by at least one hardware computer processor, perform any of the methods described herein.

[0014] In some embodiments, a tumor-informed assay for monitoring the presence of cancer in a patient is provided, wherein the tumor-informed assay is configured to detect at least one genetic variant from an output set of genetic variants according to any of the methods described herein. [Brief explanation of the drawings]

[0015] Various non-limiting embodiments of the present technology are described with reference to the following drawings, which should be understood as not necessarily being to scale.

[0016] [Figure 1]1 is an exemplary representation of a mutation catalog for a tumor sample that may be used to identify genetic variants for tumor-informed assays, according to some embodiments of the present disclosure. [Figure 2A-1] 1 is an exemplary representation of a single-base substitution mutation signature that may be used to identify a set of genetic variants for a tumor-informed assay in some embodiments of the present disclosure. [Figure 2A-2] 1 is an exemplary representation of a single-base substitution mutation signature that may be used to identify a set of genetic variants for a tumor-informed assay in some embodiments of the present disclosure. [Figure 2B-1] 1 is an exemplary representation of a double-base substitution mutation signature that may be used to identify a set of genetic variants for a tumor-informed assay in some embodiments of the present disclosure. [Figure 2B-2] 1 is an exemplary representation of a double-base substitution mutation signature that may be used to identify a set of genetic variants for a tumor-informed assay in some embodiments of the present disclosure. [Figure 3] 1 is a flowchart of a process for obtaining a set of genetic variants for use in creating a tumor-informed assay, according to some embodiments of the present disclosure. [Figure 4] 1 is a flowchart of a process for obtaining a set of genetic variants for use in creating a tumor-informed assay, according to some embodiments of the present disclosure. [Figure 5A-1] 1 is an exemplary representation of a flat mutational signature that may be present in a database of mutational signatures, according to some embodiments of the present disclosure. [Figure 5A-2] 1 is an exemplary representation of a flat mutational signature that may be present in a database of mutational signatures, according to some embodiments of the present disclosure. [Figure 5B-1] 1 is an exemplary representation of a characteristic mutation signature that may be selected for use in the matching process, according to some embodiments of the present disclosure. [Figure 5B-2]1 is an exemplary representation of a characteristic mutation signature that may be selected for use in the matching process, according to some embodiments of the present disclosure. [Figure 6-1] 1 is an exemplary representation of selecting multiple single base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 6-2] 1 is an exemplary representation of selecting multiple single base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 6-3] 1 is an exemplary representation of selecting multiple single base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 6-4] 1 is an exemplary representation of selecting multiple single base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 6-5] 1 is an exemplary representation of selecting multiple single base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 6-6] 1 is an exemplary representation of selecting multiple single base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 7-1] 10 is an exemplary representation of selecting single-base substitution mutation signatures and double-base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 7-2] 10 is an exemplary representation of selecting single-base substitution mutation signatures and double-base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 7-3] 10 is an exemplary representation of selecting single-base substitution mutation signatures and double-base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 7-4] 10 is an exemplary representation of selecting single-base substitution mutation signatures and double-base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 7-5] 10 is an exemplary representation of selecting single-base substitution mutation signatures and double-base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 7-6] 10 is an exemplary representation of selecting single-base substitution mutation signatures and double-base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 7-7] 10 is an exemplary representation of selecting single-base substitution mutation signatures and double-base substitution mutation signatures for use in the matching process, according to some embodiments of the present disclosure. [Figure 8-1] 1 is an example of a process for determining whether a genetic variant resembles one or more known mutational signatures, according to some embodiments of the present disclosure. [Figure 8-2] 1 is an example of a process for determining whether a genetic variant resembles one or more known mutational signatures, according to some embodiments of the present disclosure. [Figure 8-3] 1 is an example of a process for determining whether a genetic variant resembles one or more known mutational signatures, according to some embodiments of the present disclosure. [Figure 9] 1 is an exemplary representation of a mutational signature for an FFPE process that may be used to filter artifactual variants, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0017] Aspects of the technology described herein relate to techniques for determining a set of genetic variants to be included in a tumor-specific assay (e.g., an MRD assay) for a patient. The inventors recognize that at least some of the genetic variants identified by sequencing a patient's tumor tissue may be artifactual mutations caused by processes such as cytosine deamination, initial integration errors and polymerase bias due to PCR enrichment, and integration errors due to sequencing. Including such genetic variants in a tumor-informed assay for a patient may reduce the sensitivity of the assay because the variants are not present in the patient's cancer. Some embodiments of the present disclosure are directed to techniques for identifying a set of genetic variants for use in a tumor-informed assay that improves the sensitivity of the assay, for example, in detecting residual cancer cells.

[0018] Some embodiments of the present disclosure relate to a tumor-informed variant filtering method that applies mutational signature analysis to select good genetic variants that are most likely to be present in a tumor or cancer type. Genetic mutations occur in human DNA through various mutation processes, including but not limited to, intrinsic errors in DNA replication and exposure to various physical, chemical, or biological mutagens. Different mutation processes may generate specific combinations of genetic variants, and as a result, unique mutational signatures may be associated with the mutational process. The existence of such mutational signatures for various cancer-related mutational processes has been studied in recent years and published in publicly accessible databases (e.g., the COSMIC database, available at https: / / cancer.sanger.ac.uk / cosmic).

[0019] The inventors recognize and understand that mutational signatures associated with different forms of cancer can be used to filter genetic variants present in samples from patients. The filtered set of genetic variants can be used, for example, in tumor-informed assays for the patient. As described in more detail below, some embodiments attempt to match genetic variants in a mutation catalog determined for a tumor sample from a patient to multiple mutational signatures associated with the patient's cancer. Genetic variants associated with the signature for the tumor sample (e.g., variants that match the predicted cancer type) can be prioritized for incorporation into the tumor-informed assay, improving the sensitivity of the assay in terms of detecting cancer in future samples collected from the patient. Furthermore, by preferentially selecting genetic variants based on their association with known cancer mutational signatures, additional normal or healthy patient samples may not be required to identify tumor-associated variants compared to tumor tissue. The techniques described herein may also be applied to quality control, for example, by rejecting samples that appear to be from different cancer types (indicating potential handling issues) or, as explained in more detail below, by filtering out genetic variants that appear to be associated with ex vivo mutational processes that are unlikely to be relevant to the patient's cancer.

[0020] A mutation catalog is a mutation profile containing a set of genetic variants observed in a single biological sample, e.g., a tumor sample. In some embodiments, a mutation catalog may include a context, such as a trinucleotide context, which identifies different combinations of nucleotides immediately 5' (e.g., preceding) and 3' (e.g., following) of a given mutation in a DNA sequence. FIG. 1 illustrates an example of a mutation catalog generated for a tumor sample according to some embodiments of the present disclosure. As shown, if six variants (C>A, C>G, C>T, T>A, T>C, T>G) and 16 trinucleotide contexts for each variant are considered, the mutation catalog provides 96 "channels" of information. In some embodiments, a mutation catalog for a tumor sample may be determined by identifying multiple genetic variants relative to a reference sequence in the tumor sample and grouping the variants according to their context. As shown in FIG. 1, the mutation catalog may be represented as a histogram, and the height of the bars in the histogram may correspond to either a mutation count (as shown in FIG. 1) or a ratio (e.g., the percentage of single-base substitutions). Thus, a mutation catalog for a tumor sample can indicate whether certain variants are over- or under-represented in the sample and can provide insight into the genetic variants that are most frequently found in tumor samples.

[0021] As described above, mutational signatures related to various mutational processes (e.g., mutational processes associated with cancer) have been studied and published in various databases (e.g., via the COSMIC database available at https: / / cancer.sanger.ac.uk / cosmic). Mutational signatures may be generated from multiple samples from different individuals. For example, mutational signatures may represent recurring patterns of sequence context-dependent gene variants observed in patients with similar etiologies. Several unique mutational signatures associated with different cancer types and causes (e.g., exposure to different mutagens) have been identified. Mutational signatures associated with different cancer types can be generated based on whole genome sequencing (WGS) of cancer tissues from thousands of samples, and the results are clustered to generate characteristic mutational signatures for each cancer type. Each mutational signature may represent the overall picture of both passenger and driver gene mutations associated with that cancer. Mutational signatures can be associated with disease (eg, lung, colon, or skin cancer) and / or underlying causes of disease (eg, oxidation, microsatellite instability, exposure to ultraviolet light).

[0022] Similar to mutation catalogs for samples (e.g., the mutation catalog shown in Figure 1), mutation signatures determined based on population data can be represented as histograms of genetic mutations. A histogram of genetic mutations for an exemplary mutation signature (labeled SBS4) is shown in Figure 2A. The SBS4 mutation signature is associated with lung cancer due to smoking-induced oxidation and is characterized by an excess of C>A and G>T mutations (as well as CC>AA and GG>TT double substitutions and C / G deletions, not shown). As shown in Figure 2A, genetic mutations in a mutation signature can be further characterized by context, such as trinucleotide context (i.e., different combinations of 5' and 3' nucleotides adjacent to any given genetic mutation). When trinucleotide context is characterized, characterization of genetic mutations generates 96 classes of profiles for any given sample.

[0023] While the exemplary mutational signature shown in Figure 2A includes genetic mutations that are single-base substitutions, it should be understood that mutational signatures can also include genetic mutation types other than single-base substitutions, including, but not limited to, double-base substitutions, insertions and deletions, copy number variations, and microsatellites. A histogram of genetic mutations for an exemplary mutational signature (labeled DBS1) that includes double-base substitutions is shown in Figure 2B. As shown, the melanoma-associated DBS1 mutational signature is characterized primarily by CC>NN mutations, particularly CC>NN mutations with a TT context (i.e., TCCT>TNNT). Other contexts for genetic mutations can also be used. For example, additional nucleotides located 5' and / or 3' of the mutation may be used to provide context, such as dinucleotide, trinucleotide, tetranucleotide, pentanucleotide, or hexanucleotide. Insertions and deletions may be further characterized by homopolymer length, number of repeat units, and microhomology length. Copy number variations may be further characterized by length (e.g., 0-100 kb, 100 kb-1 Mb, greater than 1 Mb, 100 kb-1 Mb, 1 Mb-10 Mb, 10 Mb-40 Mb, greater than 40 Mb), hyperdiploidy, loss of heterozygosity, or heterozygosity. Additionally, although further resolution may be gained by classifying genetic variants by context, in some embodiments, additional nucleotide context may not be used and variants may be classified simply by their base change (e.g., C to A, C to G, T to C, etc.).

[0024] As described above, mutational signatures may be generated from multiple samples from different individuals. For example, mutational signatures may be determined by identifying recurring patterns in multiple mutation catalogs from individual patient samples. For example, consider a large number (e.g., 1,000) of mutation catalog datasets generated from biopsies of lung, skin, and breast cancer samples. As shown in FIG. 1, each mutation catalog represents the count or frequency of different genetic variants (e.g., single base substitutions, double base substitutions, insertions, deletions, etc.) associated with, for example, the trinucleotide context of the genetic variant, providing 96 different channels of information for the sample (6 variants across 16 different trinucleotide contexts). By examining recurring patterns across mutation catalogs within a dataset (e.g., whether a C-to-A base change in a particular context occurs across multiple samples), mutational signatures associated with different processes and / or cancer types may be identified, with each mutational signature indicating the frequency of a recurrent variant observed in the sample set. Any suitable pattern recognition technique may be used to identify recurring patterns of genetic variants in the mutation catalog dataset. Such techniques include, but are not limited to, machine learning techniques such as non-negative matrix factorization (NMF), principal component analysis (PCA), and vector quantization (VQ). Conventional software packages are available for extracting mutational signatures from large datasets of mutation catalogs (e.g., the open-source MutationalPatterns R package, available at https: / / bioconductor.org / packages / release / bioc / html / MutationalPatterns.html).

[0025] While users can define parameters for extracting mutational signatures from mutation catalog datasets, mutational signatures are most commonly extracted for individual cancer types (e.g., lung, breast) and / or associated exposures (e.g., smoking, aging). Thus, each mutational signature extracted in this way may represent the mutational processes present in a tumor sample set. While mutational signatures and mutation catalogs may appear to provide similar information, there are important differences between them. For example, each mutational signature represents the proportion of genetic mutations across the mutation catalog dataset from which it was extracted, but includes only recurring patterns observed in that dataset, rather than the count or ratio of all genetic mutations in the sample, as represented in the mutation catalog. Mutational signatures may be extracted, for example, by clustering mutation catalogs from many samples and selecting only samples with recurring, commonly occurring profiles. An example of such mutational signature extraction is described in Degasperi et al., Substitution mutational signatures in whole-genome-sequenced cancers in the UK population, Science 376:6591(2022).

[0026] The inventors recognize and understand that the genetic mutation information in a mutational signature can be leveraged to select a set of genetic variants for use in a tumor-informed assay. Accordingly, some embodiments of the present disclosure align or "match" a mutation catalog determined for a tumor sample to a mutational signature set, thereby enabling a determination of which mutational processes (represented in the signature) are most closely associated with the sample.

[0027] FIG. 3 illustrates a process (300) for determining a set of genetic variants for use in a tumor-informed assay, according to some embodiments of the present disclosure. In step (310), a mutation catalog is generated for a sample from a patient. For example, as described in further detail below, the sample may be processed and DNA sequenced to generate a text-based file describing the genetic mutations in the sample relative to a reference genetic profile. The occurrence of genetic mutations, and optionally their context, may be formulated as a mutation catalog for the sample. Process (300) then proceeds to step (312), in which a mutation signature set associated with a specific cancer for the sample is selected. For example, the June 2022 version of the COSMIC database v.3.3 contains approximately 130 distinct mutation signatures extracted from mutation catalogs of numerous tumor samples for different cancers. In step (312), a mutation signature set most likely associated with the specific cancer associated with the sample analyzed in step (310) may be selected.

[0028] Process (300) then proceeds to step (314), where the mutation catalog generated in step (310) and a mutation signature set within the set selected in step (312) are used to determine a set of genetic variants that are most likely to be true somatic variants associated with the sample. For example, as described in further detail below, the mutation catalog for the sample may be matched to the mutation signature set to determine a weighted sum of mutation signatures that represent the exposures most likely to result in the mutation catalog. By way of example, a mutation catalog from a lung cancer patient may include mutations associated with aging (e.g., C to T changes due to deamination), smoking, and other mutational processes. Due to differences in exposure to these mutational processes, different genetic mutations associated with the signatures may be present in some lung cancer samples but absent in others. As described herein, matching the mutation catalog to the mutation signature set may deconvolute the mutational processes active in a particular patient's tumor. By matching the mutation catalog to a mutation signature set, each genetic variant in the mutation catalog can be associated with a specific mutation signature, thereby providing insight into the mutational processes associated with the evolution of that tumor sample.

[0029] After the set of genetic variants characterizing the sample is determined in step (314), process (300) proceeds to step (316), where the set of genetic variants is output for use in creating tumor-informed assays for the patient. By detecting genetic variants that are not artifactual and are likely to be cancer-associated, and selecting variants that characterize the underlying mutational processes associated with the sample, assays that are sensitive to those genetic variants can be generated.

[0030] Figure 4 shows a process (400) for creating a tumor-informed assay based on a set of genetic variants selected by the techniques described herein. In step (410), a sample (e.g., a blood sample, tumor biopsy, or other tissue sample) may be processed (e.g., using whole-exome sequencing or other next-generation sequencing technology) to determine genetic sequence variants in the sample relative to a reference genome, and the output is a text-based variant call format (VCF) file. Process (400) may then proceed to step (412), in which the genetic variants identified in the VCF file may optionally be subjected to one or more filtering criteria to identify low-quality genetic variants, such as low coverage, low base or low mapping quality, redundant surrounding sequence, or presence in online databases (e.g., having a frequency of greater than 0.05 in dbSNP, indicating a high likelihood of being a germline variant). Process (400) then proceeds to step (414), where a mutation catalog for the sample is generated using the genetic variants identified in the VCF file that were not filtered in step (412). For example, a matrix containing genetic variants and their counts or proportions in the mutation profile of the sample may be generated as the mutation catalog for the sample. An exemplary mutation catalog for a sample is shown in FIG. 1 above. While filtering low-quality variants may be optional in some embodiments of the present disclosure, removing these variants may be useful because these variants may affect the signature matching process described herein.

[0031] Process (400) then proceeds to step (420), in which the mutation catalog for the sample is matched to one or more mutation signature sets. As described above, matching the mutation catalog to one or more mutation signatures attempts to reveal the underlying mutational processes the sample has undergone by mapping variants in the mutation catalog for the sample to variants in known genetic profiles for different mutational processes. For example, consider a mutation catalog containing 1,000 mutations. After matching in step (420), 500 mutations may be attributable to the SBS4 signature (smoking), 200 mutations may be attributable to SBS3 (defective homologous recombination DNA damage repair), and the remaining mutations may be associated with other signatures or no signatures at all. Each event attributable to a particular signature may be considered an "exposure" to the process associated with that signature. Thus, by matching a mutation catalog for a sample from a patient to a population-based signature set according to some embodiments of the present disclosure, it may be possible to indicate that the patient has been exposed to a particular mutational process defined by a particular cancer signature.

[0032] The matching technique attempts to determine the contribution of each mutational process to the genetic variants expressed in the mutation catalog by quantifying the presence and frequency of each mutational signature in the mutation catalog. In some embodiments, matching may be performed by creating two matrices: a sample matrix M and a signature matrix P. The sample matrix M may contain 96 rows for each mutation / trinucleotide context combination and n columns for each mutation catalog. The signature matrix P may contain k rows for each mutation signature and 96 columns for each trinucleotide context. Given these two matrices, a weight matrix E may be determined, containing k columns for each signature and n rows for each sample, such that the reconstructed tumor sample matrix R (calculated as M-(P*E)) minimizes a given error threshold e. In other words, the weight matrix E may be determined so that matrix E best replicates M by finding the optimal exposure combination that minimizes the difference between (P*E) and M. For a given sample n, the highest weight of the corresponding row in E may be selected to understand which mutational process is most likely to cause each mutation type in the mutation catalog. It should be understood that other techniques, including but not limited to other minimization techniques, golden search, minimal quadratic, etc., can alternatively be used to fit the mutation catalog to the mutation signature set in step (420).

[0033] The inventors recognize and appreciate that the accuracy of the genetic variants output from the matching process in step (420) varies significantly depending on the mutational signatures provided as input to the matching process. Accordingly, some embodiments of the present disclosure include one or more signature curation steps for selecting a set of signatures from a database of signatures for use in the matching step (420). As shown in FIG. 4 , the signature curation step may include a step (416) of selecting signatures expected to be most common or otherwise observed in a particular cancer type associated with the sample. For example, the inventors recognize that due to the sensitivity of the matching algorithm, a matching result may associate a mutational process that is unlikely to occur in a particular patient, e.g., by matching the mutation catalog to a mutational signature associated with a mutational process that is not associated with the particular cancer type associated with the sample. In some cases, the matching algorithm may force a particular variant to match a signature even if the sample was not exposed to that mutational process. Therefore, the accuracy of the genetic variants output from the matching process in step (420) can be improved by pre-selecting specific signatures that are expected to be observed in such samples.

[0034] In one example, the "smoking signature" SBS4 shown in FIG. 2A is found in various tumor types, even when tobacco carcinogens are unlikely to reach the site (e.g., in prostate cancer). This may be because the SBS4 signature is dominated by C>A and G>T mutations, which are found in many cancers. Thus, if the SBS4 signature were included in the mutational signature set used to analyze a prostate cancer sample, the matching process in step (420) may preferentially match variants to SBS4 rather than other signatures due to this strong phenotype, even though the patient has never smoked. Similarly, mutational signatures associated with mutational processes common across all sample types may not yield as meaningful results in the matching process as mutational signatures associated with more specific processes that are not common across all or many sample types. For example, the SBS1 mutational signature (which is related to aging and is therefore not limited to somatic changes), and the SBS3, SBS5, and SBS8 mutational signatures, which include mutations common across all cancer types, may not be preferred candidates for use in the matching process of step (420).

[0035] The inventors also recognize that fitting techniques using minimization tend to favor "flat" signatures in which variants are evenly distributed. An example of a flat signature (labeled SBS3) is shown in FIG. 5A. As can be observed in FIG. 5A, each variant in the SBS3 signature has a similar contribution to the overall mutation profile (e.g., no variant has a frequency greater than 1%, 2%, 2.5%, 3%, 3.5%, 4%, or 4.5%). When included in the signature set provided as input to the fitting process in step (420), many or all variants in many mutation catalogs may be associated with this flat signature, which may result in overfitting to the signature. In an effort to address the incorporation of such flat signatures, in some embodiments of the present disclosure, as shown in step (418) of process (400), only "distinctive" signatures may be included in the signature set provided as input to the fitting process in step (420). An example of a mutational signature (labeled SBS7a) that may be considered distinctive is shown in FIG. 5B. As shown, the SBS7a signature shows a clear preference for a particular mutation type (in this case, a C>T substitution) compared to other mutation types. Other non-limiting examples of distinctive mutational signatures include SBS2 (activity of the APOBEC family of cytidine deaminases) and SBS10d (deficiency in POLD1 proofreading).

[0036] In some embodiments, flat signatures may be identified visually, empirically (e.g., by fitting multiple mutation catalogs to a signature and verifying that the signature tends to select for most variants across the mutation catalogs), and by setting a threshold for the ratio of contributions of a given mutation type required for incorporation. For example, in some embodiments, to be considered a distinctive signature (as opposed to a flat signature), a signature may have a given mutation type that represents at least 5% of all genetic mutations in the signature's mutation profile. In some embodiments, this threshold may be at least 10%, at least 15%, at least 20%, at least 25%, at least 50%, or at least 75% of all genetic mutations in the signature's mutation profile.

[0037] An example of selecting characteristic mutational signatures for use in the matching process of step (420) in Figure 4 is shown in Figure 6. In the example of Figure 6, the tumor sample is a urothelial cancer sample. In step (416) of process (400), four mutational signatures (SBS1, SBS13, SBS2, and SBS5) are determined to be frequently observed for this cancer type. While each of these signatures is associated with urothelial cancer, SBS5 may be characterized as a flat (e.g., non-distinctive) signature and, because it is not distinctive, may be removed from the signature set used in the matching process in step (418) of process (400), and SBS1 may be removed in step (416) of process (400) because it is observed in many cancer types and may not be informative. Thus, the resulting mutational signature set used for matching includes only the SBS2 and SBS13 signatures, as shown in Figure 6. Visual comparison of the SBS2 and SBS13 mutation signatures with the mutation catalog of tumor samples confirms that the mutation profiles of the SBS2 and SBS13 signatures contain relevant genetic mutations for characterizing tumor samples and have strong mutation phenotypes. In particular, Figure 6 shows that tumor samples are essentially a combination of the SBS2 C>T mutation profile and the SBS13 C>G mutation profile, but the ratios are different. This may be because the mutation catalog indicates the amount of exposure that the sample has received from each of the mutation processes related to SBS2 and SBS13. In the example of Figure 6, nearly 80% of the mutation catalog for the sample can be explained by the SBS2 and SBS13 signatures, of which 60% is the result of exposure to SBS2 and 40% is the result of exposure to SBS13.

[0038] Another example of selecting characteristic mutational signatures for use in the matching process of step (420) in Figure 4 is shown in Figure 7. In the example of Figure 7, the tumor sample is a melanoma sample. As shown in Figure 7, the mutational catalog for the melanoma sample includes both single base substitutions (top) and double base substitutions (bottom). It may be determined that one characteristic single base substitution cancer mutational signature (SBS7a) and one characteristic double base substitution cancer mutational signature (DBS1) are frequently observed in melanoma cancer types, and this mutational signature set may be used for matching in step (420) of process (400) shown in Figure 4. Visual comparison of the SBS7a and DBS1 mutational signatures with the mutational catalog of the melanoma tumor sample confirms that the mutational profile of the SBS7a and DBS1 signatures contains relevant genetic mutations for characterizing the tumor sample.

[0039] In some embodiments, the number of signatures used in the matching process is between 2 and 5. In some cases, using additional signatures may affect the results of the matching process, for example, by over-fitting.

[0040] Returning to process (400) shown in FIG. 4, after matching is completed in step (420), process (400) proceeds to step (422), where it may be determined whether the sample's exposure to particular signatures determined during the matching process exceeds a threshold. If step (422) determines that the sample has an exposure to at least one signature that exceeds the threshold, process (400) proceeds to step (424), where a context (e.g., trinucleotide context) is added to each variant in the set of variants. This context may be used to assign a context probability (step 430) from a relevant subset of the signatures identified in step 422 (step 426), such that each variant is annotated with the likelihood that it is caused by exposure to one of the cancer mutation signatures used in the matching process. For example, in the example urothelial cancer sample in Figure 6, the SBS13 signature C>A (TCC context) has a lower proportion than the C>G (TCA context). Therefore, it can be inferred that a C>G (TCA) change in the mutation catalog is more likely to be associated with urothelial cancer than a C>A (TCC) change. In some embodiments, the context probability may be determined as the sum of the associated exposures. For example, if it is determined that 60% of C>T (TCA) variants are the result of exposure to SBS2 and 40% are the result of exposure to SBS13, the context probability for C>T (TCA) may be calculated as the sum of 0.6*C>T (TCA) frequency in SBS2 and 0.4*C>T (TCA) frequency in SBS13, adjusting for the determined exposure to each cancer mutation signature.

[0041] After assigning context probabilities to variants in step (430), process (400) proceeds to step (432), in which variants are sampled based on their assigned context probabilities to generate a set of genetic variants in step (442) that are most likely to be associated with the tumor sample. In particular, sampling may generate genetic variants that correspond to peaks in the mutational signature. Alternatively, genetic variants may be ranked based on context probabilities, and the top numbers (e.g., 6, 12, 16, 20, 48, 96) may be selected for inclusion in the set of variants determined in step (442). By considering mutational processes likely involved in tumor development and allowing probabilistic association of variants to these processes, the selected variants are more likely to be true somatic alterations, as opposed to variants introduced by artifactual processes.

[0042] If it is determined in step (422) that the matching process did not identify any exposures above the threshold, process (400) may proceed to step (440), in which the variants contained in the VCF file are instead ranked based on one or more quality criteria. For example, variants may be ranked based on coverage depth in normal or healthy samples, the number of reads between tumor and normal samples, allele frequency, the probability that the variant is a germline variant, etc. Additional examples of such quality criteria may be found in International Patent Publication WO2022029688A1, the contents of which are incorporated herein by reference. The ranked variants may then be used to determine a set of genetic variants in step (442) that can be used to create a tumor-informed assay.

[0043] In some embodiments, variants may also be weighted based on their likelihood of being driver mutations (i.e., mutations likely associated with causing cancer) or passenger mutations. The latter may be particularly useful for identifying minimal residual disease because they are less likely to be eliminated by targeted treatment, and therefore may be even more useful for identifying resistant clones that have lost driver mutations. In some embodiments, variants may also be weighted or ranked based on other characteristics, such as clonality (i.e., clonal or subclonal), mappability, or likelihood of being amplified, as further described in International Patent Publication WO2022029688A1.

[0044] After the set of genetic variants is determined in step (442) (e.g., by ranking the variants in step (440) or sampling variants based on contextual probability in step (432)), process (400) proceeds to step (444), where the set of variants may be used to create a tumor-informed assay. The tumor-informed assay may be designed to preferentially amplify the set of variants in subsequent plasma samples to detect residual disease. Examples of tumor-informed assays can be found, for example, in International Patent Publication WO2022-029688A1 and U.S. Patent Publication 2020 / 0157604, the contents of which are incorporated herein by reference.

[0045] While the techniques described herein may prioritize cancer-associated variants based on context (e.g., by sampling variants based on contextual probability in step (432) of process (400)), some artifactual variants may still be selected for incorporation into the assay (although typically may be of lower priority than cancer-based variants). In such cases, it may be advantageous to specifically filter out all variants that are not associated with any previously seen cancer signatures, since genetic variants not previously seen in any known cancer mutational signatures are unlikely to be observed in the mutation catalog. Filtering of variants likely to be artifactual in the mutation catalog may be performed by comparing each variant in the mutation catalog to various known cancer signatures, for example, by cosine similarity. In some embodiments, all available mutational signatures (e.g., not limited to a particular cancer type) may be used to filter artifactual variants, since filtering may determine whether a given mutational profile resembles any previously observed mutational process (e.g., signature). Any variants that are not similar to previously observed signatures are likely the result of some other process that introduces errors, which may be due to sample preparation or amplification, for example. Returning to the urothelial cancer sample example in Figure 6, the artifactual variant filtering process is unlikely to exclude C->T and C->G mutations, because these mutation profiles are present in the SBS2 and SBS13 mutation signatures, respectively. However, other variants (e.g., T->C mutations) may be excluded if they are not sufficiently similar to some other mutation processes observed in the mutation signature.

[0046] In some embodiments, rather than analyzing all variants in the mutation catalog, samples showing an excess of mutations (e.g., greater than 50%, greater than 60%, greater than 70%, greater than 80%) for a particular mutation type (e.g., C>T, including all contexts) may be flagged for further analysis. The mutation profile for a given flagged mutation may be compared to all (or a subset) of mutation signatures that also show a significant ratio for this mutation type, using, for example, cosine similarity, an example of which is shown in FIG. 8. If the mutation profile is very low similarity to any signature, the mutation type may be removed from further consideration for inclusion in the final set of genetic variants. As shown in the example of FIG. 8, samples with an excess of C>T mutations may be flagged and then compared pairwise with multiple mutation signatures. Using cosine similarity, the mutation profile for the C>T mutation may be determined to be sufficiently similar to the corresponding mutation profiles in the SBS1, SBS2, and SBS7a mutation signatures (e.g., using cosine similarity and a threshold of 0.85), and therefore, this variant type may be retained for further consideration in the variant selection process.

[0047] In some embodiments of the present disclosure, artifactual variants in a mutation catalog may be identified and removed by comparing the variants to "artifactual signatures" generated from the mutation catalog corresponding to samples containing a large number of artifactual mutations. For example, mutations introduced by the formalin-fixed, paraffin-embedded (FFPE) process can be identified and removed based on their similarity to FFPE mutation signatures. Figure 9 shows an example of an artifactual signature for FFPE-based mutations. Such signatures may be used to identify and remove specific variants in the mutation catalog that are likely due to the FFPE process rather than exposure to the mutation process for a particular type of cancer associated with the sample. Similarly, artifactual signatures can be generated from samples with PCR amplification errors or other types of errors to provide a filter for removing specific variants from the mutation catalog.

[0048] In some embodiments, the mutational signature set includes at least one mutational signature associated with exposure to a treatment, such as chemotherapy. Some mutational exposures occur earlier in cancer development than others; therefore, some signatures are likely to represent clonal mutations, while other signatures are likely to be subclonal mutations (i.e., signatures associated with earlier exposures are likely to be clonal, while signatures associated with certain later events, such as chemotherapy, will be subclonal). Using this information, it is possible to select variants that are more likely to be clonal. This is beneficial, for example, when a patient has already received some treatment (e.g., a breast cancer sample obtained after neoadjuvant therapy). Mutations from a tumor can be prioritized based on the signature by selecting variants that a) are more likely to be clonal and b) are more likely to be true somatic changes. Similarly, variants may be excluded if they are more likely to be clonal.

[0049] Having thus described several aspects and embodiments of the technology described in this disclosure, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art.

[0050] Such changes, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. For example, those skilled in the art will readily conceive of various other means and / or structures for performing the functions and / or results and / or obtaining one or more of the advantages described herein. Each of such variations and / or modifications is deemed to be within the scope of the embodiments described herein. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Accordingly, the above-described embodiments are presented by way of example only, and within the scope of the appended claims and their equivalents, it is to be understood that embodiments of the invention may be practiced otherwise than as specifically described herein. Furthermore, any combination of two or more features, systems, articles, materials, kits, and / or methods described herein, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.

[0051] The above-described embodiments may be implemented in any of numerous ways. One or more aspects and embodiments of the present disclosure involving the execution of a process or method may utilize program instructions executable by a device (e.g., a computer, processor, or other device) to perform or control the execution of the process or method. In this regard, various inventive concepts may be embodied as a computer-readable storage medium (or multiple computer-readable storage media) (e.g., computer memory, one or more hard drives, flash memory, circuitry in a field programmable gate array or other semiconductor device, or other tangible computer storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods implementing one or more of the various embodiments described above. The one or more computer-readable media may be transportable such that the stored program(s) can be loaded into one or more different computers or other processors to implement the various aspects described above. In some embodiments, the computer-readable medium may be non-transitory.

[0052] The terms "program" or "software" are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be used to program a computer or other processor to implement various aspects as described above. Furthermore, it should be understood that, according to one aspect, one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but may be distributed modularly among many different computers or processors to implement various aspects of the present disclosure.

[0053] Computer-executable instructions may take many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.

[0054] Additionally, data structures may be stored on computer-readable media in any suitable format. For simplicity of explanation, data structures may be depicted as having fields that are related through their locations within the data structure. Such relationships may also be achieved by allocating storage for the fields with locations in the computer-readable media that convey the relationship between the fields. However, any suitable mechanism may be used to establish relationships between information within fields of a data structure, including the use of pointers, tags, or other mechanisms for establishing relationships between data elements.

[0055] The above-described embodiments of the present technology may be implemented in any of numerous ways. For example, embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be executed on any suitable processor or collection of processors, whether provided on a single computer or distributed among multiple computers. It should be understood that any component or collection of components performing the above functions may generally be considered a controller that controls the above functions. The controller may be implemented in numerous ways, such as dedicated hardware or general-purpose hardware (e.g., one or more processors) programmed with microcode or software to perform the above functions, and may be implemented in a combination of multiple ways if the controller corresponds to multiple components of a system.

[0056] Furthermore, it should be understood that a computer may be embodied in any of a number of forms, such as, by way of non-limiting example, a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer may be embedded in a device not generally considered a computer, but having suitable processing capabilities, including a personal digital assistant (PDA), a smartphone, or any other suitable portable or fixed electronic device.

[0057] A computer may also have one or more input / output devices. These devices may be used, among other things, to present a user interface. Examples of output devices that may be used to provide a user interface include a printer or display for visual presentation of output and a speaker or other sound-generating device for audible presentation of output. Examples of input devices that may be used for a user interface include a keyboard and pointing devices such as a mouse, touchpad, and digitizing tablet. As another example, a computer may receive input information via voice recognition or in other audible formats.

[0058] Such computers may be interconnected by one or more networks in any suitable form, including local area networks or wide area networks, such as enterprise networks, and intelligent networks (IN) or the Internet. Such networks may be based on any suitable technology and operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.

[0059] Also, as described, some aspects may be implemented as one or more methods. Steps performed as part of a method may be ordered in any suitable manner. Thus, embodiments may be constructed such that steps are performed in an order different from that shown, which may include some steps being performed simultaneously even though shown as sequential steps in an exemplary embodiment.

[0060] All definitions defined and used herein should be understood to supersede dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0061] The indefinite articles "a" and "an," as used in the specification and claims, unless the context clearly indicates otherwise, should be understood to mean "at least one."

[0062] The phrase "and / or," as used in the specification and claims, should be understood to mean "either or both" of the elements so conjugated, i.e., elements that are sometimes present conjointly and other times present separately. Multiple elements listed with "and / or" should be construed in the same manner, i.e., to mean "one or more" of the elements so conjugated. Other elements other than those specifically identified by the "and / or" clause may optionally be present, whether related to the elements specifically identified. Thus, as a non-limiting example, a reference to "A and / or B," when used in conjunction with open-ended language such as "comprising," can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements), etc.

[0063] As used in this specification and claims, the phrase "at least one," referring to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed in the list of elements, nor excluding any combination of elements in the list of elements. This definition also allows for the optional presence of elements other than those specifically identified in the list of elements to which the phrase "at least one" refers, whether related to the identified elements or not. Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B" or, equivalently, "at least one of A and / or B") can refer in one embodiment to no B (and optionally including elements other than B), at least one, optionally more than one A; in another embodiment to no A (and optionally including elements other than A) and at least one, optionally more than one B; in yet another embodiment to at least one, optionally more than one A and at least one, optionally more than one B (and optionally including other elements); etc.

[0064] Also, the phrasing and terminology used herein is for purposes of description and should not be regarded as limiting. The use of "including," "comprising," "having," "containing," "involving," and variations thereof, is meant to encompass the items listed thereafter and equivalents thereof, as well as additional items.

[0065] In the claims, as well as in the above specification, all transitional phrases such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "composed of," and the like, shall be understood to be open-ended, i.e., meaning including, but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively.

[0066] The use of counters such as "first," "second," and "third" in the claims to modify claim elements does not in itself imply any priority, precedence, or order of one claim element over another, or any chronological order in which method steps are performed; rather, the counters are used to distinguish claim elements merely as labels to distinguish one claim element having a particular name from another element having the same name (excluding the use of the counter).

Claims

1. receiving a sample taken from a patient, said sample being related to a cancer type; generating a mutation catalog for the sample, the mutation catalog indicating the proportion of genetic variants observed in the sample; selecting a signature set associated with said cancer type, said signature set comprising one or more signatures, each signature comprising a mutation profile; determining a set of genetic variants that are most likely to be true somatic variants associated with the sample based on the signature set and the mutation catalogue associated with the cancer type; outputting the set of genetic variants for use in creating a tumor-informed assay for the patient; A method for selecting genetic variants for a tumor-informed assay, comprising:

2. selecting a signature set associated with the cancer type comprises accessing a database configured to store a plurality of signatures associated with the cancer type; 10. The method of claim 1, wherein the plurality of signatures is a population-level signature determined from a plurality of cancer samples.

3. 3. The method of claim 2, further comprising including in said signature set only those signatures from said database that have one or more genetic variants that account for at least 5% of all genetic variants in said variant profile.

4. 4. The method of claim 3, further comprising including in said signature set only those signatures from said database that have one or more genetic variants that account for at least 10% of all genetic variants in said variant profile.

5. 5. The method of claim 4, further comprising including in said signature set only those signatures from said database that have one or more genetic variants that account for at least 25% of all genetic variants in said variant profile.

6. 6. The method of claim 1, wherein determining a set of genetic variants comprises matching the mutation profile in the signature set to the mutation catalogue to determine the corresponding amount of each signature observed in the sample.

7. 7. The method of claim 6, wherein determining the set of genetic variants further comprises determining whether the amount observed of each signature in the sample exceeds a threshold.

8. determining the set of genetic variants for each of the genetic variants in the sample, associating a context probability based on the frequency of the genetic variant in the signature set and the determined amount that each signature was observed in the sample; and Sampling the genetic variants weighted by associated context probabilities to determine which genetic variants to include in the set of genetic variants. The method of claim 6 further comprising:

9. The method of claim 8 , wherein the context probability is based on a trinucleotide context.

10. filtering the set of genetic variants based at least in part on the suitability of each of the genetic variants in the set of genetic variants for use in the tumor-informed assay; generating the tumor-informed assay based on the set of genetic variants in the filtered set of genetic variants; The method of any one of claims 1 to 9, further comprising:

11. The method of any one of claims 1 to 10, wherein the set of signatures associated with the cancer type comprises at least two signatures.

12. wherein selecting the signature set associated with the cancer type further comprises selecting a mutation profile associated with a source of error introduced during a sequencing process, the method comprising:

12. The method of any one of claims 1 to 11, further comprising the step of removing from the set of genetic variants all genetic variants determined to result from said error source.

13. 13. The method of claim 12, wherein the source of error is associated with a formalin-fixed paraffin-embedded (FFPE) process, an amplification process, or a sequencing process.

14. wherein selecting the signature set associated with the cancer type further comprises selecting a mutation profile associated with a treatment, the method comprising:

14. The method of any one of claims 1 to 13, further comprising the step of excluding from the set of genetic variants all genetic variants determined to be attributable to the treatment.

15. 15. The method of claim 14, wherein the treatment is chemotherapy.

16. 16. The method of any one of claims 1 to 15, wherein generating a mutation catalogue for the sample comprises performing whole exome sequencing on the sample.

17. The method of any one of claims 1 to 16, wherein the signature set comprises at least one double base substitution signature.

18. 18. The method of any one of claims 1 to 17, wherein at least some of the mutation profiles are associated with different exposure types.

19. 19. The method of any one of claims 1 to 18, wherein the mutation profile indicates the proportion of genetic variants observed in a population.

20. 20. The method of claim 19, wherein the genetic variant comprises a trinucleotide context.

21. The method of any one of claims 1 to 20, wherein the signature is indicative of exposure to a mutational process.

22. 22. The method of claim 21, wherein the mutational process is associated with cancer.

23. receiving a sample collected from a patient; generating a mutation catalogue from the set of genetic variants, said mutation catalogue indicating the proportion of genetic variants observed in said sample; selecting a signature set, the signature set comprising one or more signatures, each signature comprising a mutation profile; determining a set of genetic variants; removing from said set of genetic variants all genetic variants determined to result from said one or more signatures; outputting the set of genetic variants for use in generating an assay for the patient; A method for selecting genetic variants, comprising:

24. 24. The method of claim 23, wherein the signature set is associated with an error source.

25. 26. The method of claim 25, wherein the source of error is associated with a formalin-fixed paraffin-embedded (FFPE) process, an amplification process, or a sequencing process.

26. The method of any one of claims 23 to 25, wherein the signature set is therapeutically relevant.

27. 27. The method of claim 26, wherein the treatment is chemotherapy.

28. 27. The method of claim 26, wherein the excluding step further comprises excluding from the set of genetic variants all genetic variants determined to be subclonal genetic variants relevant to the treatment.

29. The signature set further comprises at least one mutation profile associated with a cancer type, and the determining step comprises:

29. The method of any one of claims 23-28, further comprising determining the set of genetic variants associated with the sample based on at least one mutation profile associated with the cancer type and the mutation catalogue.

30. A system comprising at least one hardware computer processor programmed to carry out any one of the methods according to claims 1 to 29.

31. A computer readable medium encoded with a plurality of instructions which, when executed by at least one hardware computer processor, performs the method of any one of claims 1 to 29.

32. 30. A tumor-informative assay for monitoring the presence of cancer in a patient configured to detect at least one genetic variant from an output set of genetic variants according to the method of any one of claims 1 to 29.