Clustered Mutations for Cancer Treatment
Patent Information
- Application Number
- JP2024535241
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-12-14
- Filing Date
- 2022-12-13
- Publication Date
- 2025-11-17
AI Technical Summary
Existing cancer treatments lack comprehensive analysis of clustered mutations, particularly clustered indels and driver mutations, which are crucial for understanding cancer progression and determining treatment strategies.
A method for treating cancer by analyzing clustered mutations in genes such as TP53, EGFR, KIT, KMT2C, ELF3, APC, and ARID1A, and BRAF, and administering targeted therapies based on the presence or absence of these mutations to inhibit cancer cell growth or select appropriate treatment strategies.
This approach allows for personalized cancer treatment by identifying patients with better overall survival outcomes and tailoring therapies to their specific mutation profiles, potentially reducing treatment aggressiveness and improving patient outcomes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Application No. 63 / 289,601, filed December 14, 2021, the contents of which are incorporated herein by reference in their entirety. [Background technology]
[0002] background Throughout this disclosure, technical and patent literature is referenced. In some aspects, the literature is referenced by Arabic numerals, and its full bibliographic citation appears immediately before the claims. These disclosures are provided to describe the state of the art to which this disclosure pertains, and are incorporated herein by reference.
[0003] The genomes of cancer cells contain somatic mutations imprinted by the activity of various mutational processes 1,2 Although most single base substitutions, as well as small insertions and deletions (indels), are scattered independently across the genomic landscape, a subset of substitutions and indels tend to cluster together. 3,4 This clustering is due to, among other things, a combination of heterogeneous mutation rates across the genomic landscape, biophysical signatures of exogenous carcinogens, dysregulation of endogenous processes, and the occurrence of larger events associated with genomic instability. 4~16 Previous analyses of clustered mutations have focused on single-base substitutions, not double and multi-base substitutions. 1、9、16~19 , widespread hypermutation called omikli 15 , and a longer event called kataegis 12、14、16、20 We uncovered several classes of clustering events, including kataegic events, which are usually defined as sharing the same strand and reference allele. 2、16 Previous studies have identified nine clustered mutational signatures of different mutational processes. 4, as well as clustered driver replacement by APOBEC3-associated mutagenesis 15 or oncogenic POLH mutagenesis 4 Clustered driver substitutions by . To the applicant's knowledge, an analysis of clustered indels or a comprehensive survey of clustered driver mutations has never been performed. Summary of the Invention
[0004] Disclosure Summary Double base substitutions have been extensively investigated, revealing multiple intrinsic and extrinsic mutational processes that can cause these events, including failure of DNA repair pathways and exposure to environmental mutagens. 1、2、16、17 In contrast, multi-base substitutions have not been comprehensively investigated, probably due to their low number in most cancer genomes. Moreover, only a few reported processes are related to omikli and kataegic events, and the majority of these processes are attributed to the AID / APOBEC3 family of deaminases. 4、5、12、14~16、21~24 For example, in B cell lymphomas, clustered tracks of C>T and C>G mutations in the WRCY motif are the result of direct duplication to the AID lesion. 21 Alternatively, AID-induced lesions can be processed by the mismatch repair pathway that recruits the error-prone DNA polymerase η, resulting in non-canonical AID mutations. 21 In addition to AID, APOBEC3 enzymes are typically involved in antiviral responses and restricting the mobility of mobile elements. 25~31 , which is a substantial source of clustered mutational events 2、4、12、14~16、24、32 Specifically, APOBEC3 enzymes require single-stranded DNA as a substrate, resulting in omikli and kataegis. 14、15、24、32 Omikli was found to be enriched in regions of early replication and prevalent in microsatellite-stable tumors, indicating a role for mismatch repair in exposing short stretches of single-stranded DNA while processing mismatched bases during replication. 15Furthermore, differential activity of mismatch repair in gene-rich regions leads to increased mutational burden of omikli mutations in cancer driver genes. 15 Kataegis are less common than omikli because they may rely on longer tracks of single-stranded DNA. 12~14 Such tracks are typically available during repair of double-strand breaks, with the majority of kataegis observed within 10 kb of the detected breakpoint. 11 .
[0005] Amplification of known oncogenes through double-strand breaks and complex rearrangements is known to drive tumorigenesis in many cancer types. 33 Recent studies have elucidated high copy number states of circular extrachromosomal DNA (ecDNA), which often contain known oncogenes, and are found in most human cancers. 33~36 The circular nature of ecDNAs and their rapid replication patterns mimic double-stranded DNA viral pathogens and represent a potential substrate for APOBEC3 mutagenesis, which may ultimately contribute to the subclonal diversification of ecDNA-harboring tumors through accelerated diversification of extrachromosomal oncoproteins.
[0006] Described herein is a comprehensive examination of clustered substitutions and clustered indels across 2,583 cancer genomes across 30 different tumor types. Results elucidate numerous mutational processes that give rise to clustered mutations, including clustered driver mutations associated with altered differential gene expression and overall survival, and reveal recurrent APOBEC3 mutagenesis, termed kyklonas, that drive ecDNA evolution.
[0007] The application of these findings is further provided herein. In one embodiment, a method of treating the inhibition of cancer cell growth or a method of treating cancer in a subject in need of cancer treatment is disclosed, wherein the subject has one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A in a sample isolated from the subject, or lacks clustered mutations in the BRAF gene. The method comprises, consists of, or consists essentially of administering an aggressive treatment to the subject, thereby inhibiting the growth of cancer cells in the subject or treating cancer in the subject. If the subject does not have clustered mutations in one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A, or lacks clustered mutations in the BRAF gene, a less aggressive treatment can be administered.
[0008] In yet another embodiment, a method for selecting cancer patients for active treatment is disclosed.The method comprises, consists of, or essentially consists of, assaying and / or detecting at least one clustered mutation in the gene selected from TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A, and / or no clustered mutation in BRAF gene in the sample isolated from the subject, and when one or more clustered mutations are found in TP53, EGFR, KIT, KMT2C, ELF3, APC and / or ARID1A in the sample isolated from the cancer patient, and / or when no clustered mutation in BRAF gene is detected, the subject is selected for treatment.
[0009] Cancer patients determined to have a mutation with a better predicted outcome, such as longer overall survival, may in one embodiment receive treatment, but may choose a less aggressive treatment for initial or subsequent treatments.
[0010] In yet another embodiment, a method for identifying whether a cancer patient is likely to experience a longer overall survival or a shorter overall survival is disclosed.This method comprises, consists of, or consists essentially of, assaying and / or detecting at least one clustered mutation in the gene selected from TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A or BRAF gene in the sample isolated from the patient, and when clustered mutation is detected in BRAF or when clustered mutation is not detected in one clustered mutation in the gene selected from TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A, the patient may experience a longer overall survival, and when clustered mutation is detected in one clustered mutation in the gene selected from TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A or when clustered mutation is not detected in BRAF gene, the patient may experience a shorter overall survival. BRIEF DESCRIPTION OF THE DRAWINGS [Brief description of the drawings]
[0011] [Figure 1]Figure 1A-B show the landscape of clustered mutations across human cancers. Figure 1A: Pan-cancer distribution of clustered substitutions subclassified into double base substitutions (DBS), multi-base substitutions (MBS), omikli, kataegis and other clustered mutations. Top panel: Each black dot represents a single cancer genome. Grey bars reflect the clustered median tumor mutation burden (TMB) for the cancer type. Center panel: Clustered TMB normalized to genome-wide TMB, which reflects the contribution of clustered mutations to the overall TMB of a given sample. Grey bars reflect the median contribution of the cancer type. Bottom panel: Proportion of each subclass of clustered events for a given cancer type with the total number of samples with at least one clustered event relative to the total number of samples in a given cancer cohort. Figure 1B: Pan-cancer distribution of clustered small insertions and deletions. Top and center panels have the same information as the panels in Figure 1A. Bottom panel: Proportion of each cluster type of indels for a given cancer type, with the total number of samples with at least one clustering indel relative to the total number of samples in a given cancer cohort. All 2,583 whole-genome sequenced samples from the PCAWG were included in the analysis, but cancers with fewer than 10 samples were removed from the main plot and included in Figure 6D.
[0012] [Diagram 2] Figure 2 shows the mutational processes underlying the clustering events. Each circle represents the activity of a signature for a given cancer type, the radius of the circle determines the proportion of samples with more than a given number of mutations specific to each subclass, and the color reflects the median number of mutations per cancer type. A minimum of two samples per cancer type is required for visualization.
[0013] [Diagram 3]Figures 3A-G show a panorama of clustered driver mutations in human cancers. The percentage of clustered mutations (top) compared to the percentage of clustered driver events (bottom) for substitutions in Figure 3A and indels in Figure 3B. Figure 3C: Frequency of clustered driver events across known cancer genes. The radius of the circle is proportional to the number of samples with clustered driver mutations in the gene, and the greyscale reflects the clustered mutation burden. All clustered driver events are classified into one of five clustered classes, and the number of clustered driver substitutions and the total number of driver substitutions are shown on the right. Figure 3D: Clustered indel drivers are displayed as in Figure 3C. Figure 3E: Odds ratios of clustered substitutions (top) and clustered indels (bottom) resulting in deleterious changes (n=192 clustered substitutions; n=54 clustered indels) or synonymous changes (n=5 clustered substitutions; n=5 clustered indels) within a given driver gene compared to non-clustered driver mutations (n=771 deleterious substitutions and n=237 synonymous substitutions; n=111 deleterious indels and n=50 synonymous indels). All events were overlaid with the PCAWG consensus list of driver events and annotated using VEP. Odds ratios are shown with 95% confidence intervals. Figure 3F: Kaplan-Meier survival curves comparing outcomes of samples with clustered mutations vs. non-clustered mutations in BRAF (top), TP53 (middle) and EGFR (bottom) across the TCGA cohort. Only cohorts with >5 samples with clustered mutations within a given gene were included. Figure 3G: Kaplan-Meier survival curves comparing outcomes of samples with clustered versus non-clustered mutations in the same gene across the MSK-IMPACT cohort. Log10(hazard ratios) are shown in Figure 3F and Figure 3G with their 95% confidence intervals. Cox regressions were adjusted for age (TCGA only), mutation burden and cancer type (see experiment number 1 below). In Figures 3A, 3B and 3E, P values were calculated using two-tailed Fisher's exact test and corrected for multiple hypothesis testing.
[0014] [Figure 4]Figure 4A-4F show kataegic events co-localized with most forms of structural variation. Figure 4A: Proportion of all kataegic events per cancer type consistent with distinct amplifications or structural variations. Figure 4B: Distance to the nearest breakpoint of all kataegic mutations (turquoise), kyklonas (gold) and non-clustered mutations (red). Kataegic distances were modeled as Gaussian mixtures with three components (blue line). Figure 4C: Left: Volcano plot showing samples statistically enriched for kyklonas (red; q-values from FDR-corrected Z-test). Center left: Proportion of samples with ecDNA co-occurring with kataegis. Center right: Mutational spectrum of all kyklonas. Right: Proportion of kyklonic events attributable to SBS2 and SBS13. Cosine similarity was calculated between kyklonic and reconstructed spectra constructed using SBS2 and SBS13 (p-values from Z-score test). Figure 4D: Rainfall plot showing the IMD distribution for a given sample with genomic location of the ecDNA breakpoints (grayscale). Figure 4E: YTCA enrichment vs. RTCA enrichment per sample with kyklonas, where YTCA and RTCA enrichment suggest higher APOBEC3A or APOBEC3B activity, respectively. Genetic mutations were separated into transcriptional mutations (template strand) and coding mutations. RTCA / YTCA fold enrichment was compared to that of non-clustered mutations. Figure 4F: Relative expression of APOBEC3A and APOBEC3B in samples with ecDNA with kyklonas (n=59) compared to samples without ecDNA (n=1,364) (left) and in samples without kyklonas (n=98) compared to samples without ecDNA (n=1,364). Expression values were normalized using upper quartile normalization obtained from FPKM and PCAWG emission. P values in Figures 4E and 4F were obtained using a two-tailed Mann-Whitney U test and are FDR corrected using the Benjamini-Hochberg procedure. For each box plot, the center line reflects the median, the lower and upper limits of the box correspond to the first and third quartiles, and the lower and upper whiskers extend from the box by 1.5 times the interquartile range.
[0015] [Diagram 5] Figure 5A-E show recurrent APOBEC3 hypermutations in ecDNA. Figure 5A: Number of clustered events consistent with a single amplicon or structural variation (SV) event; each dot represents an amplicon or SV (n=84 circular; n=275 linear; n=111 severely rearranged; n=62 BFB; and n=11,139 SV). A 10 kb window was used to determine co-occurrence of kataegis with SV breakpoints (**q-value <0.01; ****q-value <0.0001). Figure 5B: Left: Normalized distribution of variant allele frequency (VAF) for all clustered mutations except kataegis, all non-ecDNA kataegis and kyklonas, as shown in grayscale. Right: Normalized VAF distribution for kyklonic-ecDNA with oncogene and kyklonic-ecDNA without oncogene. Figure 5C: All kataegis and kyklonas repeat frequencies using a 10 Mb sliding genomic window. Figure 5D: Number of kyklonic events and kyklonic mutations per ecDNA region containing (n=137) or not containing (n=134; left and right, respectively). Figure 5E: Total number of clustered and kataegic mutations found in samples with ecDNA containing oncogenes (n=67 samples) compared to samples with ecDNA not containing oncogenes (n=44; left and right, respectively). P values in Figure 5A, Figure 5D, and Figure 5E were obtained using a two-tailed Mann-Whitney U test and are FDR corrected using the Benjamini-Hochberg procedure. For data represented as box plots in Figure 5A, Figure 5D, and Figure 5E, the center line reflects the median, the lower and upper limits of the box correspond to the first and third quartiles, and the lower and upper whiskers extend from the box by 1.5 times the interquartile range (IQR).
[0016] [Figure 6-1]Figures 6A-K show the identification and clinical association of clustered events. Figure 6A: Schematic for isolating clustered mutations in samples. Figure 6B: Subclassification of clustered substitutions and clustered indels. Expected IMD obtained using steps 2 and 3 (panel a). Figure 6C: Distribution of indels present in a single clustered event. Figure 6D: Distribution of clustered substitutions (left) and clustered indels (right) across cancers with less than 10 samples subdivided into different categories. Figure 6E: Correlation between tumor mutation burden (TMB) of each sample, TMB in the exome, or TMB (left) and indels (right) for each class of clustered substitution. Figure 6F: Distribution of variant allele frequencies for all clustered substitution classes (left; DBS: 1,215 samples; MBS: 851; omikli: 1,466; kataegis: 1,108; others: 335) with average fold enrichment (right) compared to non-clustered mutations. For each box plot, the center line reflects the median, the lower and upper limits correspond to the first and third quartiles, and the lower and upper whiskers extend from the box by 1.5 times the interquartile range (IQR). Figure 6G: Kaplan-Meier curves between samples with high (top 80th percentile) and low (bottom 20th percentile) amounts of clustered substitutions (left) or indels (right) in PCAWG ovarian cancer. Figure 6H: Cox regression performed on PCAWG cancer types while correcting for age (n=20 top clustered substitutions and n=21 bottom clustered substitutions; n=49 top clustered indels and n=49 bottom clustered indels). Figure 6I: Kaplan-Meier survival curves for TCGA cancer types with different patient outcomes associated with the detection of any clustered mutations. Cox regression was performed on TCGA samples (OV: n = 111 top clustered substitutions, n = 159 bottom clustered substitutions; UCEC: n = 322 top, n = 64 bottom; ACC: n = 24 top, n = 67 bottom) while adjusting for age (Figure 6J) and total mutational burden (Figure 6K). PCAWG ovarian cancer was included in k). Measurement centers for each Cox regression reflect log10(hazard ratio) with 95% confidence intervals in Figure 6H-Figure 6K. [Figure 6-2] Same as above.
[0017] [Figure 7] Figures 7A-7E show the mutation process of clustered driver events. Figure 7A: Proportion of clustered driver substitutions and indels within each cancer type. All 2,583 whole genome sequenced samples from PCAWG with detected driver events are included, but cancer types with less than 10 samples are not shown. Figure 7B: Proportion of clustered driver mutations per cancer gene compared between cancer genes (n=19 genes) vs. tumor suppressor genes (n=30 genes) and genes with high number of isoforms (n=17) vs. genes with low number of isoforms (n=23; top and bottom quartiles of isoforms across all cancer drivers). Figure 7C: Proportion of clustered driver mutations for a given subclass per cancer gene compared between cancer genes (n=17 genes with clustered substitutions and n=13 genes with clustered indels) vs. tumor suppressor genes (n=28 genes with clustered substitutions and n=70 genes with clustered indels). Figure 7D: Relative expression of driver genes with clustered events versus non-clustered events. All expression values were normalized using FPKM normalization and upper quartile normalization obtained from the official PCAWG emission, and then normalized using the mean expression of wild-type genes. A value of 1 (dashed line) reflects no difference in expression compared to wild-type genes. Figure 7E: Proportional activity of mutation signatures contributing to clustered driver events within each subclass. Multi-base substitutions (MBS) did not contribute to the reported driver events. For the analyses in b-d), p-values were generated using a two-tailed Mann-Whitney U test (*P<0.05; p=0.03 for STAT6; p=0.04 for CTNNB1; p=0.02 for BTG1). For each box plot, the center line reflects the median, the lower and upper limits of the box correspond to the first and third quartiles, and the lower and upper whiskers extend from the box by 1.5 times the interquartile range (IQR).
[0018] [Figure 8-1] Figures 8A-E show the repeat mutagenesis and functional effect of kyklonas. Figure 8A: Total number of repeat-mutated ecDNAs shown as a percentage of the total number of ecDNAs with kyklonas for a given cancer type. The total number of ecDNAs containing kyklonas is displayed above each bar graph for each cancer type. All ecDNAs with repeat hypermutation were considered enriched for kyklonic events after correcting for multiple hypothesis testing (Z-score test; q-value <0.05). Figure 8B: Percentage of samples with ecDNAs exclusively divided into samples with co-occurring kataegis, samples with no kataegis overlap, and samples with no kataegis detected across the genome. The number of samples included in each cancer type is listed. For certain cancer types, as little as one sample may represent the entire proportional breakdown (e.g., bone-osteosarcoma or bone-epithelial). Figure 8C: A single sarcoma genome and (Figure 8D) a single head squamous cell carcinoma genome show matching ecDNA regions and kataegis displayed as rainfall (top left) with a single zoomed-in ecDNA represented using a circos plot (top right). Bottom: Two regions of ecDNA with matching kyklonic events. Variant allele frequencies are shown per event (orange). Figure 8E: Kyklonic substitutions resulting in recurrent coding mutations within known cancer genes. [Figure 8-2] Same as above.
[0019] [Figure 9]Figures 9A-B show the determination of the number of mutations that distinguish omikli from kataegis. Figure 9A: Modeling the number of mutations per event using a mixture of two Poisson distributions. The first component, representing omikli, has an average IMD of 2.1, and the second component, representing kataegis, has an average IMD of 4.4. The estimated contribution of mutations in each component is shown as a bar for each corresponding event size. Figure 9B: Distribution of IMD per event across events of different sizes (n=199,912 events with 2 mutations; n=35,576 events with 3 mutations; n=15,320 events with 4 mutations; n=9,613 events with 5 mutations). The chosen cutoff between omikli and kataegis was 4 mutations. For each box plot, the midline reflects the median, the lower and upper limits of the box correspond to the first and third quartiles, and the lower and upper whiskers extend from the box by 1.5 times the interquartile range (IQR).
[0020] [Figure 10] Figure 10 shows the distribution of low-confidence clustered indels. The number of clustered indels that fall within regions of the genome with low mapping scores comprises approximately 1% of all clustered indels. Within these 1% of mutations with low mapping scores, only 30% of events have an intermutation distance of less than 10 (0.3% of all clustered indels), while lbp indels that fall within low mapping regions comprise only 0.5% of all clustered indels.
[0021] [Figure 11-1] 11A-11B are Kaplan-Meier survival curves comparing outcomes of samples with clustered versus non-clustered mutations in the same gene across the MSK-MET cohort, composed of targeted sequencing from both primary (FIG. 11A) and metastatic (FIG. 11B) cancers. Log10 transformed hazard ratios (log10(HR)) are shown with 95% confidence intervals. Cox regressions were adjusted for age, tumor mutation burden, and sex. Each comparison was obtained using a single cancer type. [Figure 11-2]Same as above.
[0022] [Figure 12] Figure 12A-12B show de novo signatures of double base (DBS) and multi-base (MBS) signatures. Figure 12A: Activity of DBS de novo signature (top) and corresponding signatures extracted from prostate, skin, stomach and uterine cancers that could not be accurately reproduced using known COSMIC mutation signatures (bottom). Figure 12B: Activity of MBS de novo signature (top) and corresponding signatures extracted from colon, esophagus and head and neck cancers that could not be accurately reproduced using known COSMIC mutation signatures (bottom).
[0023] [Figure 13]Figures 13A-13D show experimental validation and epidemiological associations of clustered mutational processes. Figure 13A: Experimental validation of three omikli processes. Specifically, APOBEC3-associated omikli were validated using clonally expanded BT-474 breast cancer cell lines (top), omikli events resulting from exposure to benzo[a]pyrene were validated using iPSC cells (middle), and omikli events resulting from exposure to ultraviolet light were validated using iPSC cells (bottom). Figure 13B: Mutational processes of chain-cooperative kataegic events. Figure 13C: Epidemiological associations comparing the ratio of clustered tumor mutation burden (TMB) to total TMB for a given sample between drinkers (n=25) and non-drinkers (n=61); smokers (n=68) and non-smokers (n=11); homologous recombination deficient (HR deficient; n=25) and homologous recombination competent samples (HR competent; n=64). For each box plot, the center line reflects the median, the lower and upper limits of the box correspond to the first and third quartiles, and the lower and upper whiskers extend from the box by 1.5 times the interquartile range (IQR). P values were calculated using a two-tailed Mann-Whitney U test. Figure 13D: Mutational process of clustered events with inconsistent variant allele frequencies (VAFs) classified as other clustered substitutions. A minimum of two samples per cancer type are required for visualization.
[0024] [Figure 14]Figure 14A-B are examples of clustered mutation signatures. (Figure 14A) Two samples showing intramutation distance (IMD) distribution of substitutions across genomic coordinates, where each dot represents the minimum distance to a neighboring mutation for a selected mutation colored based on the subclassification of the corresponding event (rainfall plot; left). The red line indicates the sample-dependent IMD threshold for each sample. Certain clustered mutations may exceed this threshold based on correction for regional mutation density. Mutation spectrum of different catalogs of clustered and non-clustered substitutions for each sample (right; MBS not shown). (Figure 14B) Two samples showing IMD distribution of indels across a given genome, where the IMD indel threshold is shown in red (left). Non-clustered and clustered indel catalogs for each sample (right).
[0025] [Figure 15]Figures 15A-D show clustering events and structural variations. Figure 15A: Percentage of all clustering events co-located with structural variations across all cancer types (left) and each cancer type (right). Figure 15B: Distance to nearest neighbor structural variation for each class of clustered mutations (grayscale) and non-clustered mutations (red). The distribution of each class of clustering events was modeled using Gaussian mixtures (grayscale). DBS and MBS were modeled using a single distribution, while omikli, others, and indels were modeled using two components reflecting the smallest distribution of overlap with structural variations. Figure 15C: Mutational signatures active in ecDNA clustering events. Figure 15D: YTCA enrichment versus RTCA enrichment per sample within non-ecDNA kataegis (top) and non-SV-associated kataegis (bottom), where YTCA and RTCA enrichment suggest APOBEC3A or APOBEC3B activity, respectively. Genetic variants were divided into transcriptional (template strand) and coding variants. RTCA / YTCA fold enrichment was compared to the fold enrichment of non-clustered variants (p-values were calculated using a two-tailed Mann-Whitney U test and corrected for multiple hypothesis testing using the Benjamini-Hochberg false discovery rate procedure).
[0026] [Figure 16]Figure 16A-E show validation of APOBEC3 hypermutations in ecDNA in three independent cohorts. Figure 16A: Distribution of clustered substitutions (left) and clustered indels (right) across three validation cohorts. Clustered substitutions were subclassified into double base substitutions, multi-base substitutions, omikli, kataegis and other clustered mutations. Top: Each black dot represents a single cancer genome. Greyscale bars reflect the median clustered tumor mutation burden (TMB) and the proportion of clustered mutations contributing to the overall TMB of a given sample for each cancer type. Center: Proportion of each subclass of clustered events for a given cancer type with the total number of samples with at least one clustered event relative to the total number of samples in a given cancer cohort. Bottom: Proportion of clustered mutations compared to the proportion of clustered driver events for substitutions (left) and indels (right). P-values were calculated using Fisher's exact test and corrected for multiple hypothesis testing using the Benjamini-Hochberg false discovery rate procedure. Figure 16B: Left: Mutation spectrum of all kyklonas across validation cohorts. Right: Proportion of kyklonic events attributable to SBS2 and SBS13 (p-values determined using Z-score test). Figure 16C: Proportion of samples with ecDNA that co-occur with kataegis, do not co-occur with kataegis, or have no kataegic activity detected across each cohort. Figure 16D: YTCA enrichment vs. RTCA enrichment per sample with kyklonas, where YTCA and RTCA enrichment suggest higher APOBEC3A or APOBEC3B activity, respectively. RTCA / YTCA fold enrichment was compared to fold enrichment of non-clustered mutations (p-values were calculated using a two-tailed Mann-Whitney U test). Figure 16E: Proportion of ecDNA with kyklonas with multiple kyklonic events. The total number of ecDNA with kyklonas is displayed above each bar graph for each cancer type.
[0027] [Figure 17]Figures 17A-B show that kyklonas occur distally from structural breakpoints across three independent cohorts. Figure 17A: Distance to the nearest breakpoint for all kataegic mutations (grayscale), kyklonas (grayscale) and non-clustered mutations (grayscale) across three validation cohorts. Figure 17B: Distance to the nearest SV breakpoint was normalized by calculating the distance that a mutation would be expected to fall from the breakpoint given the number of detected breakpoints per chromosome and the total length of the chromosome across the validation cohort (grayscale) and PCAWG (grayscale). A value of 1 (dashed line) reflects the distance that would be expected based on random placement of mutations across chromosomes, while values less than 1 reflect mutations occurring closer than would be expected by random chance. The distribution of kataegic mutations was modeled using a Gaussian mixture model (grayscale line) with an automated selection criterion for the number of components using the minimum Bayesian information criterion (BIC).
[0028] [Figure 18]Figures 18A-C are examples of kyklonas in three independent cohorts. Figure 18A: A single undifferentiated sarcoma genome showing kataegis matches with ecDNA regions displayed as rainfall (left) and a single expanded ecDNA represented using a circos plot (middle). The outer track of the circos plot represents the reference genome of ecDNA with proximal known cancer driver genes. The middle track reflects a circular rainfall plot where each dot represents the IMD around a single mutation colored based on substitution change. The innermost track shows the average variant allele frequency (VAF) for each kyklonic event. Right: Two smaller regions of selected ecDNA containing a single kyklonic event in the ZNF536 region resulting in a large number of missense and stop-gained mutations, and a single kyklonic event in the adjacent promoter with average VAF per event. Figure 18B: A single lung adenocarcinoma genome showing kataegis matches with ecDNA regions (left) and a single expanded ecDNA with TBC1D15 and two different kyklonic events (middle) represented using a circos plot. Right: Two kyklonic events matching with upstream regions and TBC1D15. Figure 18C: A single esophageal squamous cell carcinoma genome showing kataegis matches with ecDNA regions (left) and a single expanded ecDNA with PRKAA2 and DAB1 and three different kyklonic events (middle). Right: Two kyklonic events matching with DAB1. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0029] Detailed Description of the Disclosure It is to be understood that the present disclosure is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure may be limited only by the appended claims.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this technology belongs.Although any method and material similar or equivalent to those described herein can be used to practice or test this technology, the preferred method, device and material are described herein.All technical publications and patent publications cited herein are incorporated herein by reference in their entirety.Nothing herein should be construed as an admission that the technology is not entitled to antedate such disclosure by prior invention.
[0031] The practice of the present technology employs, unless otherwise indicated, conventional techniques of tissue culture, immunology, molecular biology, microbiology, cell biology, and recombinant DNA that are within the skill of one in the art. See, e.g., Sambrook and Russell, eds. (2001) Molecular Cloning: A Laboratory Manual, 3rd ed.; series Ausubel et al., eds. (2007) Current Protocols in Molecular Biology; series Methods in Enzymology (Academic Press, Inc., New York); MacPherson et al. (1991) PCR 1: A Practical Approach (IRL Press of Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach; Harlow and Lane, eds. (1999) Antibodies, A Laboratory Manual; Freshney (2005) Culture of Animal Cells: A Manual of Basic Technique, 5th ed.; Gait, ed. (1984) Oligonucleotide Synthesis; U.S. Patent No. 4,683,195; Hames and Higgins, eds. (1984) Nucleic Acid Hybridization; Anderson (1999) Nucleic Acid Hybridization; Hames and Higgins, eds. (1984) Transcription and Translation; Immobilized Cells and Enzymes (IRL Press (1986)); Perbal (1984) A Practical Guide to Molecular Cloning; Miller and Calos, eds. (1987) Gene Transfer See Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Gene Transfer and Expression in Mammalian Cells (ed. Makrides, 2003); Immunochemical Methods in Cell and Molecular Biology (eds. Mayer and Walker, 1987) (Academic Press, London); and Weir's Handbook of Experimental Immunology (eds. Herzenberg et al., 1996).
[0032] All numerical designations, including ranges, e.g., pH, temperature, time, concentration, and molecular weight, are approximations that vary (+) or (-) by increments of 1.0 or 0.1, or alternatively by a variation of + / - 15%, or alternatively by 10%, or alternatively by 5%, or alternatively by 2%, as appropriate. It is to be understood, although not always expressly stated, that all numerical designations are preceded by the term "about." It is also to be understood, although not always expressly stated, that the reagents described herein are merely exemplary, and that equivalents of such reagents are known in the art.
[0033] Where the present technology relates to polypeptides, proteins, polynucleotides, or antibodies, it should be presumed, without explicit recitation or other intended purpose, that equivalents or biological equivalents of such are intended within the scope of the present technology. definition
[0034] As used in this specification and claims, "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a cell" includes a plurality of cells, including mixtures thereof.
[0035] As used herein, the term "comprising" is intended to mean that the compositions and methods include the recited elements but do not exclude other elements. "Consisting essentially of," when used to define compositions and methods, is intended to mean excluding other elements of any essential importance to the combination for the intended use. For example, a composition consisting essentially of the elements defined herein would not exclude trace contaminants from the isolation and purification methods, as well as pharma- ceutically acceptable carriers such as phosphate buffered saline, preservatives, and the like. "Consisting of" is intended to mean excluding trace elements of other components and substantial method steps for administering the compositions disclosed herein. Embodiments defined by each of these transition terms are within the scope of the present disclosure.
[0036] As used herein, the term "animal" refers to living multi-cellular vertebrate organisms, a category that includes mammals and birds. The term "mammal" includes both human and non-human mammals.
[0037] In one aspect, the term "equivalent" or "biologically equivalent" of an antibody refers to the ability of an antibody to selectively bind to its epitope protein or fragment thereof as measured by ELISA or other suitable method. Biologically equivalent antibodies include, but are not limited to, antibodies, peptides, antibody fragments, antibody variants, antibody derivatives, and antibody mimetics that bind to the same epitope as the reference antibody.
[0038] In one embodiment, the term "equivalent" of a "chemical equivalent" of a chemical refers to the ability of a chemical to selectively interact with its target protein, DNA, RNA or fragments thereof, as measured by inactivation of the target protein, incorporation of the chemical into DNA or RNA, or other suitable methods. Chemical equivalents include agents with the same or similar biological activity, including, but not limited to, pharma- ceutically acceptable salts or mixtures thereof that interact with and / or inactivate the same target protein, DNA or RNA as the reference chemical.
[0039] The term "allele", which is used interchangeably with "allelic variant" herein, refers to alternative forms of a gene or a part thereof. Alleles occupy the same locus or position on homologous chromosomes. If a subject has two identical alleles of a gene, the subject is said to be homozygous for that gene or allele. If a subject has two different alleles of a gene, the subject is said to be heterozygous for that gene. Alleles of a particular gene may differ from each other by a single nucleotide or several nucleotides, and may include nucleotide substitutions, deletions, and insertions. An allele of a gene may also be a form of a gene that includes a mutation.
[0040] The term "genetic marker" refers to an allelic variant of a polymorphic region of a gene of interest and / or the expression level of the gene of interest.
[0041] The term "polymorphism" refers to the coexistence of more than one form of a gene or a portion thereof. A portion of a gene that has at least two different forms, i.e., two different nucleotide sequences, is called a "polymorphic region of a gene." A polymorphic region can be a single nucleotide whose identity differs in different alleles.
[0042] The term "genotype" refers to the particular allelic composition of a cell or a particular gene, and in some embodiments, the particular polymorphisms associated with that gene, and the term "phenotype" refers to the detectable outward manifestation of a particular genotype.
[0043] As used herein, the term "isolated" refers to a molecule or biological material or cellular material that is substantially free of other materials. In one aspect, the term "isolated" refers to a nucleic acid, such as DNA or RNA, or a protein or polypeptide, or a cell or organelle, or a tissue or organ, respectively, that is separated from other DNA or RNA, or proteins or polypeptides, or cells or organelles, or tissues or organs that are present in the natural source. The term "isolated" also refers to a nucleic acid or peptide that is substantially free of cellular material, viral material or medium if produced by recombinant DNA technology, or chemical precursors or other chemicals if chemically synthesized. Furthermore, "isolated nucleic acid" is meant to include nucleic acid fragments that are not naturally occurring as fragments and would not be found in the natural state. The term "isolated" is also used herein to refer to a polypeptide that is isolated from other cellular proteins, and is meant to encompass both purified and recombinant polypeptides. The term "isolated" is also used herein to refer to a cell or tissue that is isolated from other cells or tissues, and is meant to encompass both cultured and engineered cells or tissues.
[0044] As used herein, "treating" a disease in a subject or "treatment" of a disease in a subject refers to (1) preventing a symptom or disease from occurring in a subject who is predisposed to the disease or does not yet show symptoms of the disease; (2) inhibiting or arresting the development of the disease; or (3) improving or regressing the disease or symptoms of the disease. As understood in the art, "treatment" is an approach to obtain beneficial or desired results, including clinical results. For the purposes of this technology, beneficial or desired results may include, but are not limited to, alleviation or amelioration of one or more symptoms, whether detectable or undetectable, attenuation of the severity of a condition (including a disease), stabilization (i.e., not worsening) of a condition (including a disease), delay or slowdown of a condition (including a disease), progression, improvement or alleviation of a condition (including a disease), state and remission (whether partial or total) of a condition (including a disease). In one embodiment, treatment excludes prevention.
[0045] As used herein, "active treatment" or "active chemotherapy" may refer to any one or combination of therapeutic cancer therapies, including but not limited to any form of chemo-drug therapy meant to destroy rapidly growing / proliferating cancer cells in the body. "Active chemotherapy" refers to any treatment that may extend beyond the first-line treatment or standard treatment regimen for any particular cancer or tumor. "Active chemotherapy" may include, but is not limited to, adoptive cellular therapy, immune checkpoint blockade including PD1, PD-L1, and CTLA4, pre-targeted radioimmunotherapy, oncolytic virus therapy, or cancer vaccines.
[0046] Where the disease is cancer, the following clinical endpoints are non-limiting examples of treatments: (1) elimination of cancer in the subject, or in a tissue / organ of the subject, or at a cancer locus; (2) reduction in tumor burden (e.g., number of cancer cells, number of cancer foci, number of cancer cells within foci, size of solid tumors, enrichment of liquid cancer in bodily fluids, and / or amount of cancer in the body); (3) progression and / or metastasis of cancer, including, but not limited to, growth and / or division of cancer cells, size growth of solid tumors or cancer loci, progression of cancer and / or metastasis (time to form new metastases, number of total metastases, size of metastases, and various tissues / organs to house metastatic cells, etc.). (4) stabilization or delay or slowing or inhibition of the development of cancer; (5) lower risk of having cancer growth and / or progression; (6) induction of a patient's immune response against cancer, such as a higher number of tumor infiltrating immune cells, a higher number of activated immune cells, or a higher number of cancer cells expressing an immunotherapy target, or a higher expression level of an immunotherapy target in cancer cells; (7) increased probability of survival and / or increased survival, such as increased overall survival (OS, which may be expressed as 1-year, 2-year, 5-year, 10-year, or 20-year survival rate), increased progression-free survival (PFS), increased disease-free survival (DFS), increased time to tumor recurrence (TTR), and increased time to tumor progression (TTP). In some embodiments, the subject following treatment experiences one or more endpoints selected from tumor response, reduction in tumor size, reduction in tumor burden, increased overall survival, increased progression-free survival, inhibition of metastasis, improved quality of life, minimization of drug-related toxicity, and avoidance of side effects (e.g., reduction in treatment-emergent adverse events). In some embodiments, improving quality of life includes, but is not limited to, elimination or amelioration of cancer-specific symptoms, such as fatigue, pain, nausea / vomiting, loss of appetite, and constipation; improving or maintaining psychological well-being (e.g., levels of irritability, depression, memory loss, tension, and anxiety); improving or maintaining social well-being (e.g., reduced need for assistance with eating, dressing, or using the toilet; improving or maintaining the ability to engage in usual leisure, hobbies, or social activities; improving or maintaining relationships with family).In some embodiments, improvement in patient quality of life, as measured qualitatively through patient narrative, is measured quantitatively using validated quality of life tools known to those skilled in the art, or a combination thereof. Further non-limiting examples of endpoints include reduced hospitalizations, reduced use of medications to treat side effects, longer drug holidays, and earlier return to work or caregiving responsibilities. In one aspect, prophylaxis or preventative measures are excluded from treatment.
[0047] As used herein, immune cells are cells of the immune system, including but not limited to lymphocytes (e.g., T cells, B cells, natural killer (NK) cells and natural killer T (NKT) cells), bone marrow derived cells (e.g., granulocytes (basophils, eosinophils, neutrophils, mast cells), monocytes, macrophages and dendritic cells (DCs). T cells are divided into two broad categories, namely, CD8+ T cells or CD4+ T cells, based on which proteins are present on the surface of the cells. CD8+ T cells are also called cytotoxic T cells or cytotoxic lymphocytes (CTLs). The four major CD4+ T cell subsets are TH1, TH2, TH17 and Treg, with "TH" referring to "T helper cells". T cells may also refer to gamma delta T cells. Dendritic cells (DCs) are important antigen presenting cells (APCs) and can arise from monocytes. In some embodiments, immune cells refer to killer cells, including but not limited to cytotoxic T cells, gamma delta T cells, NK cells and NK-T cells. In one embodiment, immune cells are CD45+ cells.
[0048] The terms "subject," "host," "individual," and "patient" are used interchangeably herein to refer to animals, typically mammals. Any suitable mammal can be treated by the methods described herein. Non-limiting examples of mammals include humans, non-human primates (e.g., apes, gibbons, chimpanzees, orangutans, monkeys, macaques, etc.), domestic animals (e.g., dogs and cats), farm animals (e.g., horses, cows, goats, sheep, pigs), and laboratory animals (e.g., mice, rats, rabbits, guinea pigs). In some embodiments, the mammal is a human. The mammal can be of any age or at any stage of development (e.g., adult, adolescent, child, infant, or intrauterine mammal). The mammal can be male or female. In some embodiments, the subject is a human. In some embodiments, the subject has cancer, is diagnosed with cancer, or is suspected of having cancer. The subject can be male or female.
[0049] In certain embodiments, the terms "disease", "disorder" and "condition" are used interchangeably herein and refer to cancer, the status of having been diagnosed with cancer, or the status of being suspected of having cancer. "Cancer", also referred to herein as "tumor", is medically known as the uncontrolled division of abnormal cells in a part of the body, benign or malignant. In one embodiment, cancer refers to a broad group of diseases including malignant neoplasms, unregulated cell division and growth, and invasion of nearby parts of the body. Non-limiting examples of cancer include carcinomas, sarcomas, leukemias and lymphomas, such as colon cancer, colorectal cancer, rectal cancer, gastric cancer, esophageal cancer, head and neck cancer, breast cancer, brain cancer, lung cancer, stomach cancer, liver cancer, gallbladder cancer or pancreatic cancer. In one embodiment, the term "cancer" refers to solid tumors, which are abnormal masses of tissue that usually do not contain cysts or liquid areas, including, but not limited to, sarcomas, carcinomas, and certain lymphomas (such as non-Hodgkin's lymphoma). In another embodiment, the term "cancer" refers to cancers present in bodily fluids (e.g., blood and bone marrow), such as leukemia (cancer of the blood) and certain lymphomas.
[0050] Additionally or alternatively, cancer may refer to localized cancer (which is an invasive malignant cancer that is entirely confined to the organ or tissue in which it originated), metastatic cancer (which refers to a cancer that has spread from its site of origin to another part of the body), non-metastatic cancer, primary cancer (a term used to describe the first cancer a subject experiences), secondary cancer (which refers to a metastasis from a second cancer unrelated to the primary or first cancer), advanced cancer, unresectable cancer, or recurrent cancer. As used herein, advanced cancer refers to a cancer that has progressed after receiving one or more of a first-line treatment, a second-line treatment, or a third-line treatment.
[0051] The term "chemotherapy" includes cancer therapy using chemical or biological agents or other therapies, such as radiation therapy, small molecule drugs or large molecules, such as antibodies, immunotherapy, RNAi and gene therapy.Non-limiting examples of chemotherapy are provided below.Although not always explicitly stated, when specific treatments are noted, it should be understood that the scope of the present disclosure includes equivalents unless excluded.
[0052] The term "contacting" refers to a direct or indirect binding or interaction between two or more. A specific example of a direct interaction is binding. A specific example of an indirect interaction is when one entity acts on an intermediate molecule, which then acts on a second referenced entity. Contacting as used herein includes in solution, in solid phase, in vitro, ex vivo, in a cell, and in vivo. In vivo contacting can be referred to as administering or administration.
[0053] As used herein, the terms "administration" and "administering" are used to mean introducing a drug into a subject. Routes of administration include, but are not limited to, oral (such as tablets, capsules, or suspensions), topical, transdermal, intranasal, vaginal, rectal, subcutaneous intravenous, intravenous, intraarterial, intramuscular, intraosseous, intraperitoneal, intraocular, subconjunctival, subtenon, intravitreal, retrobulbar, intracameral, intratumoral, epidural, and intrathecal.
[0054] "Immunotherapeutic agent" refers to a type of cancer treatment that uses the patient's own immune system to fight cancer, including but not limited to physical intervention, chemicals, biological molecules or particles, cells, tissues or organs, or any combination thereof, to enhance or activate or initiate the patient's immune response against cancer. Non-limiting examples of immunotherapeutic agents include antibodies, immunomodulators, checkpoint inhibitors, antisense oligonucleotides (ASOs), RNA interference (RNAi), clustered regularly interspaced short palindromic repeats (CRISPR) systems, viral vectors, anti-cancer cell therapy (e.g., transplanting anti-cancer immune cells that have been expanded and / or activated in vivo, if necessary, or administering immune cells expressing chimeric antigen receptors (CARs), CAR therapy, and cancer vaccines. As used herein, unless otherwise specified, an immunotherapeutic agent is not an inhibitor of thymidylate biosynthesis or an anthracycline or other topoisomerase II inhibitor. As used herein, immune checkpoint refers to a regulator and / or modulator of the immune system (e.g., immune response, antitumor immune response, neoplastic antitumor immune response, antitumor immune cell response, antitumor T cell response, and / or antigen recognition of T cell receptors in the course of immune response). Their interaction activates either inhibitory or activating immune signaling pathways. Thus, checkpoints can include one of two signals: stimulatory immune checkpoints that stimulate immune response, and inhibitory immune checkpoints that inhibit immune response. In some embodiments, immune checkpoints are important for self-tolerance, which prevents the immune system from attacking cells indiscriminately. However, some cancers can protect themselves from attack by stimulating immune checkpoint targets. In some embodiments, immune checkpoints are present on T cells, antigen-presenting cells (APCs) and / or tumor cells.
[0055] One target of immunotherapeutic agents is tumor-specific antigens, and immunotherapy directs or enhances the immune system to recognize and attack tumor cells. Non-limiting examples of such agents include cancer vaccines that present tumor-specific antigens to the patient's immune system, monoclonal antibodies or antibody-drug conjugates that specifically bind to tumor-specific antigens, bispecific antibodies that specifically bind to tumor-specific antigens and immune cells (such as T cell engagers or NK cell engagers), immune cells (such as killer cells) that specifically bind to tumor-specific antigens (such as CAR-T cells, CAR-NK cells, CAR-NKT cells), polynucleotides (or vectors containing them) that transfect / transduce immune cells to express tumor-specific antibodies of their antigen-binding fragments (e.g., CARs), or polynucleotides (or vectors containing them) that transfect / transduce cancer cells to express antigens or markers that can be recognized by immune cells.
[0056] Another exemplary target is an inhibitory immune checkpoint that suppresses nascent anti-tumor immune responses, such as A2AR, B7-H3, B7-H4, BTLA, CTLA-4, CTLA-4 / B7-1 / B7-2, IDO, KIR, LAG3, NOX2, PD-1, PD-L1, and TIM-3, VISTA, SIGLEC7 (sialic acid-binding immunoglobulin-type lectin 7, also known as CD328), and SIGLEC9 (sialic acid-binding immunoglobulin-type lectin 9, also known as CD329). Non-limiting examples of such agents include antagonists or inhibitors of inhibitory immune checkpoints, agents that reduce the expression and / or activity of inhibitory immune checkpoints (e.g., via antisense oligonucleotides (ASOs), RNA interference (RNAi), or clustered regularly interspaced short palindromic repeats (CRISPR) systems), antibodies or antibody-drug conjugates or ligands that specifically bind to inhibitory immune checkpoints and reduce (or inhibit) the activity of inhibitory immune checkpoints, immune cells with reduced (or inhibited) inhibitory immune checkpoints (which optionally specifically bind tumor-specific antigens, such as CAR-T cells, CAR-NK cells, and CAR-NKT cells), and polynucleotides (or vectors containing same) that transfect / transduce immune cells or cancer cells to reduce or inhibit their inhibitory immune checkpoints. Such reduced expression or activity of inhibitory immune checkpoints enhances the patient's immune response to cancer.
[0057] Further possible immunotherapeutic targets are stimulatory checkpoint molecules (including but not limited to 4-1BB, CD27, CD28, CD40, CD122, CD137, OX40, GITR and ICOS), and immunotherapeutic agents activate or enhance anti-tumor immune responses. Non-limiting examples of such agents include agonists of stimulatory checkpoints, agents that increase the expression and / or activity of stimulatory immune checkpoints, antibodies or antibody-drug conjugates or ligands that specifically bind to and activate or enhance the activity of stimulatory immune checkpoints, immune cells with increased expression and / or activity of stimulatory immune checkpoints (such as CAR-T cells, CAR-NK cells and CAR-NKT cells, which optionally specifically bind to tumor-specific antigens), and polynucleotides (or vectors containing same) that transfect / transduce immune cells or cancer cells to express the stimulatory immune checkpoint.
[0058] Additional or alternative targets may be exploited by immunotherapeutic agents such as immune regulating agents, including but not limited to agents that activate immune cells, recruit immune cells to cancer or cancer cells, or increase immune cells infiltrating solid tumors and / or cancer loci. Non-limiting examples of such agents are immuneregulators or variants, mutants, fragments, or equivalents thereof.
[0059] In some embodiments, the immunotherapeutic agent utilizes one or more targets, such as a bispecific T cell engager, a bispecific NK cell engager, or a CAR cell therapy, In some embodiments, the immunotherapeutic agent targets one or more immune regulatory or effector cells.
[0060] As used herein, the term "antibody" refers collectively to immunoglobulin or immunoglobulin-like molecules, including, by way of example only, IgA, IgD, IgE, IgG, and IgM, combinations thereof, and similar molecules produced during the immune response in any vertebrate, including mammals such as humans, goats, rabbits, rats, dogs, donkeys, mice, camelids (e.g., dromedaries, llamas, and alpacas), as well as non-mammalian species, such as shark immunoglobulins. Unless otherwise specified, the term "antibody" includes intact immunoglobulins and "antibody fragments" or "antigen-binding fragments" that specifically bind to a molecule of interest (or a group of closely similar molecules of interest) and substantially exclude binding to other molecules (e.g., antibodies and antibody fragments have a binding constant at least 10 greater than that for other molecules in a biological sample). 3 M -1 Large, at least 10 4 M -1 Greater than or at least 10 5 M -1(having a large binding constant for the molecule of interest). The term "antibody" also includes genetically engineered forms such as chimeric antibodies (e.g., murine or humanized non-primate antibodies), heteroconjugate antibodies (e.g., bispecific antibodies). See also Pierce Catalog and Handbook, 1994-1995 (Pierce Chemical Co., Rockford, Ill.); Owen et al., Kuby Immunology, 7th ed., WH Freeman & Co, 2013; Murphy, Janeway's Immunobiology, 8th ed., Garland Science, 2014; Male et al., Immunology (Roitt), 8th ed., Saunders, 2012; Parham, The Immune System, 4th ed., Garland Science, 2014. The term "antibody" includes any protein- or peptide-containing molecule that contains at least a portion of an immunoglobulin molecule, such as a whole antibody and any antigen-binding fragment or single chain thereof. The terms "antibody", "antibodies" and "immunoglobulin" also include immunoglobulins of any isotype, fragments of antibodies that retain specific binding to an antigen, including, but not limited to, Fab, Fab', F(ab)2, Fv, scFv, dsFv, Fd fragments, dAb, VH, VL, VhH, and V-NAR domains; minibodies, diabodies, triabodies, tetrabodies and kappabodies; antibody fragments and multispecific antibody fragments formed from one or more isolated antibody fragments. Such examples include, but are not limited to, the complementarity determining regions (CDRs) or ligand binding portions of the heavy or light chains, heavy or light chain variable regions, heavy or light chain constant regions, framework (FR) regions, or any portion thereof, at least a portion of a binding protein, chimeric antibodies, humanized antibodies, single chain antibodies, and fusion proteins comprising the antigen-binding portion of an antibody and a non-antibody protein. The variable regions of the heavy and light chains of an immunoglobulin molecule contain the binding domains that interact with the antigen. The constant region of the antibody (Ab) may mediate the binding of the immunoglobulin to host tissue.Antibodies can be polyclonal, monoclonal, multispecific (eg, bispecific antibodies) and antibody fragments, so long as they exhibit the desired biological activity.
[0061] As used herein, the term "monoclonal antibody" refers to an antibody produced by a single clone of B lymphocytes or by a cell transfected with the light and heavy chain genes of a single antibody. Monoclonal antibodies are produced by methods known to those skilled in the art, such as by creating hybrid antibody-forming cells from the fusion of myeloma cells and immune spleen cells. Monoclonal antibodies include humanized monoclonal antibodies.
[0062] In some embodiments, the antibody is a bispecific immune cell engager, which refers to a bispecific monoclonal antibody that can recognize and specifically bind to tumor antigens (such as CD19, EpCAM, MCSP, HER2, EGFR, or CS-1) and immune cells, and direct immune cells to cancer cells, thereby treating cancer. Non-limiting examples of such antibodies include bispecific T cell engagers, bispecific cytotoxic T lymphocyte (CTL) engagers, and bispecific NK cell engagers. In one embodiment, the engager is a fusion protein consisting of two single chain variable fragments (scFv) of different antibodies. Additionally or alternatively, the immune cell is a killer cell, including, but not limited to, cytotoxic T cells, gamma delta T cells, NK cells, and NK-T cells.
[0063] As used herein, the term "antigen-binding domain" refers to any protein or polypeptide domain capable of specifically binding to an antigen target.
[0064] As used herein, the term "chimeric antigen receptor" (CAR) refers to a fusion protein that includes an extracellular domain capable of binding to an antigen, a transmembrane domain derived from a polypeptide different from the polypeptide from which the extracellular domain is derived, and at least one intracellular domain. A "chimeric antigen receptor (CAR)" may also be referred to as a "chimeric receptor," "T-body," or "chimeric immune receptor (CIR)." An "extracellular domain capable of binding to an antigen" refers to any oligopeptide or polypeptide that can bind to a specific antigen. An "intracellular domain" or "intracellular signaling domain" refers to any oligopeptide or polypeptide that is known to function as a domain that transmits a signal that causes activation or inhibition of a biological process within a cell. In certain embodiments, the intracellular domain may include, or consist essentially of, or even further include, in addition to a primary signaling domain, one or more costimulatory signaling domains. A "transmembrane domain" refers to any oligopeptide or polypeptide that is known to span a cell membrane and that can function to link the extracellular domain and the signaling domain. A chimeric antigen receptor may optionally include a "hinge domain" that serves as a linker between the extracellular domain and the transmembrane domain.
[0065] As used herein, CAR therapy can refer to administering immune cells expressing a CAR to a subject, as well as contacting immune cells (such as in vivo) with a vector expressing a CAR.
[0066] As used herein, the term "NK cells", also known as natural killer cells, refers to a type of lymphocyte that originates from bone marrow and plays a key role in the innate immune system. NK cells provide a rapid immune response against virus-infected cells, tumor cells or other stressed cells, even in the absence of antibodies and major histocompatibility complexes on the cell surface. NK cells for use in cell therapy and / or CAR therapy may be isolated or obtained from commercially available sources. Non-limiting examples of commercially available NK cell lines include the NK-92 line (ATCC® CRL-2407™), the NK-92MI line (ATCC® CRL-2408™). Further examples include, but are not limited to, the NK lines HANK1, KHYG-1, NKL, NK-YS, NOI-90 and YT. Non-limiting exemplary sources of such commercially available cell lines include the American Type Culture Collection, or ATCC (http: / / www.atcc.org / ) and the German Collection of Microorganisms and Cell Cultures (https: / / www.dsmz.de / ).
[0067] As used herein, the term "T cell" refers to a type of lymphocyte that matures in the thymus. T cells play an important role in cell-mediated immunity and are distinguished from other lymphocytes, such as B cells, by the presence of a T cell receptor on the cell surface. T cells for use in cell therapy and / or CAR therapy can be isolated or obtained from commercial sources. "T cells" include all types of immune cells that express CD3, including T helper cells (CD4+ cells), cytotoxic T cells (CD8+ cells), natural killer T cells, T regulatory cells (Tregs) and gamma delta T cells. "Cytotoxic cells" include CD8+ T cells, natural killer (NK) cells, and neutrophils, which can mediate cytotoxic responses. Non-limiting examples of commercially available T cell lines include the lines BCL2(AAA) Jurkat (ATCC® CRL-2902™), BCL2(S70A) Jurkat (ATCC® CRL-2900™), BCL2(S87A) Jurkat (ATCC® CRL-2901™), BCL2 Jurkat (ATCC® CRL-2899™), Neo Jurkat (ATCC® CRL-2898™), TALL-104 cytotoxic human T cell line (ATCC#CRL-11386).Further examples include mature T cell lines such as, for example, Deglis, EBT-8, HPB-MLp-W, HUT 78, HUT 102, Karpas 384, Ki 225, My-La, Se-Ax, SKW-3, SMZ-1, and T34; and immature T cell lines such as, for example, ALL-SIL, Be13, CCRF-CEM, CML-T1, DND-41, DU.528, EU-9, HD-Mar, HPB-ALL, H-SB2, HT-1, JK-T1, Jurkat, Karpas 45, KE-37, KOPT-K1, K-T1, L-KAW, Loucy, MAT, MOLT-1, MOLT 3, MOLT-4, MOLT 13, MOLT-16, MT-1, MT-ALL, P12 / Ichikawa, Peer, PER0117, PER-255, PF-382, PFI-285, RPMI-8402, ST-4, SUP-T1~T14, TALL-1, TALL-101, T ALL-103 / 2, TALL-104, TALL-105, TALL-106, TALL-107, TALL-197, TK-6, TLBR-1, -2, -3, and -4, CCRF-HSB-2 (CCL-120.1), J.RT3-T3.5 (ATCC TIB-153), J45.01 (ATCC CRL-1990), J.CaM1.6 (ATCC CRL-2063), RS4;11 (ATCC CRL-1873), CCRF-CEM (ATCC CRM-CCL-119); and cutaneous T-cell lymphoma lines, such as HuT78 (ATCC CRM-TIB-161), MJ[G11] (ATCC CRL-8294), HuT102 (ATCC TIB-162). Null leukemia cell lines, including but not limited to REH, NALL-1, KM-3, L92-221, are another commercially available source of immune cells for use in CAR therapy, as are cell lines derived from other leukemias and lymphomas, such as K562 erythroleukemia, THP-1 monocytic leukemia, U937 lymphoma, HEL erythroleukemia, HL60 leukemia, HMC-1 leukemia, KG-1 leukemia, U266 myeloma.Non-limiting exemplary sources of such commercially available cell lines include the American Type Culture Collection, or ATCC (http: / / www.atcc.org / ) and the German Collection of Microorganisms and Cell Cultures (https: / / www.dsmz.de / ).
[0068] As used herein, "tumor-specific antigen" refers to an antigenic substance produced in tumor cells that can induce an immune response in a subject. In some embodiments, such tumor-specific antigens are not expressed on or in cells of a subject that are not cancer cells. In some embodiments, such tumor-specific antigens may still be expressed in or on some non-cancer cells. For example, tumor-specific antigens may not be expressed on the cell surface of non-cancer cells in a subject. In one embodiment, tumor-specific antigens may be expressed on or in non-cancer cells of a subject, but at a much lower level compared to cancer cells. In another embodiment, tumor-specific antigens may be expressed on or in non-cancer cells of a subject that are not adjacent to cancer or cancer cells.Non-limiting examples of tumor-specific antigens include: alpha-fetoprotein (AFP), beta-2-microglobulin (B2M), beta-human chorionic gonadotropin (β-hCG), bladder tumor antigen (BTA), C-kit / CD117, CA15-3 / CA27.29, CA19-9, CA-125, CA27.29, calcitonin, carcinoembryonic antigen (CEA), chromogranin A (CgA), cytokeratin fragment 21-1, Des-gamma-carboxyprothrombin (DCP), estrogen receptor (ER) / progestin. Theron receptor (PR), epithelial tumor antigen (ETA), fibrin / fibrinogen, gastrin, HE4, overexpression of HER2 / neu, 5-HIAA, lactate dehydrogenase, melanoma-associated antigen (MAGE), MUC-1, neuron-specific enolase (NSE), nuclear matrix protein 22, programmed death ligand 1 (PD-L1), prostate-specific antigen (PSA), prostatic acid phosphatase (PAP), soluble mesothelin-related peptide (SMRP), somatostatin receptor, tyrosinase, thyroglobulin, and aberrant production of ras. p53, alpha folate receptor, 5T4, ανβ6 integrin, BCMA, B7-H3, B7-H6, CAIX, CD16, CD19, CD20, CD22, CD25, CD30, CD33, CD44, CD44v6, CD44v7 / 8, CD70, CD79a, CD79b, CD123, CD138, CD171, CEA, CSPG4, EGFR, EGFR family including ErbB2 (HER2), EGFRvni, EGP2, EGP40, EPCAM, EphA2, EpCAM, FAP, fetal AchR, FRoc, GD2, GD3, glypican-3 (GPC3), HLA-A1+MAGE1, HLA-A2+MAGE1, HLA-A3+MAGE1, HLA-Al+NY-ESO-1, HLA-A2+NY-ESO-1, HLA-A3+NY-ESO-1, IL-1lRoc, IL-13R a2, Lambda, Lewis-Y, Kappa, mesothelin, Mucl, Mucl6, NCAM, NKG2D ligand, NY-ESO-1, PRAME, PSCA, PSMA, ROR1, SSX, Survivin, TAG72, TEM, VEGFR2 and WT-1.
[0069] An "effective amount" or "therapeutically effective amount" is intended to refer to the amount of a compound or agent administered or delivered to a patient that is most likely to result in a desired response to treatment. The amount is empirically determined by the patient's clinical parameters, including, but not limited to, stage of disease, age, sex, histology, and likelihood of tumor recurrence.
[0070] As used herein, a "patient" is intended to mean an animal patient, a mammalian patient, or even a human patient. By way of example only, mammals include, but are not limited to, simian, murine, bovine, equine, porcine or ovine subjects. Patients may be female or male.
[0071] The terms "clinical outcome", "clinical parameter", "clinical response" or "clinical endpoint" refer to any clinical observation or measurement of a patient's response to treatment. Non-limiting examples of clinical outcomes include tumor response (TR), overall survival (OS), progression-free survival (PFS), disease-free survival, time to tumor recurrence (TTR), time to tumor progression (TTP), relative risk (RR), objective response rate (RR or ORR), toxicity or side effects.
[0072] The term "suitable for treatment" or "suitably treated with treatment" is intended to mean that a patient is likely to exhibit one or more favorable clinical outcomes compared to a patient with the same disease and receiving the same treatment, but with a different characteristic under consideration for comparison purposes. In one embodiment, the characteristic under consideration is a gene polymorphism or somatic mutation. In another embodiment, the characteristic under consideration is the expression level of a gene or polypeptide. In one embodiment, the more favorable clinical outcome is a relatively high or relatively good chance of tumor response, such as tumor burden reduction. In another embodiment, the more favorable clinical outcome is a relatively long overall survival. In yet another embodiment, the more favorable clinical outcome is a relatively long progression-free survival or time to tumor progression. In yet another embodiment, the more favorable clinical outcome is a relatively long disease-free survival. In yet another embodiment, the more favorable clinical outcome is a relative reduction or delay of tumor recurrence. In another embodiment, the more favorable clinical outcome is a relatively reduced metastasis. In another embodiment, the more favorable clinical outcome is a relatively low relative risk. In yet another aspect, the more desirable clinical outcome is a relative reduction in toxicity or side effects. In some embodiments, more than one clinical outcome is considered simultaneously. In one such aspect, a patient with a characteristic, such as a genotype of a gene polymorphism, can show more than one desirable clinical outcome compared to a patient without this characteristic, who has the same disease and is receiving the same treatment. As defined herein, the patient is considered suitable for treatment. In another such aspect, a patient with a characteristic can show one or more desirable clinical outcomes, but at the same time show one or more less desirable clinical outcomes. The clinical outcomes are then considered together, and a decision is made accordingly as to whether the patient is suitable for treatment, taking into account the patient's particular situation and the relevance of the clinical outcomes. In some embodiments, progression-free survival or overall survival is weighted more heavily than tumor response in collective decision-making.
[0073] Response criteria can be based on RECIST criteria (Therasse and Arbuck et al., 2000, New Guidelines to Evaluate Response to Treatment in Solid Tumors, J Natl Cancer Inst, 92:205-16). A "complete response" (CR) to treatment refers to the clinical state of a patient with evaluable but non-measurable disease, in which the tumor and all evidence of disease have disappeared following administration of treatment. In this context, a "partial response" (PR) refers to a response that is less than a complete response. "Stable disease" (SD) indicates that the patient remains stable following treatment. "Progressive disease" (PD) indicates that the tumor has grown (i.e., gotten larger) or spread (i.e., metastasized to another tissue or organ), or the cancer overall has worsened following treatment. For example, tumor growth of more than 20 percent after initiation of treatment typically indicates progressive disease. A "non-response" (NR) to treatment refers to a patient state in which the tumor or evidence of disease remains constant or progresses.
[0074] "Overall survival" (OS) refers to the length of time a cancer patient remains alive after cancer treatment.
[0075] "Progression-free survival" (PFS) or "time to tumor progression" (TTP) refers to the length of time after treatment that a cancer patient's tumor does not grow. Progression-free survival includes the time during which a patient experiences a complete response, partial response, or stable disease.
[0076] "Disease-free survival" refers to the length of time after treatment that a cancer patient lives without signs of cancer or tumor.
[0077] "Time to tumor recurrence (TTR)" refers to the length of time before a tumor reappears (returns) after cancer treatment, such as surgical removal or chemotherapy. The tumor may return in the same location as the original (primary) tumor or in another location in the body.
[0078] "Relative risk" (RR) in statistics and mathematical epidemiology refers to the risk of an event (or disease development) given an exposure. Relative risk is the ratio of the odds of an event occurring in an exposed group versus an unexposed group.
[0079] "Objective response rate" refers to the proportion of responders (patients with either a partial response (PR) or a complete response (CR)) compared with non-responders (patients with either SD or PD). Duration of response can be measured from the time of initial response to documented tumor progression.
[0080] The term "identify" or "identifying" refers to closely relating or associating a patient with a group or population of patients likely to experience the same or similar clinical response to a treatment.
[0081] The term "selecting" a patient for treatment refers to indicating that the selected patient is suitable for treatment. Such indication may be in writing, for example, by a handwritten prescription or a computerized report making a corresponding prescription or recommendation.
[0082] "Normal cells corresponding to a tumor tissue type" refers to normal cells derived from the same tissue type as the tumor tissue. Non-limiting examples are normal lung cells from a patient with a lung tumor, or normal colon cells from a patient with a colon tumor.
[0083] As used herein, the term "amplification" or "amplifying" refers to one or more methods known in the art for copying a target nucleic acid, thereby increasing the number of copies of a selected nucleic acid sequence. Amplification can be exponential or linear. The target nucleic acid can be either DNA or RNA. The sequence thus amplified forms an "amplicon." The exemplary methods described below relate to amplification using polymerase chain reaction ("PCR," although numerous other methods for amplifying nucleic acids are known in the art (e.g., isothermal, rolling circle, etc.). Those skilled in the art will understand that these other methods can be used in place of or in conjunction with PCR methods.
[0084] As used herein, the term "complement" refers to a complementary sequence to a nucleic acid according to standard Watson / Crick base pairing rules. A complementary sequence may also be a sequence of RNA that is complementary to a DNA sequence or its complementary sequence, or may be cDNA. As used herein, the term "substantially complementary" means that two sequences hybridize under stringent hybridization conditions. Those skilled in the art will understand that substantially complementary sequences need not hybridize along their entire length. In particular, a substantially complementary sequence includes a contiguous sequence of bases that does not hybridize to a target or marker sequence, located 3' or 5' to a contiguous sequence of bases that hybridize to a target or marker sequence under stringent hybridization conditions.
[0085] As used herein, the term "hybridize" or "specifically hybridize" refers to the process in which two complementary nucleic acid strands anneal to each other under appropriately stringent conditions. Hybridization is typically performed using a probe-length nucleic acid molecule. Nucleic acid hybridization techniques are well known in the art. Those skilled in the art know how to estimate and adjust the stringency of hybridization conditions so that sequences with at least the desired level of complementarity will stably hybridize, while sequences with lower complementarity will not hybridize. For examples of hybridization conditions and parameters, see, for example, Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor, Plainview, NY; Ausubel, FM et al., 1994, Current Protocols in Molecular Biology. John Wiley & Sons, Secaucus, NJ.
[0086] As used herein, a "primer" refers to an oligonucleotide that can act as an initiation point for synthesis when placed under conditions that initiate primer extension (e.g., primer extension in the context of applications such as PCR). A primer is complementary to a target nucleotide sequence and hybridizes to a substantially complementary sequence in the target, resulting in the addition of a nucleotide to the 3' end of the primer in the presence of a DNA or RNA polymerase. The 3'-nucleotide of a primer should generally be complementary to the target sequence at the corresponding nucleotide position for optimal expression and amplification. An oligonucleotide "primer" can be naturally occurring, as in the case of a purified restriction digest, or can be synthetically produced. As used herein, the term "primer" includes all forms of primers that can be synthesized, including peptide nucleic acid primers, locked nucleic acid primers, phosphorothioate modified primers, labeled primers, and the like.
[0087] Primers are typically about 5 to about 100 nucleotides in length, such as about 15 to about 60 nucleotides in length, such as about 20 to about 50 nucleotides in length, such as about 25 to about 40 nucleotides in length. In some embodiments, primers can be at least 8, at least 12, at least 16, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60 nucleotides in length. The optimal length for a particular primer application can be readily determined by the methods described in H. Erlich, PCR Technology. Principles and Application for DNA Amplification (1989).
[0088] As used herein, "probe" refers to a nucleic acid that interacts with a target nucleic acid through hybridization. A probe can be fully or partially complementary to a target nucleic acid sequence. The level of complementarity generally depends on many factors based on the function of the probe. One or more probes can be used to detect the presence or absence of a mutation in a nucleic acid sequence, for example, by the sequence characteristics of the target. A probe can be labeled or unlabeled, or modified in any of several ways well known in the art. A probe can specifically hybridize to a target nucleic acid.
[0089] The probe can be DNA, RNA or RNA / DNA hybrid. The probe can be an oligonucleotide, an artificial chromosome, a fragmented artificial chromosome, a genomic nucleic acid, a fragmented genomic nucleic acid, an RNA, a recombinant nucleic acid, a fragmented recombinant nucleic acid, a peptide nucleic acid (PNA), a locked nucleic acid, an oligomer of a cyclic heterocycle, or a conjugate of a nucleic acid. The probe can include modified nucleic acid bases, modified sugar moieties, and modified internucleotide linkages. The probe can be fully complementary to the target nucleic acid sequence or can be partially complementary. The probe can be used to detect the presence or absence of the target nucleic acid. The probe is typically at least about 10, 15, 21, 25, 30, 35, 40, 50, 60, 75, 100 nucleotides or more in length.
[0090] As used herein, "detecting" refers to determining the presence of a nucleic acid of interest in a sample or the presence of a protein of interest in a sample. Detection does not require a method to provide 100% sensitivity and / or 100% specificity.
[0091] As used herein, "detectable label" refers to a molecule or compound or a group of molecules or compounds used to identify a nucleic acid or protein of interest. In some cases, the detectable label can be directly detected. In other cases, the detectable label can be part of a binding pair that can be subsequently detected. The signal from the detectable label can be detected by various means and depends on the nature of the detectable label. The detectable label can be an isotope, a fluorescent moiety, a colored substance, etc. Examples of means for detecting a detectable label include, but are not limited to, spectroscopic, photochemical, biochemical, immunochemical, electromagnetic, radiochemical, or chemical means such as fluorescence, chemiluminescence, or chemiluminescence, or any other suitable means.
[0092] As used herein, "TaqMan® PCR detection system" refers to a method for real-time PCR. In this method, a PCR reaction mix includes a TaqMan® probe that hybridizes to an amplified nucleic acid segment. The TaqMan® probe includes a donor and a quencher fluorophore at both ends of the probe that are close enough to each other that the fluorescence of the donor is captured by the quencher. However, once the probe hybridizes to the amplified segment, the 5'-exonuclease activity of Taq polymerase cleaves the probe, thereby allowing the donor fluorophore to emit detectable fluorescence.
[0093] As used herein, the term "sample" or "test sample" refers to any liquid or solid material that contains nucleic acid. In suitable embodiments, the test sample is obtained from a biological source (i.e., a "biological sample"), such as a cell in culture or a tissue sample from an animal, preferably a human. In an exemplary embodiment, the sample is a tumor or liquid biopsy sample.
[0094] As used herein, "target nucleic acid" refers to a segment of a chromosome, a complete gene with or without intergenic sequences, segments or portions of a gene with or without intergenic sequences, or a sequence of a nucleic acid for which a probe or primer is designed. Target nucleic acid may include wild-type sequences, nucleic acid sequences containing mutations, deletions or matches, tandem repeat regions, genes of interest, regions of genes of interest or any upstream or downstream regions thereof. Target nucleic acid may represent alternative sequences or alleles of a particular gene. Target nucleic acid may be derived from genomic DNA, cDNA or RNA. As used herein, target nucleic acid may be native DNA or PCR amplification product.
[0095] As used herein, the term "stringency" is used in reference to the conditions of temperature, ionic strength, and the presence of other compounds under which nucleic acid hybridization is carried out. Under high stringency conditions, nucleic acid base pairing occurs only between nucleic acids that have sufficiently long segments with a high frequency of complementary base sequences. Exemplary hybridization conditions are as follows: High stringency generally refers to conditions that allow hybridization of only nucleic acid sequences that form stable hybrids in 0.018M NaCl at 65°C. High stringency conditions can be provided, for example, by hybridization in 50% formamide, 5x Denhardt's solution, 5x SSC (saline sodium citrate) 0.2% SDS (sodium dodecyl sulfate) at 42°C, followed by washing in 0.1x SSC and 0.1% SDS at 65°C. Moderate stringency refers to conditions equivalent to hybridization in 50% formamide, 5x Denhardt's solution, 5x SSC, 0.2% SDS at 42° C., followed by a wash in 0.2x SSC, 0.2% SDS at 65° C. Low stringency refers to conditions equivalent to hybridization in 10% formamide, 5x Denhardt's solution, 6x SSC, 0.2% SDS, followed by a wash in 1° SSC, 0.2% SDS at 50° C.
[0096] As used herein, the term "substantially identical" refers to a polypeptide or nucleic acid that exhibits at least 50%, 75%, 85%, 90%, 95%, or even 99% identity to a reference amino acid or nucleic acid sequence over a comparison region. For polypeptides, the length of the comparison sequence is generally at least 20, 30, 40, or 50 or more amino acids, or the full length of the polypeptide. For nucleic acids, the length of the comparison sequence is generally at least 10, 15, 20, 25, 30, 40, 50, 75, or 100 or more nucleotides, or the full length of the nucleic acid.
[0097] The "TP53 gene" or "tumor protein P53 gene" is a gene that provides instructions for making the tumor suppressor protein p53. The protein p53 plays a role in regulating cell division by preventing cells from growing or proliferating too fast. P53 binds directly to DNA when DNA damage is detected, and p53 determines whether the DNA is repaired or undergoes apoptosis. If the cell can be repaired, p53 activates DNA repair genes to repair the damage. P53 is important in preventing the development of tumors. Mutations in the TP53 gene are universal across cancer types. TP53 mutations have been correlated with the development of various cancers, including but not limited to breast cancer, bladder cancer, cholangiocarcinoma, lung cancer, melanoma, and ovarian cancer.
[0098] "EGFR gene" or "epidermal growth factor receptor gene" is a gene that codes for the EGFR protein. EGFR is a transmembrane glycoprotein and a protein kinase. Mutations in the EGFR gene are correlated with many types of cancer, including but not limited to non-small cell lung cancer, glioblastoma, and basal-like breast cancer. Tyrosine kinase inhibitors have shown efficacy in EGFR-amplified tumors. Thus, TK inhibitors may be an aggressive treatment for poor prognosis cancers that have EGFR as a marker for treatment.
[0099] The "BRAF gene" or "B-Raf proto-oncogene" is a gene that encodes the RAF serine / threonine protein kinase. BRAF plays a role in regulating cell division, differentiation, and secretion. Mutations in BRAF are also frequently correlated with cancer-causing mutations in melanoma and other forms of cancer.
[0100] The "KIT" gene (also known as c-Kit) encodes a receptor tyrosine kinase. As disclosed by the National Library of Medicine (https: / / www.ncbi.nlm.nih.gov / gene / 3815, last accessed December 12, 2022), this gene was initially identified as a homolog of the feline sarcoma viral oncogene v-kit and is often referred to as the proto-oncogene c-Kit. The canonical form of this glycosylated transmembrane protein has an N-terminal extracellular region with five immunoglobulin-like domains, a transmembrane region, and an intracellular tyrosine kinase domain at the C-terminus. Upon activation by its cytokine ligand, stem cell factor (SCF), the protein phosphorylates multiple intracellular proteins that play a role in the proliferation, differentiation, migration and apoptosis of many cell types, thereby playing an important role in hematopoiesis, stem cell maintenance, gametogenesis, melanogenesis, as well as the development, migration and function of mast cells. The protein can be a membrane-bound or soluble protein. Mutations in this gene are associated with gastrointestinal stromal tumors, mast cell disease, acute myeloid leukemia, and piebold syndrome. Multiple transcript variants have been found for this gene that encode different isoforms. See also Gene Cards (https: / / www.genecards.org / cgi-bin / carddisp.pl?gene=KIT, last accessed December 12, 2022).
[0101] The "KMT2C" gene is a member of the myeloid / lymphoid or mixed leukemia (MLL) family and encodes a nuclear protein with an AT-hook DNA-binding domain, a DHHC-type zinc finger, six PHD-type zinc fingers, a SET domain, a post-SET domain and a RING-type zinc finger. The protein is a member of the ASC-2 / NCOA6 complex (ASCOM), which has histone methylation activity and is involved in transcriptional coactivation. Sequence information for the gene and encoded protein can be found in GeneCards (https: / / www.genecards.org / cgi-bin / carddisp.pl?gene=KMT2C, last accessed December 12, 2022).
[0102] The "ELF3" gene enables DNA-binding transcription activator activity, RNA polymerase II-specific and sequence-specific double-stranded DNA binding activity. It is involved in inflammatory responses; negative regulation of transcription, DNA templates; and positive regulation of transcription by RNA polymerase II. It is located in the Golgi apparatus; in the cytosol; and in the nucleoplasm. Sequence information for the gene and its encoded protein can be found in GeneCards (https: / / www.genecards.org / cgi-bin / carddisp.pl?gene=ELF3, last accessed on December 12, 2022).
[0103] The "APC" gene encodes a tumor suppressor protein that acts as an antagonist of the Wnt signaling pathway. The APC gene is also involved in other processes, including cell migration and adhesion, transcriptional activation, and apoptosis. Defects in this gene cause familial adenomatous polyposis (FAP), an autosomal dominant premalignant disease that usually progresses to malignant tumors. Mutations in the APC gene have been found to occur in most colorectal cancers, where disease-associated mutations tend to cluster in small regions, termed mutation cluster regions (MCRs), resulting in truncated protein products. Sequence information for the gene and encoded protein can be found in GeneCards (https: / / www.genecards.org / cgi-bin / carddisp.pl?gene=APC, last accessed December 12, 2022).
[0104] The "AIRD1A" gene encodes a member of the SWI / SNF family, whose members have helicase and ATPase activity and are thought to regulate the transcription of certain genes by altering the chromatin structure around those genes. The encoded protein is part of the large ATP-dependent chromatin remodeling complex SNF / SWI, which is required for transcriptional activation of genes normally repressed by chromatin. It has at least two conserved domains that may be important for its function. Two transcript variants have been found for this gene that encode different isoforms. Sequence information for the gene and the encoded protein can be found in GeneCards (https: / / www.genecards.org / cgi-bin / carddisp.pl?gene=ARID1A, last accessed on December 12, 2022).
[0105] A "composition" typically contemplates a combination of an active agent, e.g., a compound or composition, with an inert (e.g., detectable agent or label) or active naturally occurring or non-naturally occurring carrier, such as an adjuvant, diluent, binder, stabilizer, buffer, salt, lipophilic solvent, preservative, adjuvant, and the like, and includes pharma- ceutically acceptable carriers. Carriers also include pharmaceutical excipients and additives proteins, peptides, amino acids, lipids, and carbohydrates (e.g., sugars, including monosaccharides, di-, tri-, tetra-oligosaccharides, and oligosaccharides; derivatized sugars, such as alditols, aldonic acids, esterified sugars, and the like; and polysaccharides or sugar polymers), which may be present alone or in combination, and comprise 1-99.99% by weight or volume, alone or in combination. Exemplary protein excipients include serum albumins, such as human serum albumin (HSA), recombinant human albumin (rHA), gelatin, casein, and the like. Representative amino acids / antibody components that may also function in a buffering capacity include alanine, arginine, glycine, arginine, betaine, histidine, glutamic acid, aspartic acid, cysteine, lysine, leucine, isoleucine, valine, methionine, phenylalanine, aspartame, etc. Carbohydrate excipients are also contemplated to be within the scope of this technology, examples of which include, but are not limited to, monosaccharides such as fructose, maltose, galactose, glucose, D-mannose, sorbose, etc.; disaccharides such as lactose, sucrose, trehalose, cellobiose, etc.; polysaccharides such as raffinose, melezitose, maltodextrin, dextran, starch, etc.; and alditols such as mannitol, xylitol, maltitol, lactitol, xylitol sorbitol (glucitol) and myo-inositol.
[0106] As used herein, the terms "nucleic acid sequence" and "polynucleotide" are used interchangeably to refer to any length of polymeric form of nucleotides, either ribonucleotides or deoxyribonucleotides.Thus, this term includes, but is not limited to, single-stranded, double-stranded or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers that contain purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural or derivatized nucleotide bases.
[0107] The term "encoding" as applied to a nucleic acid sequence refers to a polynucleotide that is said to "encode" a polypeptide if, in its native state, or when manipulated by methods well known to those of skill in the art, it can be transcribed and / or translated to produce an mRNA for the polypeptide and / or fragment thereof. The antisense strand is the complement of such a nucleic acid, from which the coding sequence can be deduced.
[0108] As used herein, the term "vector" refers to a nucleic acid construct designed for transfer between different hosts, including but not limited to plasmids, viruses, cosmids, phages, BACs, YACs, etc. In some embodiments, a plasmid vector can be prepared from a commercially available vector. In other embodiments, a viral vector can be produced from a baculovirus, retrovirus, adenovirus, AAV, etc., according to techniques known in the art. In one embodiment, the viral vector is a lentiviral vector. It should be understood that the vector contains regulatory elements necessary for the replication or expression of the inserted polynucleotide, including, for example, promoter or enhancer elements.
[0109] As used herein, the term "promoter" refers to any sequence that regulates the expression of a coding sequence, such as a gene. A promoter can be, for example, constitutive, inducible, repressible, or tissue-specific. A "promoter" is a control sequence that is a region of a polynucleotide sequence where the initiation and rate of transcription are controlled. A promoter can include genetic elements to which regulatory proteins and molecules, such as RNA polymerase and other transcription factors, can bind.
[0110] As used herein, the term "isolated cells" generally refers to cells that are substantially separated from other cells of a tissue. "Immune cells" include, for example, white blood cells (leukocytes), lymphocytes (T cells, B cells, natural killer (NK) cells), and bone marrow-derived cells (neutrophils, eosinophils, basophils, monocytes, macrophages, dendritic cells) derived from hematopoietic stem cells (HSCs) produced in the bone marrow. "T cells" include all types of immune cells that express CD3, including T helper cells (CD4+ cells), cytotoxic T cells (CD8+ cells), natural killer T cells, T regulatory cells (Tregs), and gamma delta T cells. "Cytotoxic cells" include CD8+ T cells, natural killer (NK) cells, and neutrophils, which are capable of mediating cytotoxic responses.
[0111] The term "transduce" or "transduction" as applied to the production of chimeric antigen receptor cells refers to the process by which an exogenous nucleotide sequence is introduced into a cell. In some embodiments, this transduction is accomplished via a vector.
[0112] As used herein, the term "autologous" with respect to cells refers to cells that are isolated and infused back into the same subject (recipient or host). "Allogeneic" refers to non-autologous cells.
[0113] "Effective amount" or "effective amount" refers to the amount of an agent, or a combined amount of two or more agents, that when administered for treatment of a mammal or other subject is sufficient to effect such treatment for a disease. An "effective amount" varies depending on the agent(s), the disease and its severity, and the age, weight, etc., of the subject being treated.
[0114] A "solid tumor" is an abnormal mass of tissue that usually does not contain cysts or liquid areas. Solid tumors can be benign or malignant. Various types of solid tumors are named for the type of cells that form them. Examples of solid tumors include sarcomas, carcinomas, and lymphomas.
[0115] As used herein, the term "label" refers to a directly or indirectly detectable compound or composition that is directly or indirectly conjugated to a composition to be detected to produce a "labeled" composition, such as an N-terminal histidine tag (N-His), a magnetically active isotope, e.g. 115 Sn, 117 Sn and 119 Sn, non-radioactive isotopes, e.g. 13 C and 15N, a polynucleotide or a protein, such as an antibody, is intended. The term also includes sequences conjugated to a polynucleotide that provide a signal upon expression of the inserted sequence, such as green fluorescent protein (GFP). The label may be detectable by itself (e.g., a radioisotope label or a fluorescent label) or, in the case of an enzymatic label, may catalyze a chemical change of a substrate compound or composition that is detectable. The label may be suitable for small-scale detection or may be more suitable for high-throughput screening. Suitable labels thus include, but are not limited to, magnetically active isotopes, non-radioactive isotopes, radioisotopes, fluorescent dyes, chemiluminescent compounds, dyes, and proteins, including enzymes. The label may be simply detected or quantified. A response that is simply detected generally includes a response whose presence is merely confirmed, whereas a response that is quantified generally includes a response that has a quantifiable (e.g., numerically reportable) value, such as intensity, polarization, and / or other properties. In luminescent or fluorescent assays, the detectable response may be generated directly using a luminophore or fluorophore associated with the assay component actually involved in binding, or indirectly using a luminophore or fluorophore associated with another (e.g., reporter or indicator) component. Examples of luminescent labels that generate a signal include, but are not limited to, bioluminescence and chemiluminescence. A detectable luminescent response generally involves a change or occurrence of a luminescent signal. Suitable methods and luminophores for luminescent labeling of assay components are known in the art and are described, for example, in Haugland, Richard P. (1996) Handbook of Fluorescent Probes and Research Chemicals (6 th Examples of luminescent probes include, but are not limited to, aequorin and luciferase.
[0116] Examples of suitable fluorescent labels include, but are not limited to, fluorescein, rhodamine, tetramethylrhodamine, eosin, erythrosine, coumarin, methylcoumarin, pyrene, Malacite green, stilbene, Lucifer Yellow, Cascade Blue™, and Texas Red. Other suitable optical dyes are described in Haugland, Richard P. (1996) Handbook of Fluorescent Probes and Research Chemicals (6 th (ed.).
[0117] In another embodiment, the fluorescent label is functionalized to facilitate covalent attachment to cellular components present in or on the surface of cells or tissues, such as cell surface markers. Suitable functional groups include, but are not limited to, isothiocyanate groups, amino groups, haloacetyl groups, maleimides, succinimidyl esters, and sulfonyl halides, all of which can be used to attach the fluorescent label to a second molecule. The choice of functional group of the fluorescent label depends on the site of attachment to either the linker, drug, marker, or second labeling agent.
[0118] As used herein, the term "immunoconjugate" includes an antibody or antibody derivative associated or linked to a second agent, such as a cytotoxic agent, a detectable agent, a radioactive agent, a targeting agent, a human antibody, a humanized antibody, a chimeric antibody, a synthetic antibody, a semi-synthetic antibody, or a multispecific antibody.
[0119] "Immune response" refers broadly to an antigen-specific response of lymphocytes to a foreign substance. The terms "immunogen" and "immunogenic" refer to a molecule that has the ability to induce an immune response. Although all immunogens are antigens, not all antigens are immunogenic. The immune response disclosed herein can be humoral (through antibody activity) or cell-mediated (through T cell activation). The response can occur in vivo or in vitro. Those skilled in the art will appreciate that a variety of macromolecules, including proteins, nucleic acids, fatty acids, lipids, lipopolysaccharides, and polysaccharides, have the potential to be immunogenic. Those skilled in the art will further appreciate that a nucleic acid that encodes a molecule capable of eliciting an immune response necessarily encodes an immunogen. Those skilled in the art will further appreciate that an immunogen is not limited to a full-length molecule, but can include a partial molecule.
[0120] A host cell can be a eukaryotic or prokaryotic cell. "Eukaryotic cells" includes all kingdoms of life except the kingdom Monera. They can be easily distinguished by a membrane-bound nucleus. Animals, plants, fungi and protists are organisms in which the eukaryotes or cells are organized into complex structures by internal membranes and a cytoskeleton. The most characteristic membrane-bound structure is the nucleus. Unless otherwise specified, the term "host" includes eukaryotic hosts including, for example, yeast, higher plants, insects and mammalian cells. Non-limiting examples of eukaryotic cells or hosts include monkeys, cows, pigs, mice, rats, birds, reptiles and humans.
[0121] "Prokaryotic cells" usually lack a nucleus or any other membrane-bound organelles and are divided into two domains: bacteria and archaea. In addition to chromosomal DNA, these cells can also contain genetic information in circular loops called episomes. Bacterial cells are very small, roughly the size of an animal mitochondrion (approximately 1-2 μm in diameter and 10 μm in length). Prokaryotic cells are characterized by three main shapes: rod-shaped, spherical, and spiral. Instead of undergoing an elaborate replication process like eukaryotes, bacterial cells divide by binary fission. Examples include, but are not limited to, Bacillus, E. coli, and Salmonella.
[0122] As used herein, the term "detectable marker" refers to at least one marker capable of directly or indirectly generating a detectable signal. This non-exhaustive list of markers includes, for example, enzymes that generate a signal detectable by colorimetry, fluorescence, luminescence such as horseradish peroxidase, alkaline phosphatase, β-galactosidase, glucose-6-phosphate dehydrogenase, chromophores such as fluorescent dyes, luminescent dyes, groups with electron density that are detected by electron microscopy or by electrical properties such as conductivity, amperometry, voltammetry, impedance, etc., such as detectable groups whose molecules are of sufficient size to induce detectable modifications in their physical and / or chemical properties, such detection being achieved by optical methods such as diffraction, surface plasmon resonance, surface fluctuations, contact angle changes, or atomic force spectroscopy, tunneling effect, or the like. 32 P, 35 S or 125 This can be achieved by physical methods such as by radioactive molecules such as I.
[0123] As used herein, the term "purification label" refers to at least one marker useful for purification or identification. A non-exhaustive list of this marker includes His, lacZ, GST, maltose binding protein, NusA, BCCP, c-myc, CaM, FLAG, GFP, YFP, cherry, thioredoxin, poly(NANP), V5, Snap, HA, chitin binding protein, Softag1, Softag3, Strep or S protein. Suitable direct or indirect fluorescent markers include FLAG, GFP, YFP, RFP, dTomato, cherry, Cy3, Cy5, Cy5.5, Cy7, DNP, AMCA, biotin, digoxigenin, Tamra, Texas Red, rhodamine, Alexa fluor, FITC, TRITC or any other fluorescent dye or hapten.
[0124] As used herein, the term "expression" refers to the process by which a polynucleotide is transcribed into mRNA and / or the transcribed mRNA is subsequently translated into a peptide, polypeptide or protein. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in eukaryotic cells. The expression level of a gene may be determined by measuring the amount of mRNA or protein in a cell or tissue sample. In one embodiment, the expression level of a gene from a sample may be directly compared to the expression level of that gene from a control or reference sample. In another embodiment, the expression level of a gene from a sample may be directly compared to the expression level of that gene from the same sample after administration of a compound.
[0125] As used herein, "homology" or "identical", percent "identity" or "similarity", when used in the context of two or more nucleic acid or polypeptide sequences, refers to two or more sequences or subsequences that are the same, or a specified percentage of nucleotides or amino acid residues that are the same, e.g., two or more sequences or subsequences that have at least 60% identity, preferably at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity over a specified region (e.g., a nucleotide sequence encoding an antibody described herein or an amino acid sequence of an antibody described herein). Homology can be determined by comparing positions in each sequence that may be aligned for purposes of comparison. If a position in the compared sequences is occupied by the same base or amino acid, then the molecules are homologous at that position. The degree of homology between sequences is a function of the number of matching or homologous positions shared by the sequences. Alignment and percent homology or percent sequence identity can be determined using software programs known in the art, such as those described in Current Protocols in Molecular Biology (Ausubel et al., eds. 1987), Supplement 30, section 7.7.18, table 7.7.1. Preferably, default parameters are used for alignment. A preferred alignment program is BLAST, using default parameters. In particular, preferred programs are BLASTN and BLASTP, using the following default parameters: genetic code=standard; filter=none; strand=both; cutoff=60; expectation=10; matrix=BLOSUM62; description=50 sequences; sort=high score; database=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProtein+SPupdate+PIR. Details of these programs can be found at the following internet address: ncbi.nlm.nih.gov / cgi-bin / BLAST.The terms "homology" or "identical", percent "identity" or "similarity" can also refer to or be applied to the complement of a test sequence. The terms also include sequences that have deletions and / or additions, as well as sequences that have substitutions. As described herein, preferred algorithms can account for gaps, etc. Preferably, the identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is at least 50-100 amino acids or nucleotides in length. An "unrelated" or "non-homologous" sequence shares less than 40% identity, or alternatively less than 25% identity, with one of the sequences disclosed herein.
[0126] "Administration" can be performed in one dose, continuously or intermittently throughout the course of treatment. Methods for determining the most effective means and dosages of administration are known to those of skill in the art and will vary with the composition used for the treatment, the purpose of the treatment, the target cells being treated, and the subject being treated. Single or multiple administrations can be performed with dose levels and patterns selected by the treating physician. Appropriate dosage formulations and methods for administering agents are known in the art. Routes of administration can also be determined, and methods for determining the most effective route of administration are known to those of skill in the art and will vary with the composition used for the treatment, the purpose of the treatment, the health or disease stage of the subject being treated, and the target cells or tissues. Non-limiting examples of routes of administration include oral administration, nasal administration, infusion, injection, and topical application. As will be appreciated by those of skill in the art, the treatment can be co-administered with other treatments, such as immuno-oncology or chemotherapy. The treatments can be administered simultaneously or in conjunction.
[0127] The phrases "first-line" or "second-line" or "third-line" refer to the order of treatments a patient has undergone. A first-line treatment regimen is the first treatment given, whereas a second-line or third-line treatment is given after a first-line treatment or after a second-line treatment, respectively. The National Cancer Institute defines first-line treatment as the "first treatment for a disease or condition." In patients with cancer, the first treatment can be surgery, chemotherapy, radiation therapy, or a combination of these therapies. First-line treatment is also referred to by those skilled in the art as "primary and primary treatment." See the National Cancer Institute website at www.cancer.gov, last visited May 1, 2008. Typically, a patient is given a subsequent chemotherapy regimen because the patient did not show a positive clinical or subclinical response to the first-line treatment or because the first-line treatment has been stopped.
[0128] In one aspect, the term "equivalent" or "biologically equivalent" of an antibody refers to the ability of an antibody to selectively bind to its epitope protein or fragment thereof as measured by ELISA or other suitable method. Biologically equivalent antibodies include, but are not limited to, antibodies, peptides, antibody fragments, antibody variants, antibody derivatives, and antibody mimetics that bind to the same epitope as the reference antibody.
[0129] Where the present disclosure relates to a polypeptide, protein, polynucleotide or antibody, it should be assumed, without express recitation and unless otherwise intended, that equivalents or biological equivalents of such are intended within the scope of the present disclosure. As used herein, the term "biological equivalent thereof" is intended to be synonymous with "equivalent thereof" when referring to a reference protein, antibody, polypeptide or nucleic acid, and is intended to have minimal homology while maintaining the desired structure or functionality. Unless specifically recited herein, any polynucleotide, polypeptide or protein referred to herein is also intended to include its equivalent. For example, an equivalent is intended to have at least about 70% homology or identity, or at least 80% homology or identity, and alternatively at least about 85%, or alternatively at least about 90%, or alternatively at least about 95%, or alternatively at least 98% percent homology or identity, and exhibits substantially equivalent biological activity to the reference protein, polypeptide or nucleic acid. Alternatively, when referring to a polynucleotide, the equivalent is a polynucleotide that hybridizes under stringent conditions to a reference polynucleotide or its complement.
[0130] A polynucleotide or polynucleotide region (or polypeptide or polypeptide region) having a certain percentage (e.g., 80%, 85%, 90%, or 95%) of "sequence identity" to another sequence means that, when aligned, that percentage of bases (or amino acids) are the same when comparing the two sequences. Alignment and percent homology or percent sequence identity can be determined using software programs known in the art, such as those described in Current Protocols in Molecular Biology (Ausubel et al., eds. 1987), Supplement 30, section 7.7.18, table 7.7.1. Preferably, default parameters are used for alignment. A preferred alignment program is BLAST, using default parameters. Particularly preferred programs are BLASTN and BLASTP using the following default parameters: genetic code=standard; filter=none; strand=both; cutoff=60; expectation=10; matrix=BLOSUM62; description=50 sequences; sort=high score; database=non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProtein+SPupdate+PIR. Details of these programs can be found at the following internet address: ncbi.nlm.nih.gov / cgi-bin / BLAST.
[0131] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex stabilized through hydrogen bonds between the bases of nucleotide residues. The hydrogen bonds can occur by Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific manner. The complex can include two strands forming a duplex structure, three or more strands forming a multistranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction can constitute a step in a more extensive process, such as the initiation of a PCR reaction or the enzymatic cleavage of a polynucleotide by a ribozyme.
[0132] Examples of stringent hybridization conditions include an incubation temperature of about 25°C to about 37°C; a hybridization buffer concentration of about 6xSSC to about 10xSSC; a formamide concentration of about 0% to about 25%; and a wash solution of about 4xSSC to about 8xSSC. Examples of moderate hybridization conditions include an incubation temperature of about 40°C to about 50°C; a buffer concentration of about 9xSSC to about 2xSSC; a formamide concentration of about 30% to about 50%; and a wash solution of about 5xSSC to about 2xSSC. Examples of high stringency conditions include an incubation temperature of about 55°C to about 68°C; a buffer concentration of about 1xSSC to about 0.1xSSC; a formamide concentration of about 55% to about 75%; and a wash solution of about 1xSSC, 0.1xSSC, or deionized water. Generally, hybridization incubation times are 5 minutes to 24 hours with one, two or more washing steps, and washing incubation times are about 1, 2, or 15 minutes. SSC is a 0.15M NaCl and 15mM citrate buffer. It is understood that equivalents of SSC using other buffer systems can be used.
[0133] "Normal cells corresponding to a tumor tissue type" refers to normal cells derived from the same tissue type as the tumor tissue. Non-limiting examples are normal lung cells from a patient with a lung tumor, or normal colon cells from a patient with a colon tumor.
[0134] As used herein, the term "isolated" refers to molecules or biologicals or cellular material that are substantially free of other materials. In one aspect, the term "isolated" refers to a nucleic acid, such as DNA or RNA, or a protein or polypeptide (e.g., an antibody or derivative thereof), or a cell or organelle, or a tissue or organ, that is separated from other DNA or RNA, or proteins or polypeptides, or cells or organelles, or tissues or organs, respectively, that are present in the natural source. The term "isolated" also refers to a nucleic acid or peptide that is substantially free of cellular material, viral material or medium if produced by recombinant DNA technology, or chemical precursors or other chemicals if chemically synthesized. Furthermore, "isolated nucleic acid" is meant to include nucleic acid fragments that are not naturally occurring as fragments and would not be found in the natural state. The term "isolated" is also used herein to refer to a polypeptide that is isolated from other cellular proteins, and is meant to encompass both purified and recombinant polypeptides. The term "isolated" is also used herein to refer to a cell or tissue that is isolated from other cells or tissues, and is meant to encompass both cultured and engineered cells or tissues.
[0135] As used herein, the term "monoclonal antibody" refers to an antibody produced by a single clone of B lymphocytes or by a cell transfected with the light and heavy chain genes of a single antibody. Monoclonal antibodies are produced by methods known to those skilled in the art, such as by creating hybrid antibody-forming cells from the fusion of myeloma cells and immune spleen cells. Monoclonal antibodies include humanized monoclonal antibodies.
[0136] The terms "protein", "peptide" and "polypeptide" are used interchangeably in their broadest sense to refer to a compound of two or more subunit amino acids, amino acid analogs or peptidomimetics. The subunits may be linked by peptide bonds. In alternative embodiments, the subunits may be linked by other bonds, such as esters, ethers, etc. A protein or peptide must contain at least two amino acids, and there is no limit to the maximum number of amino acids that may comprise a protein or peptide sequence. As used herein, the term "amino acid" refers to natural and / or unnatural or synthetic amino acids, including glycine and both D and L optical isomers, amino acid analogs and peptidomimetics.
[0137] The terms "polynucleotide" and "oligonucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides or their analogs. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, EST or SAGE tags), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, RNAi, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes and primers. Polynucleotides can contain modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after construction of the polynucleotide. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with a labeling component. The term also refers to both double-stranded and single-stranded molecules. Unless otherwise specified or required, any embodiment of this technology that is a polynucleotide encompasses both the double-stranded form and each of the two complementary single-stranded forms that are known or predicted to constitute the double-stranded form. As used herein, the term "purified" does not require absolute purity, but rather is intended as a relative term. Thus, for example, a purified nucleic acid, peptide, protein, biological complex or other active compound is one that has been wholly or partially isolated from proteins or other contaminants. Generally, a substantially purified peptide, protein, biological complex or other active compound for use within the present disclosure will contain greater than 80% of all macromolecular species present in the preparation prior to mixing or formulating the peptide, protein, biological complex or other active compound with pharmaceutical carriers, excipients, buffers, absorption enhancers, stabilizers, preservatives, adjuvants or other co-ingredients in a complete pharmaceutical formulation for therapeutic administration. More typically, the peptide, protein, biological complex or other active compound is purified to represent greater than 90%, and often greater than 95%, of all macromolecular species present in the purified preparation prior to mixing with other formulation components. In other cases, the purified preparation may be essentially homogeneous, with other macromolecular species undetectable by conventional techniques. Methods for carrying out the present disclosure
[0138] Treatment method
[0139] In one embodiment, disclosed is a method of inhibiting the growth of cancer cells in a subject or a method of treating cancer in a subject that needs to treat cancer, wherein the subject has clustered mutations in one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A gene(s) in a sample isolated from the subject, or has clustered mutations in two or more of them, or has clustered mutations in three or more of them, or has clustered mutations in four or more of them, or has clustered mutations in five or more of them, or has clustered mutations in six or more of them, or has clustered mutations in all seven of them, or has clustered mutations in BRAF gene, and / or has no clustered mutations in BRAF gene.The method comprises, consists of, or consists essentially of administering an active treatment to the subject, thereby inhibiting the growth of cancer cells in the subject or treating cancer in the subject.
[0140] The cancer cells can be animal or mammalian cells. Non-limiting examples of mammalian cells include human cells, non-human primate cells (e.g., apes, gibbons, chimpanzees, orangutans, monkeys, macaques, etc.), domestic animals (e.g., dogs and cats), farm animals (e.g., horses, cows, goats, sheep, pigs), and laboratory animal cells (e.g., mice, rats, rabbits, guinea pigs). In some embodiments, the cells are human cells. The mammal can be of any age or at any stage of development (e.g., adult, adolescent, child, infant, or mammal in utero). The mammal can be male or female. In some embodiments, the subject is a human. In some embodiments, the subject has cancer, is diagnosed with cancer, or is suspected of having cancer.
[0141] The subject can be any animal, typically a mammal. Any suitable mammal can be treated by the methods described herein. Non-limiting examples of mammals include humans, non-human primates (e.g., apes, gibbons, chimpanzees, orangutans, monkeys, macaques, etc.), domestic animals (e.g., dogs and cats), farm animals (e.g., horses, cows, goats, sheep, pigs), and laboratory animals (e.g., mice, rats, rabbits, guinea pigs). In some embodiments, the mammal is a human. The mammal can be of any age or at any stage of development (e.g., adult, adolescent, child, infant, or intrauterine mammal). The mammal can be male or female. In some embodiments, the subject is a human. In some embodiments, the subject has cancer, is diagnosed with cancer, or is suspected of having cancer.
[0142] In a further embodiment, the cancer cells or cancer is selected from carcinoma, sarcoma, or blood cancer. In yet a further embodiment, the cancer cells or cancer are located in the circulatory system, e.g., heart (sarcoma [angiosarcoma, fibrosarcoma, rhabdomyosarcoma, liposarcoma], myxoma, rhabdomyoma, fibroma, lipoma, and teratoma), mediastinum and pleura, and other intrathoracic organs, vascular tumors and tumor-associated vascular tissue; respiratory tract, e.g., nasal cavity and middle ear, paranasal sinuses, larynx, trachea, bronchi, and lungs (such as small cell lung cancer (SCLC), non-small cell lung cancer (NSCLC)), bronchogenic carcinoma (squamous, small undifferentiated cell, large undifferentiated cell, adenocarcinoma), alveolar (bronchiolar) carcinoma, bronchial adenoma, sarcoma, lymphoma, chondromatous hamartoma, mesothelioma; gastrointestinal system, e.g., esophagus (squamous cell carcinoma, adenocarcinoma, leiomyosarcoma, lymphoma), colon cancer, colorectal cancer, rectal cancer, stomach (carcinoma, lymphoma, leiomyosarcoma), gastric, pancreas (pancreatic ductal adenocarcinoma, insulinoma, glucagonoma, gastrinoma, carcinoid tumor, vipoma), small intestine (adenocarcinoma, lymphoma, carcinoid tumor, Karposi's sarcoma) sarcoma), leiomyoma, hemangioma, lipoma, neurofibroma, fibroma), colon (adenocarcinoma, tubular adenoma, villous adenoma, hamartoma, leiomyoma); gastrointestinal stromal tumors and neuroendocrine tumors occurring at any site; genitourinary tract, e.g. kidney (adenocarcinoma, Wilms' tumor [nephroblastoma], lymphoma, leukemia), bladder and / or urethra (squamous cell carcinoma, transitional cell carcinoma, adenocarcinoma), prostate (adenocarcinoma, sarcoma), testis (seminoma, teratoma, embryonal carcinoma, teratocarcinoma, choriocarcinoma, sarcoma, stromal cell carcinoma, fibroma, fibroadenoma, adenomatous tumor, lipoma; liver, e.g. hepatocellular carcinoma (hepatocellular carcinoma), cholangiocarcinoma, hepatoblastoma, angiosarcoma, hepatocellular adenoma, hemangioma, pancreatic endocrine tumors (pheochromocytoma, insulinoma, vasoactive intestinal peptide tumor, islet cell tumor and glucagonoma, etc.); bone, e.g. osteogenic sarcoma (osteosarcoma), fibrosarcoma, malignant fibrous histiocytoma, chondrosarcoma, Ewing's sarcoma, malignant lymphoma (reticulum cell sarcoma), multiple myeloma, malignant giant cell tumor, chordoma, osteochondroma (osteocartilaginous exostoses), benign chondroma, chondroblastoma, chondromyxofibroma, osteoid osteoma and giant cell tumor;Nervous system, e.g., neoplasms of the central nervous system (CNS), primary CNS lymphomas, skull cancer (osteoma, hemangioma, granuloma, xanthomas, osteitis deformans), meninges (meningioma, meningeal sarcoma, gliomatosis), brain cancer (astrocytoma, medulloblastoma, glioma, ependymoma, germinoma [pinealoma], glioblastoma multiforme, oligodendroglioma, schwannoma, retinoblastoma, congenital tumors), spinal neurofibroma, meningioma, glioma, sarcoma); reproductive system, e.g., gynecological system, uterus (endometrial carcinoma), cervix (cervical carcinoma, preneoplastic cervical dysplasia), ovary (ovarian carcinoma [serous cystadenocarcinoma, mucinous cystadenocarcinoma, unclassified carcinoma], granulosa-thecalcell tumor), Sertoli-Leydig cell tumor, dysgerminoma, malignant teratoma), vulva (squamous cell carcinoma, carcinoma in situ, adenocarcinoma, fibrosarcoma, melanoma), vagina (clear cell carcinoma, squamous cell carcinoma, botryoid sarcoma (embryonal rhabdomyosarcoma), fallopian tubes (carcinoma) and other sites associated with the female reproductive organs; placenta, penis, prostate, testes, and other sites associated with the male reproductive organs; blood system, e.g. blood (myeloid leukemia [acute and chronic], acute lymphoblastic leukemia, chronic lymphocytic leukemia, myeloproliferative disorders, multiple myeloma, myelodysplastic syndromes), Hodgkin's disease, non-Hodgkin's lymphoma [malignant lymphoma]; oral cavity, e.g. lips, tongue, gums, floor of the mouth, palate, and other parts of the mouth, parotid gland, and other parts of the salivary glands, tonsils, oropharynx, nasopharynx, piriform sinuses, hypopharynx, and other sites of the lips, oral cavity and pharynx; skin, e.g., malignant melanoma, cutaneous melanoma, basal cell carcinoma, squamous cell carcinoma, Kaposi's sarcoma, lentodysplastic nevi, lipomas, hemangiomas, dermatofibromas, and keloids; and connective and soft tissues, retroperitoneum and peritoneum, eye, intraocular melanoma, and other tissues, including adnexa, breast, head or / and neck, anal region, thyroid, parathyroid, adrenal glands and other endocrine glands and associated structures, secondary and unspecified malignant neoplasms of lymph nodes, secondary malignant neoplasms of the respiratory and digestive system and secondary malignant neoplasms of other sites;
[0143] Moreover, cancer can be primary cancer or metastatic cancer.Furthermore, sample can be cancer cell separated from tumor, peripheral blood sample or liquid biopsy.In further embodiment, clustering mutation or lack of clustering mutation in BRAF gene is specifically associated with certain cancer type, whether primary or metastatic (see, for example, Figure 11A and 11B).
[0144] Active treatments can be selected from adoptive cell therapy, immune checkpoint blockade including PD1, PD-L1, and CTLA4, pre-targeted radioimmunotherapy, oncolytic virotherapy, or cancer vaccines. It may also include TK inhibitors or combination chemotherapy (i.e., two or more drugs administered in combination). The specific treatment depends on the patient, the cancer of interest, and the cluster status.
[0145] In still further embodiments, the active chemotherapy optionally comprises one or more selected from a monoclonal antibody selected from a monospecific antibody, a bispecific antibody, a multispecific antibody, and a bispecific immune cell engager; an antibody-drug conjugate; optionally a CAR therapy selected from CAR NK therapy, CAR T therapy, CAR cytotoxic T therapy, CAR gamma delta T therapy, CAR NK therapy; a cell therapy; an inhibitor or antagonist of an inhibitory immune checkpoint; optionally an activator or agonist of a stimulatory immune checkpoint selected from an activating ligand; an immunomodulatory agent; a cancer vaccine; and an oncolytic virus therapy, optionally with a vector for delivering each of them to a subject.
[0146] In another embodiment, the aggressive chemotherapy comprises a checkpoint inhibitor.Non-limiting examples of such include GS4224, AMP-224, CA-327, CA-170, BMS-1001, BMS-1166, peptide-57, M7824, MGD013, CX-072, UNP-12, NP-12, or a combination of two or more thereof.
[0147] Further checkpoint inhibitors include one or more selected from anti-PD-1, anti-PD-L1, anti-CTLA-4, anti-LAG-3, anti-TIM-3, anti-TIGIT, anti-VISTA, anti-B7-H3, anti-BTLA, anti-ICOS, anti-GITR, anti-4-1BB, anti-OX40, anti-CD27, anti-CD28, anti-CD40, and anti-Siglec-15. In further embodiments, the checkpoint inhibitors include anti-PD1 or anti-PD-L1. In one embodiment, the anti-PD1 agent includes an anti-PD1 antibody or an antigen-binding fragment thereof. In further embodiments, the anti-PD1 antibody includes nivolumab, pembrolizumab, cemiplimab, spartalizumab, canrelizumab, sintilimab, tislelizumab, toripalimab, AMF 514, or a combination of two or more thereof. In another embodiment, the anti-PD-L1 agent comprises an anti-PD-L1 antibody or antigen-binding fragment thereof. In a further embodiment, the anti-PD-L1 antibody comprises avelumab, durvalumab, atezolizumab, emvafolimab, or a combination of two or more thereof.
[0148] In one embodiment, the checkpoint inhibitor comprises an anti-CTLA-4 agent. In another embodiment, the anti-CTLA-4 agent comprises an anti-CTLA-4 antibody or an antigen-binding fragment thereof. In yet a further embodiment, the anti-CTLA-4 antibody comprises ipilimumab, tremelimumab, zalifrelimab, or AGEN1181, or a combination thereof.
[0149] In one embodiment, the treatment further comprises surgical resection of the cancer, tumor or cancer cells. The treatment can be a first line, second line, third line, fourth line, fifth line treatment.
[0150] Prognostic and diagnostic methods
[0151] In yet another embodiment, a method for selecting cancer patients for aggressive treatment is disclosed.The method comprises, consists of, or consists essentially of, assaying and / or detecting at least one clustered mutation in a gene selected from one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A gene(s), or two or more of them, or three or more of them, or four or more of them, or five or more of them, or six or more of them, or all seven of them, or lacking clustered mutation in BRAF gene in sample isolated from the subject, and if clustered mutation is detected in sample isolated from the cancer patient, or if BRAF gene is not detected, the subject is selected for treatment.Non-limiting examples of aggressive treatment are disclosed herein and are incorporated herein by reference.
[0152] For example, active treatment can be selected from adoptive cell therapy, immune checkpoint blockade including PD1, PD-L1, and CTLA4, pre-targeted radioimmunotherapy, oncolytic virotherapy, or cancer vaccine. It can also include TK inhibitors or combination chemotherapy (i.e., two or more drugs administered in combination). The specific treatment depends on the patient, the cancer of interest, and the cluster status.
[0153] In still further embodiments, the active chemotherapy optionally comprises one or more selected from a monoclonal antibody selected from a monospecific antibody, a bispecific antibody, a multispecific antibody, and a bispecific immune cell engager; an antibody-drug conjugate; optionally a CAR therapy selected from CAR NK therapy, CAR T therapy, CAR cytotoxic T therapy, CAR gamma delta T therapy, CAR NK therapy; a cell therapy; an inhibitor or antagonist of an inhibitory immune checkpoint; optionally an activator or agonist of a stimulatory immune checkpoint selected from an activating ligand; an immunomodulatory agent; a cancer vaccine; and an oncolytic virus therapy, optionally with a vector for delivering each of them to a subject.
[0154] In another embodiment, the aggressive chemotherapy comprises a checkpoint inhibitor.Non-limiting examples of such include GS4224, AMP-224, CA-327, CA-170, BMS-1001, BMS-1166, peptide-57, M7824, MGD013, CX-072, UNP-12, NP-12, or a combination of two or more thereof.
[0155] Further checkpoint inhibitors include one or more selected from anti-PD-1, anti-PD-L1, anti-CTLA-4, anti-LAG-3, anti-TIM-3, anti-TIGIT, anti-VISTA, anti-B7-H3, anti-BTLA, anti-ICOS, anti-GITR, anti-4-1BB, anti-OX40, anti-CD27, anti-CD28, anti-CD40, and anti-Siglec-15. In further embodiments, the checkpoint inhibitors include anti-PD1 or anti-PD-L1. In one embodiment, the anti-PD1 agent includes an anti-PD1 antibody or an antigen-binding fragment thereof. In further embodiments, the anti-PD1 antibody includes nivolumab, pembrolizumab, cemiplimab, spartalizumab, canrelizumab, sintilimab, tislelizumab, toripalimab, AMF 514, or a combination of two or more thereof. In another embodiment, the anti-PD-L1 agent comprises an anti-PD-L1 antibody or an antigen-binding fragment thereof. In a further embodiment, the anti-PD-L1 antibody comprises avelumab, durvalumab, atezolizumab, emvafolimab, or a combination of two or more thereof. In one embodiment, the checkpoint inhibitor comprises an anti-CTLA-4 agent. In another embodiment, the anti-CTLA-4 agent comprises an anti-CTLA-4 antibody or an antigen-binding fragment thereof. In yet a further embodiment, the anti-CTLA-4 antibody comprises ipilimumab, tremelimumab, zalifrelimab, or AGEN1181, or a combination thereof. In yet another embodiment, a method is disclosed for identifying whether a cancer patient is likely to experience a relatively long or relatively short overall survival.the method comprising, consisting of or consisting essentially of assaying and / or detecting at least one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A gene(s), or alternatively two or more of them, or alternatively three or more of them, or alternatively four or more of them, or alternatively five or more of them, or alternatively six or more of them, or alternatively all seven of them, or lacking clustered mutations in the BRAF gene in a sample isolated from the patient; If clustered mutations are detected in BRAF, the patient may experience a longer overall survival; if clustered mutations are detected in at least one or more, or alternatively two or more, or alternatively three or more, or alternatively four or more, or alternatively five or more, or alternatively six or more, or alternatively all seven of the TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A gene(s), the patient may experience a shorter overall survival.
[0156] The subject can be any animal, typically a mammal. Any suitable mammal can be treated by the methods described herein. Non-limiting examples of mammals include humans, non-human primates (e.g., apes, gibbons, chimpanzees, orangutans, monkeys, macaques, etc.), domestic animals (e.g., dogs and cats), farm animals (e.g., horses, cows, goats, sheep, pigs), and laboratory animals (e.g., mice, rats, rabbits, guinea pigs). In some embodiments, the mammal is a human. The mammal can be of any age or at any stage of development (e.g., adult, adolescent, child, infant, or intrauterine mammal). The mammal can be male or female. In some embodiments, the subject is a human. In some embodiments, the subject has cancer, is diagnosed with cancer, or is suspected of having cancer.
[0157] In further embodiments, the cancer cells or cancers are selected from carcinomas, sarcomas, or hematological cancers. In still further embodiments, the cancer cells or cancers are located in the circulatory system, e.g., heart (sarcomas [angiosarcoma, fibrosarcoma, rhabdomyosarcoma, liposarcoma], myxoma, rhabdomyoma, fibroma, lipoma, and teratoma), mediastinum and pleura, and other intrathoracic organs, vascular tumors and tumor-associated vascular tissue; respiratory tract, e.g., nasal cavity and middle ear, paranasal sinuses, larynx, trachea, bronchi, and lungs (such as small cell lung cancer (SCLC), non-small cell lung cancer (NSCLC)), bronchogenic carcinoma (squamous, undifferentiated, etc.), small cell, large cell anaplastic, adenocarcinoma), alveolar (bronchiolar) carcinoma, bronchial adenoma, sarcoma, lymphoma, chondromatous hamartoma, mesothelioma; gastrointestinal system, e.g., esophagus (squamous cell carcinoma, adenocarcinoma, leiomyosarcoma, lymphoma), stomach (carcinoma, lymphoma, leiomyoma), colon cancer, colorectal cancer, rectal cancer, gastric, pancreas (pancreatic ductal adenocarcinoma, insulinoma, glucagonoma, gastrinoma, carcinoid tumor, vipoma), small intestine (adenocarcinoma, lymphoma, calc tinoid tumor, Kaposi's sarcoma, leiomyoma, hemangioma, lipoma, neurofibroma, fibroma), colon (adenocarcinoma, tubular adenoma, villous adenoma, hamartoma, leiomyoma); gastrointestinal stromal tumors and neuroendocrine tumors occurring at any site; genitourinary tract, e.g. kidney (adenocarcinoma, Wilms' tumor [nephroblastoma], lymphoma, leukemia), bladder and / or urethra (squamous cell carcinoma, transitional cell carcinoma, adenocarcinoma), prostate (adenocarcinoma, sarcoma), testis (seminoma, teratoma, embryonal carcinoma, teratocarcinoma, choriocarcinoma, sarcoma), tumors, interstitial cell carcinoma, fibroma, fibroadenoma, adenomatous tumors, lipomas; liver, e.g., hepatocellular carcinoma (hepatocellular carcinoma), cholangiocarcinoma, hepatoblastoma, angiosarcoma, hepatocellular adenoma, hemangioma, pancreatic endocrine tumors (pheochromocytoma, insulinoma, vasoactive intestinal peptide tumor, islet cell tumor and glucagonoma, etc.); bone, e.g., osteogenic sarcoma (osteosarcoma), fibrosarcoma, malignant fibrous histiocytoma, chondrosarcoma, Ewing's sarcoma, malignant lymphoma (reticulum cell sarcoma), multiple myeloma, malignant giant cell tumor chordoma, osteochondroma (osteochondroma exostosis), benign chondroma, chondroblastoma, chondromyxoid fibroma, osteoid osteoma and giant cell tumor;Nervous system, e.g., neoplasms of the central nervous system (CNS), primary CNS lymphoma, skull cancer (osteoma, hemangioma, granuloma, xanthomas, osteitis deformans), meninges (meningioma, meningeal sarcoma, gliomatosis), brain cancer (astrocytoma, medulloblastoma, glioma, ependymoma, germinoma [pinealoma], glioblastoma multiforme, oligodendroglioma, schwannoma, retinoblastoma, congenital tumors), spinal neurofibroma, meningioma, glioma, sarcoma); reproductive system, e.g., gynecological system, uterus (endometrial cancer), cervix (cervical cancer) , preneoplastic cervical dysplasia), ovaries (ovarian cancer [serous cystadenocarcinoma, mucinous cystadenocarcinoma, unclassified carcinoma], granulosa theca cell tumor, Sertoli-Leydig cell tumor, dysgerminoma, malignant teratoma), vulva (squamous cell carcinoma, carcinoma in situ, adenocarcinoma, fibrosarcoma, melanoma), vagina (clear cell carcinoma, squamous cell carcinoma, sarcoma botryoides (embryonal rhabdomyosarcoma), fallopian tubes (carcinoma) and other sites related to the female reproductive tract; placenta, penis, prostate, testes, and other sites related to the male reproductive tract; blood system, e.g. lymphoma (myeloid leukemia [acute and chronic], acute lymphoblastic leukemia, chronic lymphocytic leukemia, myeloproliferative disorders, multiple myeloma, myelodysplastic syndromes), Hodgkin's disease, non-Hodgkin's lymphoma [malignant lymphoma]; oral cavity, e.g., lips, tongue, gums, floor of the mouth, palate, and other parts of the mouth, parotid glands, and other parts of the salivary glands, tonsils, oropharynx, nasopharynx, piriform sinus, hypopharynx, and other parts of the lips, oral cavity, and pharynx; skin, e.g., malignant melanoma, cutaneous melanoma, basal cell carcinoma, squamous cell carcinoma, epithelial carcinoma, Kaposi's sarcoma, melanocortical dysplastic nevi, lipoma, hemangioma, dermatofibroma, and keloids; and secondary and unspecified malignant neoplasms of connective and soft tissues, retroperitoneum and peritoneum, eye, intraocular melanoma, and other tissues including adnexa, breast, head and / or neck, anal region, thyroid, parathyroid, adrenal and other endocrine glands and associated structures, lymph nodes, respiratory and digestive system and other sites;
[0158] The cancer may be primary or metastatic. In some embodiments, the clustered mutations are specifically associated with primary or metastatic cancer (see, e.g., Figures 11A and 11B).
[0159] Any suitable method for identifying genotype in patient samples can be used, and the disclosure described herein is not limited to these methods.For illustrative purposes only, genotype is determined by a method that includes, consists essentially of, or even further consists of sequencing, hybridization, polymerase chain reaction (PCR), real-time PCR, reverse transcriptase PCR (RT-PCR), nested PCR, ligase chain reaction, or nucleic acid amplification, including PCR-RFLP, or microarray.These methods and their equivalents or alternatives are described herein.
[0160] The information obtained using the diagnostic assays described herein is useful for determining whether a subject is likely, likely, or unlikely to respond to a given type of cancer treatment. Based on the prognostic information, a physician can recommend a therapeutic protocol useful for treating a patient to reduce a malignant mass or tumor or to treat an individual's cancer.
[0161] Furthermore, knowledge of the identity of specific alleles (genetic profile) in an individual allows customization of treatment for a particular disease to the individual's genetic profile, which is the goal of "pharmacogenomics". For example, an individual's genetic profile may allow a physician to 1) more effectively prescribe drugs that address the molecular basis of a disease or condition, 2) better determine the appropriate dosage of a particular drug, and 3) identify novel targets for drug development. The identity of an individual patient's genotype or expression pattern can then be compared to the disease's genotype or expression profile to determine the appropriate drug and dosage to administer to the patient.
[0162] The ability to target populations predicted to show the greatest clinical benefit based on their normal or disease genetic profiles may enable 1) repositioning of marketed drugs with disappointing market results, 2) salvage of drug candidates discontinued in clinical development as a result of safety or efficacy limitations that are patient subgroup specific, and 3) expedited and cheaper development of drug candidates and more optimal drug labeling.
[0163] Biological sample collection and preparation
[0164] The methods and compositions disclosed herein can be used to detect nucleic acids associated with the genetic polymorphisms identified herein using biological samples obtained from patients. Biological samples can be obtained by standard procedures and can be used immediately or stored under conditions appropriate to the type of biological sample for later use. Any liquid or solid biological material obtained from a patient that is believed to contain nucleic acids that include the region of the polymorphic region can be a suitable sample. The sample can be a tumor sample, a peripheral blood sample or a liquid biopsy.
[0165] Methods for obtaining a test sample are known to those of skill in the art and include, but are not limited to, aspiration, tissue dissection, swab, drawing of blood or other fluid, surgical or needle biopsy.
[0166] In some embodiments, the biological sample is a tissue or cell sample.Suitable patient samples in the present method include, but are not limited to, blood, plasma, serum, biopsy tissue, fine needle biopsy sample, amniotic fluid, plasma, pleural fluid, saliva, semen, serum, tissue or tissue homogenate, frozen or paraffin sections of tissue, or combinations thereof.In some embodiments, the biological sample comprises, or alternatively consists essentially of, or even consists of at least one of tumor cells, normal cells adjacent to the tumor, normal cells corresponding to the tumor tissue type, blood cells, peripheral blood lymphocytes, or combinations thereof.In some embodiments, the biological sample is a recently isolated original sample from a patient, fixed tissue, frozen tissue, resected tissue, or microdissected tissue.In some embodiments, the biological sample is processed by tissue sectioning, fractionation, purification, nucleic acid isolation, or cell organelle separation, etc.
[0167] In some embodiments, nucleic acid (DNA or RNA) is isolated from the sample according to any method known to those skilled in the art. In some aspects, genomic DNA is isolated from the biological sample. In some aspects, RNA is isolated from the biological sample. In some aspects, cDNA is generated from mRNA in the sample. In some embodiments, nucleic acid is not isolated from the biological sample (e.g., polymorphisms are detected directly from the biological sample).
[0168] Detection of clustered mutations and polymorphisms
[0169] Methods for detecting or assaying clustered mutations are known in the art and described herein.In some embodiments, detection of clustered mutations or polymorphisms can be achieved in some aspects by molecular cloning of specific alleles and then sequencing of the alleles using techniques known in the art after isolation of suitable nucleic acid samples.In some embodiments, gene sequences can be directly amplified from genomic DNA preparations from biological samples using PCR, and sequence composition is determined by sequencing the amplified products (i.e., amplicons).Alternatively, PCR products can be analyzed after digestion with restriction enzymes, a method known as PCR-RFLP.
[0170] In some embodiments, the clustered mutations or polymorphisms are detected using allele-specific hybridization using a probe that matches the polymorphic site. In some aspects, the nucleic acid probe is 5-40 nucleotides in length. In some aspects, the nucleic acid probe is about 5, about 10, about 15, about 20, about 25, about 30, about 35, or about 40 or more nucleotides adjacent to the polymorphic site.
[0171] In another embodiment of the present disclosure, several nucleic acid probes that can specifically hybridize to nucleic acids containing allelic variants are attached to a solid support, such as a "chip" or "microarray". Such gene chips or microarrays can be used to detect genetic variations by several techniques known to those skilled in the art. In one technique, oligonucleotides are arranged on the gene chip to determine DNA sequences by sequencing by hybridization methods. The probes of the present disclosure can also be used for fluorescent detection of gene sequences. The probes can also be immobilized on an electrode surface for electrochemical detection of nucleic acid sequences.
[0172] In one embodiment, a "gene chip" or "microarray" containing probes or primers for genes of interest is provided, alone or in combination with other probes and / or primers. A suitable sample is obtained from a patient extract of genomic DNA, RNA, or any combination thereof, and optionally amplified. The DNA or RNA sample is contacted with the gene chip or microarray panel under conditions suitable for hybridization of the gene(s) of interest with the probe(s) or primer(s) contained in the gene chip or microarray. The probe or primer can be detectably labeled, thereby identifying polymorphisms in the gene(s) of interest. Alternatively, chemical or biological reactions can be used to identify the probe or primer that has hybridized with the DNA or RNA of the gene(s) of interest. The patient's genetic profile is then determined using the above-mentioned devices and methods.
[0173] In some embodiments, whole genome sequencing can be used to obtain the genotype of related polymorphisms, particularly whole genome sequencing using "next generation sequencing" technology that uses massive parallel sequencing of DNA templates.Exemplary NGS sequencing platforms for generating nucleic acid sequence data include, but are not limited to, Illumina's sequencing by synthesis technology (e.g., Illumina MiSeq or HiSeq system), Life Technologies' Ion Torrent semiconductor sequencing technology (e.g., Ion Torrent PGM or Proton system), Roche (454 Life Sciences) GS series and Qiagen (Intelligent BioSystems) Gene Reader sequencing platform.
[0174] In some embodiments, a nucleic acid that contains or alternatively consists essentially of, or even further consists of, a polymorphism is amplified to produce an amplicon that contains the polymorphism. The nucleic acid can be amplified by various methods known to those skilled in the art. Nucleic acid amplification can be linear or exponential. Amplification is generally performed using polymerase chain reaction (PCR) technology. Alternative or modified PCR amplification methods can also be used, such as isothermal amplification, rolling circle PCR, hot start PCR, real-time PCR, allele-specific PCR, assembly PCR or polymerase cycling assembly (PCA), asymmetric PCR, colony PCR, emulsion PCR, fast PCR, real-time PCR, nucleic acid ligation, gap ligation chain reaction (Gap LCR), ligation-mediated PCR, multiplex ligation-dependent probe amplification (MLPA), gap extension ligation PCR (GEXL-PCR), quantitative PCR (Q-PCR), quantitative real-time PCR (QRT-PCR), multiplex PCR, helicase-dependent amplification, intersequence specific (ISSR) PCR, inverse PCR, linear-after-the-exponential-PCR (LATE-PCR), methylation-specific PCR (MSP), nested PCR, overlap extension PCR, PAN-AC assay, reverse transcription PCR (RT-PCR), rapid amplification of cDNA ends (RACE PCR), single molecule amplification PCR (SMAR). Examples of methods include PCR, thermal asymmetric interlaced PCR (TAIL-PCR), touchdown PCR, long PCR, nucleic acid sequencing (including DNA sequencing and RNA sequencing), transcription, reverse transcription, replication, DNA or RNA ligation, and other nucleic acid extension reactions known in the art. Those skilled in the art will appreciate that other methods can be used in place of or in conjunction with the PCR method, including enzymatic replication reactions that will be developed in the future.See, e.g., Saiki, "Amplification of Genomic DNA" in PCR Protocols, Innis et al., eds., Academic Press, San Diego, Calif., pp. 13-20 (1990); Wharam et al., 29(11) Nucleic Acids Res, E54-E54 (2001); Hafner et al., 30(4) Biotechniques, pp. 852-6, 858, 860 (2001).
[0175] In some embodiments, a nucleic acid that contains a polymorphism of interest, or alternatively consists essentially of, or further consists of, a polymorphism of interest, is amplified to produce an amplicon. In some embodiments, a nucleic acid containing a region of interest is amplified using a forward primer and a reverse primer that flank the region of interest. In some embodiments, an amplicon that contains a region of interest (e.g., an amplicon having a polymorphic sequence) is detected by hybridizing a nucleic acid probe that contains a polymorphism or its complement to the corresponding complementary strand of the amplicon and detecting the hybrid formed between the nucleic acid probe and the complementary strand of the amplicon. In some embodiments, an amplicon that contains a region of interest is sequenced (e.g., dideoxy chain termination method (Sanger method and its variations), Maxam & Gilbert sequencing, pyrosequencing, exonuclease digestion and next-generation sequencing).
[0176] In some embodiments, the amplification comprises a labeled primer or probe, thereby allowing detection of an amplification product corresponding to that primer or probe. In certain embodiments, the amplification can comprise multiple labeled primers or probes. Such primers can be distinguishably labeled, allowing simultaneous detection of multiple amplification products.
[0177] In some embodiments, the amplification products are detected by any of a number of methods, such as gel electrophoresis, column chromatography, hybridization with a nucleic acid probe, or sequencing of the amplicon.
[0178] Detectable labels can be used to identify primers or probes that have hybridized to the genomic nucleic acid or amplicon. Detectable labels include fluorophores, isotopes (e.g., 32 P, 33 P, 35 S, 3 H, 14 C. 125 I, 131 I) Electron-dense reagents (e.g., gold, silver), nanoparticles, enzymes commonly used in ELISAs (e.g., horseradish peroxidase, beta-galactosidase, luciferase, alkaline phosphatase), chemiluminescent compounds, colorimetric labels (e.g., colloidal gold), magnetic labels (e.g., Dynabeads®), biotin, digoxigenin, haptens, proteins for which antisera or monoclonal antibodies are available, ligands, hormones, oligonucleotides capable of forming complexes with their corresponding oligonucleotide complements.
[0179] In one embodiment, primer or probe is labeled with a fluorophore that emits detectable signals.As used herein, the term "fluorophore" refers to a molecule that absorbs light of a specific wavelength (excitation frequency) and then emits light of a longer wavelength (emission frequency).Suitable reporter dyes are fluorescent dyes, but any reporter dye that can be attached to detection reagents such as oligonucleotide probes or primers is suitable for use in the described method. Suitable fluorescent moieties include, but are not limited to, the following fluorophores acting individually or in combination: 4-acetamido-4'-isothiocyanatostilbene-2,2'disulfonic acid; acridine and derivatives such as acridine, acridine isothiocyanate; Alexa Fluors: Alexa Fluor® 350, Alexa Fluor® 488, Alexa Fluor® 546, Alexa Fluor® 555, Alexa Fluor® 568, Alexa Fluor® 594, Alexa Fluor® 647 (Molecular Probes); 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS); 4-amino-N-[3-vinylsulfonyl)phenyl]naphthalimide-3,5-disulfonate (Lucifer Yellow VS); N-(4-anilino-1-naphthyl)maleimide; anthranilamide; Black Hole Quencher™ (BHQ™) dyes (biosearch Technologies); BODIPY dyes: BODIPY® R-6G, BODIPY® 530 / 550, BODIPY® FL; Brilliant Yellow; Coumarins and derivatives: Coumarin, 7-amino-4-methylcoumarin (AMC, Coumarin 120), 7-amino-4-trifluoromethylcouluarin (Coumarin 151); Cy2®, Cy3®, Cy3.5®, Cy5®, Cy5.5®;cyanosine;4',6-diaminidino-2-phenylindole (DAPI);5',5''-dibromopyrogallol-sulfonphthalein (Bromopyrogallol Red);7-diethylamino-3-(4'-isothiocyanatophenyl)-4-methylcoumarin;diethylenetriaminepentaacetate;4,4'-diisothiocyanatodihydro-stilbene-2,2'-disulfonic acid;4,4'-diisothiocyanatostilbene-2,2'-disulfonic acid;5-[dimethylamino]naphthalene-1-sulfonyl chloride (DNS, dansyl chloride);4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL);4-dimethylaminophenylazophenyl-4'-isothiocyanate (DABITC);Eclipse™ (Epoch Biosciences Inc.);Eosin and derivatives: Eosin, Eosin isothiocyanate;Erythrosine and derivatives: Erythrosine B, Erythrosine isothiocyanate;Ethidium;Fluorescein and derivatives: 5-Carboxyfluorescein (FAM), 5-(4,6-dichlorotriazin-2-yl)aminofluorescein (DTAF), 2',7'-dimethoxy-4'5'-dichloro-6-carboxyfluorescein (JOE), Fluorescein, Fluorescein isothiocyanate (FITC), Hexachloro-6-carboxyfluorescein (HEX), QFITC (XRITC), Tetrachlorofluorescein (TET);Fluorescamine;IR144;IR1446;Lanthamide fluorophores;Malachite Green isothiocyanate;4-Methylumbelliferone;Orthocresolphthalein;Nitrotyrosine;Pararosaniline;Phenol Red; B-phycoerythrin, R-phycoerythrin; allophycocyanin; o-phthaldialdehyde; Oregon Green®; propidium iodide; pyrene and derivatives: pyrene, pyrene butyrate, succinimidyl 1-pyrene butyrate; QSY® 7; QSY® 9; QSY® 21; QSY® 35 (Molecular Probes); Reactive Red 4 (Cibacron® Brilliant Red 3B-A); Rhodamine and derivatives: 6-carboxy-X-rhodamine (ROX), 6-carboxyrhodamine (R6G), Lissamine rhodamine B sulfonyl chloride, rhodamine (Rhod), rhodamine B, rhodamine 123, rhodamine green, rhodamine X isothiocyanate, riboflavin, rosolic acid, sulforhodamine B, sulforhodamine 101, sulfonyl chloride derivative of sulforhodamine 101 (Texas Red); Terbium chelate derivatives; N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA); tetramethylrhodamine; and tetramethylrhodamine isothiocyanate (TRITC).
[0180] In some embodiments, the primer or probe is further labeled with a quencher dye such as Tamra, Dabcyl, or Black Hole Quencher® (BHQ), particularly when the reagent is used as a self-quenching probe such as TaqMan® (U.S. Pat. Nos. 5,210,015 and 5,538,848) or Molecular Beacon probe (U.S. Pat. Nos. 5,118,801 and 5,312,728), or other stemless or linear beacon probe (Livak et al., 1995, PCR Method Appl., 4:357-362; Tyagi et al., 1996, Nature Biotechnology, 14:303-308; Nazarenko et al., 1997, Nucl. Acids Res., 25:2516-2521; U.S. Pat. Nos. 5,866,336 and 6,117,635).
[0181] In some embodiments, methods for real-time PCR use fluorescent primers / probes such as TaqMan® primers / probes (Heid, et al., Genome Res 6:986-994, 1996), molecular beacons, and Scorpion™ primers / probes. Real-time PCR quantifies the initial amount of template with greater specificity, sensitivity, and reproducibility than other forms of quantitative PCR, which detect the amount of final amplification product. Real-time PCR does not detect the size of the amplicon. The probes used in Scorpion®™ and TaqMan® technologies are based on the principle of fluorescence quenching and include a donor fluorophore and a quenching moiety. As used herein, the term "donor fluorophore" refers to a fluorophore that, when in close proximity to a quencher moiety, donates or transfers emission energy to the quencher. As a result of donating energy to the quencher moiety, the donor fluorophore itself emits less light at a particular emission frequency than it would have in the absence of a quencher moiety located in close proximity. As used herein, the term "quencher moiety" refers to a molecule that, in close proximity to a donor fluorophore, captures the luminescence energy generated by the donor and either dissipates the energy as heat or emits light at a wavelength longer than the emission wavelength of the donor. In the latter case, the quencher is considered to be an acceptor fluorophore. Quenching moieties can act via proximal (i.e., collisional) quenching or by Forster or fluorescence resonance energy transfer ("FRET"). Quenching by FRET is commonly used in TaqMan® primers / probes, and proximal quenching is used in molecular beacons and Scorpion™ type primers / probes.
[0182] Detectable label can be incorporated into, associated with, or conjugated to nucleic acid primer or probe.Label can be attached by spacer arm of various lengths to reduce potential steric hindrance or influence on other useful or desired properties.See, for example, Mansfield, Mol.Cell.Probes (1995), 9:145-156.
[0183] Detectable label can be incorporated into nucleic acid probe by covalent or non-covalent means, for example, by transcription, random primer labeling, for example, using Klenow polymerase, or nick translation, or amplification, or equivalents known in the art.For example, nucleotide bases are conjugated to detectable moieties such as fluorescent dyes, for example, Cy3™ or Cy5™, and then incorporated into nucleic acid probe during nucleic acid synthesis or amplification.Therefore, nucleic acid probe can be labeled when synthesized using Cy3™- or Cy5™-dCTP conjugates mixed with unlabeled dCTP.
[0184] Nucleic acid probes can be labeled by using PCR or nick translation in the presence of labeled precursor nucleotides, for example modified nucleotides synthesized by coupling allylamine-dUTP to succinimidyl ester derivatives of fluorescent dyes or haptens such as biotin or digoxigenin, which methods allow for custom preparation of the most common fluorescent nucleotides (see, e.g., Henegariu et al., Nat. Biotechnol. (2000), 18:345-348).
[0185] Nucleic acid probes can be labeled by non-covalent means known in the art. For example, Kreatech Biotechnology's Universal Linkage System® (ULS®) provides a non-enzymatic labeling technique in which platinum groups form coordinate bonds with DNA, RNA or nucleotides by binding to the N7 position of guanosine. This technique can also be used to label proteins by binding to the nitrogen- and sulfur-containing side chains of amino acids. See, for example, U.S. Pat. Nos. 5,580,990, 5,714,327, and 5,985,566, and European Patent No. 0539466.
[0186] Labeling with detectable labels can also include nucleic acids bound to another biological molecule, such as nucleic acids, such as oligonucleotides, or nucleic acids in the form of stem-loop structures as "molecular beacons" or "aptamer beacons". Molecular beacons as detectable moieties have been described, for example, Sokol (Proc. Natl. Acad. Sci. USA (1998), 95:11538-11543) synthesized "molecular beacon" reporter oligodeoxynucleotides with matching fluorescent donor and acceptor chromophores at their 5' and 3' ends. In the absence of a complementary nucleic acid strand, the molecular beacon remains in a stem-loop conformation where fluorescence resonance energy transfer prevents signal emission. Hybridization with a complementary sequence opens the stem-loop structure, increasing the physical distance between the donor and acceptor moieties, thereby decreasing fluorescence resonance energy transfer and allowing the beacon to emit a detectable signal when excited by light of the appropriate wavelength. See, for example, Antony (Biochemistry (2001), 40:9387-9395), which describes molecular beacons consisting of triplexes of G-rich 18-mer forming oligodeoxyribonucleotides. See also U.S. Patent Nos. 6,277,581 and 6,235,504.
[0187] Aptamer beacons are similar to molecular beacons, see, e.g., Hamaguchi, Anal. Biochem. (2001), 294:126-131; Poddar, Mol. Cell. Probes (2001), 15:161-167; Kaboev, Nucleic Acids Res. (2000), 28:E94. Aptamer beacons can adopt two or more conformations, one of which allows for ligand binding. They use a fluorescence quenching pair to report the conformational change induced by ligand binding. See, e.g., Yamamoto et al., Genes Cells (2000), 5:389-396; Smimov et al., Biochemistry (2000), 39:1462-1468.
[0188] A nucleic acid primer or probe can be indirectly detectably labeled via a peptide. A peptide can be made detectable by incorporating a predetermined polypeptide epitope that is recognized by a secondary reporter (e.g., a leucine zipper pair sequence, a binding site for a secondary antibody, a transcription activator polypeptide, a metal binding domain, an epitope tag). A label can also be attached via a second peptide that interacts with the first peptide (e.g., S-association).
[0189] As can be easily recognized by those skilled in the art, the detection of the complex containing nucleic acid from the sample hybridized with the labeled probe can be achieved by using a labeled antibody against the label of the probe.In one example, the probe is labeled with digoxigenin and detected with a fluorescently labeled anti-digoxigenin antibody.In another example, the probe is labeled with FITC and detected with a fluorescently labeled anti-FITC antibody.These antibodies are readily commercially available.In another example, the probe is labeled with FITC and detected with a primary anti-FITC antibody and a secondary labeled anti-FITC antibody.
[0190] The nucleic acid can be amplified prior to detection or can be detected directly during the amplification step (i.e., "real-time" methods such as TaqMan® and Scorpion™). In some embodiments, the target sequence is amplified using labeled primers such that the resulting amplicons are detectably labeled. In some embodiments, the primers are fluorescently labeled. In some embodiments, the target sequence is amplified and the resulting amplicons are detected by electrophoresis.
[0191] With regard to exemplary primers and probes, those skilled in the art will readily recognize that nucleic acid molecules can be double-stranded molecules, and that reference to a particular site on one strand also refers to the corresponding site on the complementary strand. When defining a variant position, allele, or nucleotide sequence, reference to adenine, thymine (uridine), cytosine, or guanine at a particular site on one strand of a nucleic acid molecule also defines thymine (uridine), adenine, guanine, or cytosine (respectively) at the corresponding site on the complementary strand of the nucleic acid molecule. Thus, either strand can be referenced to refer to a particular variant position, allele, or nucleotide sequence. Probes and primers can be designed to hybridize to either strand, and the detection methods disclosed herein can generally target either strand.
[0192] In some embodiments, primers and probes contain additional nucleotides corresponding to the sequence of a universal primer (e.g., T7, M13, SP6, T3) that adds additional sequence to the amplicon during amplification to allow further amplification and / or to prime the amplicon for sequencing.
[0193] As stated above, the present disclosure further provides a method of treating a patient selected by the method of any of the above embodiments or identified by any of the above methods as likely to experience a more favorable clinical outcome following the treatment. In some embodiments, the method involves administering such a treatment to the patient. The treatment can be any one of the following group: a first line treatment, a second line treatment, a third line treatment, a fourth line treatment, or a fifth line treatment.
[0194] Compositions and Modes of Administration
[0195] A pharmaceutical agent or drug may be administered as a composition. A "composition" typically contemplates a combination of an active agent with another carrier, such as a compound or composition, inert (e.g., detectable agent or label) or active, such as an adjuvant, diluent, binder, stabilizer, buffer, salt, lipophilic solvent, preservative, adjuvant, etc., and includes a pharma- ceutically acceptable carrier. Carriers also include pharmaceutical excipients and additives proteins, peptides, amino acids, lipids, and carbohydrates.
[0196] A variety of delivery systems are known and can be used to administer the chemotherapeutic agents of the present disclosure, such as encapsulation in liposomes, microparticles, microcapsules, recombinant cell expression, and receptor-mediated endocytosis.For example, see Wu and Wu (1987) J.Biol.Chem.262:4429-4432 for the construction of therapeutic nucleic acids as part of retroviral or other vectors.Delivery methods include, but are not limited to, intraarterial, intramuscular, intravenous, intranasal, and oral routes.In certain embodiments, it may be desirable to administer the pharmaceutical compositions of the present disclosure locally to the area requiring treatment, which can be achieved, for example, but not limited to, by local infusion during surgery, injection, or by using a catheter.
[0197] Agents identified herein as effective for their intended purpose can be administered to subjects or individuals identified by the methods herein as suitable for treatment. Therapeutic amounts can be empirically determined and will vary depending on the condition being treated, the subject being treated, and the efficacy and toxicity of the agent.
[0198] Methods of administering pharmaceutical compositions are well known to those skilled in the art and include, but are not limited to, oral, microinjection, intravenous or parenteral administration. The compositions are intended for local, oral or topical administration, as well as intravenous, subcutaneous or intramuscular administration. Administration can be continuous or intermittent throughout the course of treatment. Methods of determining the most effective means and dosages of administration are well known to those skilled in the art and vary depending on the cancer being treated and the patient and subject being treated. Single or multiple administrations can be administered at dose levels and patterns selected by the treating physician.
[0199] kit
[0200] A kit or panel is provided for use in detecting a polymorphism of interest in a biological sample of a patient. In some embodiments, the kit comprises, consists essentially of, or even consists of at least one reagent required to perform an assay. For example, the kit can include enzymes, buffers, or any other required reagents (e.g., PCR reagents and buffers). For example, in some aspects, the kit includes either a hybridization assay probe, amplification primers, and / or an antibody suitable for detection in packaging materials in an amount sufficient for at least one assay.
[0201] The various components of the kit can be provided in various forms. For example, in some embodiments, the required enzymes, nucleotide triphosphates, probes, primers, and / or antibodies are provided as lyophilized reagents. These lyophilized reagents can be premixed before lyophilization so that when reconstituted, they form a complete mixture with the appropriate ratio of each of the components ready for use in the assay. The kit can also include a reconstitution reagent for reconstituting the lyophilized reagents of the kit. In an exemplary kit for amplifying target nucleic acids from colorectal cancer patients, the enzymes, nucleotide triphosphates, and cofactors required for the enzymes are provided as a single lyophilized reagent that when reconstituted, forms a suitable reagent for use in the present amplification method.
[0202] Typically, the kit also includes instructions recorded in tangible form (e.g., contained on paper or electronic media) for using the packaged probes, primers and / or antibodies in a detection assay to determine the presence or amount of a polymorphism of interest in a test sample.
[0203] In some embodiments, the kit further comprises a solid support for immobilizing the nucleic acid of interest on the solid support. The target nucleic acid can be immobilized directly on the solid support or indirectly on the solid support via a capture probe that is immobilized on the solid support and can hybridize to the nucleic acid of interest. Examples of such solid supports include, but are not limited to, beads, microparticles (e.g., gold and other nanoparticles), microarrays, microwells, and multi-well plates. The solid surface can include a first member of a binding pair, and the capture probe or the target nucleic acid can include a second member of the binding pair. Binding of the binding pair members immobilizes the capture probe or the target nucleic acid on the solid surface. Examples of such binding pairs include, but are not limited to, biotin / streptavidin, hormone / receptor, ligand / receptor, and antigen / antibody.
[0204] In one aspect, the kit further comprises, consists essentially of, or even further consists of an effective amount of a therapeutic.
[0205] The kit can include at least one probe or primer that can specifically hybridize to the gene of interest and instructions for use.For example, in some aspects, the kit includes at least one of the above-mentioned nucleic acids.An exemplary kit for amplifying at least a part of the gene of interest includes two primers.For example, in some embodiments, the kit includes, consists essentially of, or even consists of forward primer and reverse primer that flank the polymorphism.
[0206] In some embodiments, the kit further comprises, consists essentially of, or even further consists of a nucleic acid probe for detection of the amplicon. In some embodiments, the nucleic acid probe has about 5, about 10, about 15, about 20, or about 25, or about 30, about 35, or about 40 or more consecutive nucleotides. In some aspects, the nucleic acid primers and / or probes are lyophilized.
[0207] The oligonucleotides included in the kit, whether used as probes or primers, can be detectably labeled. The label can be directly or indirectly detected, for example in the case of fluorescent labels. Indirect detection can include any detection method known to those skilled in the art, including biotin-avidin interactions, antibody binding, etc. Fluorescently labeled oligonucleotides can also contain quenching molecules. The oligonucleotides can be attached to a surface. In one embodiment, the surface is silica or glass. In another embodiment, the surface is a metal electrode.
[0208] The test sample used in the diagnostic kit includes cells, protein or membrane extracts of cells, or biological fluids such as sputum, blood, serum, plasma or urine.The test sample can also be tumor cells, normal cells adjacent to tumors, normal cells corresponding to tumor tissue types, blood cells, peripheral blood lymphocytes, or combinations thereof.The test sample used in the above method varies based on the assay format, the nature of the detection method, and the tissue, cell or extract used as the sample to be assayed.Methods for preparing protein or membrane extracts of cells are known in the art and can be easily adapted to obtain a sample that is compatible with the system to be used.
[0209] The kits may include all or some of the positive controls, negative controls, reagents, primers, sequencing markers, probes and antibodies described herein for determining a subject's genotype at a polymorphic region of a gene or target region of interest.
[0210] As appropriate, the proposed kit components can be packaged in a conventional manner for use by one of skill in the art For example, the proposed kit components can be provided in a solution or as a liquid dispersion, etc.
[0211] Typical packaging materials include solid matrices such as glass, plastic, paper, foil, microparticles, etc. that can hold hybridization assay probes and / or amplification primers within a certain range. Thus, for example, packaging materials can include glass vials used to contain sub-milligram (e.g., picogram or nanogram) quantities of contemplated probes, primers or antibodies, or they can be microtiter plate wells to which probes, primers or antibodies are operatively attached, i.e., bound so that they can participate in the amplification and / or detection method.
[0212] The instructions will typically indicate at least one assay method parameter, which may be, for example, the reagents and / or concentrations of the reagents, and the relative amounts of reagents to use per amount of sample. Further details such as maintenance, duration, temperature, and buffer conditions may also be included.
[0213] The diagnostic systems contemplate a kit having any of the hybridization assay probes, amplification primers or antibodies described herein, or as identified herein, for use in determining the presence or amount of a polymorphism of interest, whether provided individually or in one of the above combinations.
[0214] The present disclosure, having now generally been described, will be more readily understood by reference to the following examples, which are included solely for the purpose of illustrating certain aspects and embodiments of the disclosure and are not intended to be limiting of the disclosure.
[0215] Experimental Method
[0216] Experiment Number 1
[0217] Clustered mutation landscape
[0218] To identify clustered mutations, a sample-dependent intramutation distance (IMD) cutoff was obtained, where mutations below the cutoff were unlikely to occur by chance (q-value <0.01). A statistical approach utilizing IMD cutoff, variant allele frequency (VAF), and correction for local sequence context was applied to each specimen (see Figure 6A). Clustered mutations with consistent VAF were subdivided into four categories (see Figure 6B). Double base substitutions (DBS) and multi-base substitutions (MBS) were characterized as 2 and ≥3 adjacent mutations (IMD=1), respectively. Multiple substitutions with IMD >1bp and below the sample-dependent cutoff, respectively, were characterized as either omikli (2–3 substitutions) or kataegis (≥4 substitutions). Clustered substitutions with inconsistent VAF were classified as other. The clustered indels were not subdivided into distinct categories, but most events resembled diffuse hypermutations, with 92.3% of events harboring only two indels (see Figure 6C).
[0219] Examination of 2,583 whole-genome sequenced cancers from the Pan-Cancer Analysis of Whole Genomes (PCAWG) project revealed a total of 1,686,013 clustered single-base substitutions and 21,368 clustered indels (Figure 1 and Figure 6D). DBS, MBS, omikli, and kataegis constituted 45.7%, 0.7%, 37.2%, and 7.0% of the clustered substitutions across all samples, respectively, and their distribution varied significantly within and between cancer types. For example, melanoma had the highest amount of clustered substitutions, with ultraviolet light-associated doublets (i.e., CC>TT) accounting for 74.2% of the clustered mutations, but these contributed only 5.3% of all substitutions in melanoma (Figure 1A). In contrast, 11.5% of all substitutions in bone leiomyosarcoma were clustered, with omikli and kataegis constituting 43.8% and 46.7% of these mutations, respectively (Figure 1A). Clustered indels showed diverse patterns within and across cancer types alike (Figure 1B). For example, the highest mutational burden of clustered indels was observed in lung and ovarian cancers. Clustered indels in lung cancer accounted for only 2.6% of all indels and were characterized by 1-bp deletions. In contrast, long indels clustered with microhomology were commonly found in ovarian and breast cancers, contributing >10% of all indels in a subset of samples (Figure 1B). A correlation between the total number of mutations and the number of clustered mutations was observed in DBS and omikli, but not in MBS, kataegis or indels (Figure 6E). In most cancers, DBS and omikli had VAFs consistent with those of non-clustered mutations, whereas MBS and kataegis tended to have lower VAFs (Figure 6F). Kataegic events contained 4–44 mutations, and 81% of events were strand-cooperative, indicating damage or enzymatic alterations on a single DNA strand.
[0220] Overall survival was compared between patients whose cancers harbored major and minor clustered mutations within whole-genome sequenced PCAWG and whole-exome sequenced TCGA cancer types 36Better overall survival was observed only in whole-genome sequenced ovarian cancers containing high levels of clustered substitutions or clustered indels (q-value < 0.05; Figure 6G and Figure 6H). Conversely, whole-exome sequenced adrenocortical carcinomas containing clustered substitutions were associated with worse overall survival (q-value = 7.2E-05; Figure 6I-K).
[0221] Clustered mutation signatures
[0222] Mutational signature analysis was performed for each category of clustered events elucidating 12 DBS, 5 MBS, 17 omikli, 9 kataegic and 6 clustered indel signatures (Figure 2). The DBS signatures have been described previously. 1 In previous analyses, DBS and MBS were combined into a single class. 1 Separating these events into individual classes revealed that while multiple processes can give rise to DBS, the majority of MBS are attributable to signatures associated with cigarette smoking (SBS4) or ultraviolet light (SBS7). Additional DBS and MBS signatures were found within a small subset of cancer types (Figure 12).
[0223] In cancer genomes, omikli has previously been attributed to APOBEC3 mutagenesis. 6 , and there is some indirect evidence from experimental models. 23、37、38 Applicant's analysis of sequencing data from clonally expanded breast cancer cell line BT-474 with active APOBEC3 mutagenesis. 39experimentally confirmed the presence of APOBEC3-associated omikli events (cosine similarity: 0.99; Figure 13A). Only 16.2% of omikli events across 2,583 cancer genomes matched APOBEC3 mutation patterns, indicating that a plethora of other processes may generate widespread clustered hypermutations. Importantly, our analysis revealed omikli due to cigarette smoking (SBS4), clock-like mutation processes (SBS5), ultraviolet light (SBS7), both direct and indirect mutations from AID (SBS9 and SBS85), and multiple mutation signatures with unknown etiology in different cancer types (SBS8, SBS12, SBS17a / b, SBS28, SBS40, and SBS41; Figure 2). Benzo[α]pyrene 40 and ultraviolet light 41 Cell lines previously exposed to confirmed the occurrence of omikli events resulting from these two environmental exposures (cosine similarity: 0.86 and 0.84, respectively; Figure 13A).
[0224] From the nine kataegic signatures, four have been previously reported, including two associated with APOBEC3 deaminases (SBS2 and SBS13) and two associated with canonical or non-canonical AID activity (SBS84 and SBS85; Figure 2). SBS5 (clock-like mutagenesis) accounted for 15.0% of the kataegis, with most events occurring in the vicinity of AID kataegis within B-cell lymphomas. The remaining four kataegic signatures accounted for only 8.9% of the kataegic mutations and included SBS7a / b (ultraviolet light), SBS9 (indirect mutations derived from AID) and SBS37 (unknown etiology). Most kataegic signatures were strand-cooperative (Figure 13B). Some samples showed consistency, while others showed distinct signatures of clustered and non-clustered mutagenesis (Figure 14). For example, in SP56533 (lung squamous cell carcinoma), most non-clustered and omikli substitutions were caused by the tobacco signature SBS4, whereas kataegic events were generated by the APOBEC3 signature (Figure 14A). In contrast, the pattern of non-clustered substitutions in SP24815 (glioblastoma) was due to the clock-like signatures SBS1 and SBS5, whereas omikli and kataegic events were mainly due to APOBEC3 (Figure 14A).
[0225] The remaining other clustered substitutions showed inconsistent VAFs that may represent the effects of co-occurring large mutational events such as mutations or copy number changes in highly mutable genomic regions (Figure 13D).
[0226] Different cancers revealed distinct trends of clustered indel mutagenesis (Figure 2). For example, clustered indels caused by ID3 (cigarette smoking; characterized by a 1 bp deletion) were found mainly in lung cancer and were significantly elevated in smokers compared to non-smokers (p-value: 0.0014; Figures 13C and 14B). Clustered indels by signatures ID6 and ID8, both of which are due to homologous recombination defects and are characterized by long indels of microhomology, were found in breast and ovarian cancers and were highly elevated in cancers with known defects in homologous recombination genes (p-value: 4.9 × 10 -11 ;Figures 13C and 14B).
[0227] A panorama of clustered driver mutations
[0228] The PCAWG project uncovered a constellation of mutations that putatively drive cancer development 10The disclosed data reveal a significant enrichment of clustered substitutions and clustered indels among these driver mutations. Specifically, they contribute to 8.4% and 6.9% of substitution and indel drivers, respectively, whereas only 3.7% of all substitutions and 0.9% of all indels are clustered events (q-value < 1e-5; Fisher's exact test; Figure 3A and Figure 3B). Omikli accounts for 50.5% of all clustered substitution drivers, while DBS, kataegis, and other clustered events accounted for 14%-18%, respectively (Figure 3C). Clustered driver substitutions vary widely between genes and across different cancers (Figure 3C and Figure 7A), with a 2.4-fold enrichment of clustered events within cancer genes compared to tumor suppressors (p-value = 5.79E-03; Figure 7B and Figure 7C). In some cancer genes, a small percentage of driver events were due to clustered substitutions; examples include TP53 (4.5% clustered driver substitutions), KRAS (3.7%), and PIK3CA (2.2%). In other genes, most detected substitution drivers were clustered events; examples include BTG1 (73.1%), SGK1 (66.6%), EBF1 (60.0%), and NOTCH2 (38.5%). Importantly, the contribution from each class of clustered events varied across driver substitutions in different genes (Figure 3C). For example, UV-related DBS constituted 93% of clustered BRAF driver events, omikli accounted for 63% of clustered BTG1 driver events, and kataegis accounted for 100% of clustered NOTCH2 driver substitutions (Figure 3C). Similar behavior was observed for clustered indel drivers, with 48.7% being single base pair indels (Figure 3D). In some cancer genes, clustered indel drivers were rare (e.g., 2.4% of indel drivers in TP53 were clustered), whereas in other cancer genes they were common (e.g., 76.6% in ALB; Figure 3D ).When compared to non-clustered drivers, clustered driver substitutions were enriched for stop-lost mutations (q value = 1.9E-02) and depleted for stop-gain mutations (q value = 3.3E-03) (Figure 3E). Furthermore, driver genes with clustered events were often differentially expressed compared to driver genes with non-clustered events (Figure 7D). For example, clustered events within CTNNB1 and BTG1 associated with increased expression (q value < 0.05) compared to both non-clustered and wild-type expression levels of each gene. The opposite effect was observed for STAT6 and RFTN1 (q value < 0.05). Taken together, these driver events were induced by the activity of multiple mutational processes, including exposure to ultraviolet light, cigarette smoke, platinum chemotherapy, and AID / APOBEC3 activity, among others (Figure 5E).
[0229] Kataegic events and local amplification
[0230] In each sample, kataegic mutations were separated into separate events based on consistent VAF across adjacent mutations and IMD distances greater than a sample-dependent IMD threshold. Applicant's analysis revealed that 36.2% of all kataegic events occurred within 10 kb of a structural breakpoint, but not in a detected local amplification (Figure 4A). Furthermore, 21.8% of all kataegic events occurred within 10 kb of either a detected local amplification or a structural breakpoint of a local amplification: 9.6% in circular extrachromosomal DNA (ecDNA), 6.3% in linear rearrangements, 3.3% within severely rearranged events, and 2.6% associated with BFBs (Figure 4A). Finally, 42.0% of kataegic events were neither within 10 kb of a structural breakpoint nor in a detected local amplification. Modeling the distribution of distances between kataegic events and the nearest structural variation revealed a multimodal distribution with three components (Figure 4B) of kataegis within 10 kb, ~10 Mb, or >1.5 Mb of the detected breakpoint. Importantly, ecDNA-associated kataegis, called kyklonas (Greek for cyclone), had an average distance of ~750 kb from the nearest breakpoint, and only 0.35% of kyklonic events occurred both on the ecDNA and within 10 kb of the breakpoint (Figure 4B). These results indicate that kyklonic events are unlikely to have arisen due to structural rearrangements during the formation of ecDNA. In most cancer types, DBS, MBS, omikli, and other cluster events were not found in the vicinity of structural variations (Figures 15A and 15B).
[0231] Repetitive kyklonic mutagenesis of ecDNA
[0232] Although only 9.6% of kataegic events occurred within ecDNA regions, more than 30% of ecDNAs had one or more associated kyklonic events (Figure 4C). Mutations within these ecDNA regions were dominated by the APOBEC3 pattern, characterized by strand-coordinated C>G and C>T mutations in a TpCpW context and attributed to signatures SBS2 and SBS13 (p-value <1E-5; Figure 4C, Figure 4D, Figure S15C). These APOBEC3-associated events accounted for 97.8% of all kyklonic events, while the remaining mutations were attributed to the clock-like signature SBS5 (1.2%) and other signatures (1.0%; Figure S15C). Furthermore, kyklonic events were characterized by APOBEC3A-preferred YTCA contexts, which were ... S15C). 7 We showed enrichment of C>T and C>G mutations in APOBEC3B-preferred RTCAs compared to non-ecDNA kataegis, indicating that APOBEC3B may play an important role in mutagenesis of circular DNA bodies (Figure 4E). Similar levels of enrichment for RTCA contexts were also observed in both non-ecDNA kataegis and non-SV-associated kataegis, suggesting that APOBEC3B generally drives many of the strand-cooperative kataegic events (Figure S15D). Elevated expression of APOBEC3B, but not APOBEC3A, was observed in cancers with ecDNA compared to samples without ecDNA (3.1-fold; q-value <1E-5; Figure 4F). Within cancers with ecDNA, no differences in APOBEC3A / B expression were observed between samples with and without kyklonic events (Figure 4F).
[0233] More recurrent APOBEC3 kataegis were observed across circular ecDNA regions compared to other forms of structural variation (Figure 5A). An average of 2.5 kyklonic events were observed within ecDNA regions (range: 0–64 kyklonic events; 0–505 mutations). Recurrent kyklonas were widespread across cancer types (Figure 8A and Figure 8B). For example, glioblastomas and sarcomas exhibited an average of 5 and 86 kyklonic mutations, respectively. The average VAF of kyklonas was significantly lower than both non-ecDNA-associated kataegis and all other clustered events (q-value < 1e-5; Figure 5B). Interestingly, a subset of kyklonas exhibited a VAF above 0.80, which may reflect early mutagenesis of genomic regions that were later amplified as ecDNA. Furthermore, kyklonic events with high VAF occurred more commonly in ecDNA harboring known cancer genes, suggesting a mechanism of positive selection (Figure 5B). Approximately 7.2% of kyklonas arose early in the evolution of a given ecDNA population within a tumor (VAF>0.80), whereas the majority of kyklonic events (approximately 82.5%; VAF<0.5) likely arose after clonal amplification due to recurrent APOBEC3 mutagenesis.
[0234] Recurrent kyklonic events were increased within or near known cancer-associated genes, including TP53, CDK4, and MDM2, among others (Figure 5C). These recurrent kyklonas were observed across many cancers, including glioblastoma, sarcoma, head and neck cancer, and lung adenocarcinoma (Figure 8C and Figure 8D). For example, in a sarcoma sample (SP121828), 10 distinct kyklonic events coincided with a single ecDNA region with recurrent APOBEC3 activity in close proximity to MDM2, resulting in a missense L230F mutation (Figure 8C). The same ecDNA region harbored additional kyklonic events occurring within intergenic regions with distinct VAF distributions indicative of recurrent mutagenesis (Figure 8C). Similarly, two distinct kyklonic events occurred on ecDNA with EGFR, resulting in a missense mutation D191N within head and neck cancer (Figure 8D). Importantly, ecDNA regions with known cancer-associated genes had significantly higher numbers of kyklonic events and mutational burden of kyklonas compared to ecDNA regions without known cancer-associated genes (q-value < 1e-5; Figure 5D). Furthermore, applicants observed higher co-occurrence of kyklonas with known cancer-associated genes, which were 2.5 times more mutated than ecDNA without cancer-associated genes (p-value = 1.2e-5; Fisher's exact test). Overall, 41% of kyklonic events were found within the footprints of known cancer driver genes (p-value < 1e-5). These enrichments cannot be explained by either an increase in overall mutations or an increase in overall clustered mutations in these samples (Figure 5E). To understand the functional effects of kyklonas, applicants annotated the predicted outcomes of each mutation. In total, 2,247 kyklonic mutations matched putative cancer-associated genes, of which 4.3% resided within coding regions (Figure 8E). Specifically, 63 resulted in missense mutations, 29 resulted in synonymous mutations, 4 introduced premature stop codons, and 1 removed a stop codon (data not shown). These downstream consequences of APOBEC3 mutagenesis suggest a contribution of specific ecDNA populations to the oncogenic evolution.
[0235] Verification of kyklonic events in ecDNA
[0236] Kyklonic events were assessed in 61 sarcomas. 44 , 280 lung cancers 45 , and 186 esophageal squamous cell carcinomas. 46 We further investigated across three additional independent cohorts, including 1.1% and 1.2% of the kyklonos in the validation cohort. Similar to the rates reported in PCAWG, comparable rates of clustered mutagenesis were found for both substitutions and indels, with 2.4- and 5.0-fold enrichment of clustered substitutions and indels within driver events, respectively (Figure 16A). Across the three cohorts, 31% of samples with ecDNA exhibited kyklonas within sarcomas, 14% within esophageal cancers, and 28% within lung cancers, confirming the rates observed in PCAWG (Figure 4C and Figure 16C). Similar to the rate observed in PCAWG (36.2%), approximately 30.1% of all kataegis occurred within 10 kb of the nearest breakpoint in the validation cohort (Figure 17A). Furthermore, only 0.34% of kyklonic events in the validation dataset occurred closer to a structural variant than expected by chance, which closely resembled the observation in the PCAWG data (0.35%; Figure 17B). Kyklonic mutations were predominantly attributable to APOBEC3 signatures SBS2 and SBS13 (p-value < 1E-05; Figure 16B), with enrichment of mutations in the RTCA context supporting the role of APOBEC3B (Figure 16D). In sarcoma, esophageal cancer, and lung cancer, extensive repetition of kyklonic events was observed in 45%, 28%, and 46% of samples with ecDNAs with multiple distinct kyklonic events (Figure 16E). Examples from each cohort were selected to illustrate multiple kyklonic events occurring within a single ecDNA validating recurrent APOBEC3 hypermutation in ecDNA (Figure 17).
[0237] Data Source
[0238] Somatic variant calls for single-base substitutions, small insertions and deletions, and structural variations along with a list of corresponding consensus driver events were downloaded for 2,583 whitelist whole-genome sequencing samples from the PCAWG 10 The epidemiological and clinical characteristics of all available samples are summarized in the official PCAWG release ( https: / / dcc.icgc.org / releases / PCAWG A collection of whole-exome sequencing samples from TCGA along with all available clinical features was downloaded from Genomic Data Commons ( https: / / gdc.cancer.gov / The data was downloaded from the MSK-IMPACT Clinical Sequencing Cohort, which consists of 10,000 clinical cases. 43 cBioPortal ( https: / / www.cbioportal.org / study / summary?id=msk_impact_2017 ) were downloaded. Subclassification of local amplifications consisting of circular extrachromosomal DNA (ecDNA), linear amplifications, break-fusion-bridge cycles (BFB), and highly rearranged events, and their corresponding genomic locations, were obtained for a subset of samples (n = 1,291) as reported. 34 .
[0239] The experimental model used to validate the clustering events is primary Hupki mouse embryonic fibroblasts (MEFs) exposed to ultraviolet light. 41 , Human induced pluripotent stem cells (iPSCs) exposed to benzo[a]pyrene 40 , and clonally expanded BT-474 human breast cancer cell line harboring epidemiologically active APOBEC3 39 were obtained from a previous study using
[0240] The independent cohort used to validate the kyklonic events was collected from multiple sources: 61 undifferentiated sarcomas. 44 and 187 high-confidence esophageal squamous cell carcinomas. 46 were downloaded from the European Genome-phenome Archive (EGAD 00001004162 and EGAD 00001006868, respectively). 280 lung adenocarcinomas 45was downloaded from dbGaP under the accession number (phs001697.v1.p1). Clustered mutations in the validation samples were analyzed using the same approach as that utilized in the original cohort.
[0241] Detecting clustering events
[0242] SigProfilerSimulator (v1.0.2) was used to determine intramutation distance (IMD) cutoffs that are unlikely to occur by chance based on the tumor mutation burden and mutation pattern for a given sample. 51 Specifically, we simulated each tumor sample while maintaining the sample's mutational burden on each chromosome, the + / - 2bp sequence context of each mutation, and the transcription strand bias ratio across all mutations. We simulated all mutations in each sample 100 times and calculated an IMD cutoff such that 90% of mutations below this cutoff could not have appeared by chance (q-value < 0.01). For example, for a sample with an IMD threshold of 500bp, we calculated the IMD cutoff within this threshold to have 100 or fewer mutations that would be expected based on the simulated data. 1,0 00 mutations can be observed (q-value <0.01). P-values were calculated using z-tests by comparing the number of actual mutations to the distribution of simulated mutations occurring below the same IMD threshold. A maximum cutoff of 10 kb was used for all IMD thresholds. By generating a background distribution that reflects the random distribution of events used to reduce the false positive rate, the model also accounts for regional heterogeneity in mutation rates due in part to replication timing and expression, as well as clonality variation by correcting for mutation-rich and mutation-poor regions within a 1 Mb window. The 1 Mb window size has been utilized and established as an appropriate measure when considering mutation rate variability associated with chromatin structure, replication timing, and genome architecture. 14、52、53A 1 Mb window ensures that subsequent mutations are likely to occur as single events, using a maximum cutoff of 0.10 for the difference in variant allele frequency (VAF). The regional IMD cutoff was determined using a sliding window approach that calculates the fold enrichment between the actual and simulated mutation densities in a 1 Mb window across the genome. To capture additional clustering events while maintaining the original criteria, the IMD cutoff was further increased for regions where the enrichment of clustered mutations was higher than 9-fold and >90% of the clustered mutations were found in the original data (less than 10% of mutations below this cutoff would occur by chance; q-value <0.01). Finally, because the VAF of mutations can confound the definition of clustering events in ecDNA, applicants calculated the distribution of inter-event distances in recurrently mutated ecDNA, ignoring the VAF of individual mutations. This resulted in the exact same separation of kataegic events using only inter-event distances as a criterion for grouping mutations into single events.
[0243] All clustered mutations with consistent VAF were then classified into one of four categories (Figure 6A). Two adjacent mutations with an IMD of 1 were classified as double-stranded base substitutions. Three or more adjacent mutations, each with an IMD of 1, were classified as multi-base substitutions. Two or three mutations with an IMD below the sample-dependent threshold and at least a single IMD above 1 were classified as omikli. Four or more mutations with an IMD below the sample-dependent threshold and at least a single IMD above 1 were classified as kataegis. The cutoff of four mutations for kataegis was chosen by fitting a Poisson mixture model to the number of mutations involved in a single event across all extended clustering events except DBS and MBS (data not shown). This model constructed two distributions with C1=2.08 and C2=4.37 representing omikli and kataegis, respectively. A cutoff of four mutations was used for kataegis based on a >95% contribution from the kataegis-related distribution with events of four or more mutations. Applicants note that there is a certain ambiguity for events with two or three mutations. Most of these events are omikli, but some of these events are likely short kataegic events (data not shown). All remaining clustered mutations with inconsistent VAF were classified as Other. Clustered indels were not classified into distinct classes. Applicants also performed additional quality checks to ensure that the majority of clustered indels map to high-confidence regions of the genome (data not shown). Specifically, all clustered indels were analyzed using ENCODE 54 Alignment against a consensus list of blacklisted genomic regions developed by revealed that only 0.5% of all clustered indels matched regions with low mappability scores.
[0244] Clustered mutational signature analysis
[0245] Clustered mutation catalogs of the examined samples were generated using the SigProfilerMatrixGenerator for each tissue type and each category of clustering events. 55 (version 1.2.0) were used to compile the SBS288 and ID83 matrices. For example, six matrices were constructed for the clustered mutations found in Breast-AdenoCA: one matrix for DBS, one matrix for MBS, one matrix for omikli, one matrix for kataegis, one matrix for other clustered substitutions, and one matrix for clustered indels. The SBS288 classification considers the 5' and 3' bases (referenced using pyrimidine bases in Watson-Crick base pairs) directly adjacent to each single base substitution, resulting in 96 individual mutation channels. In addition, this classification considers the strand orientation for mutations occurring within gene regions resulting in three possible categories; (i) transcribed; the pyrimidine base is present on the template strand; (ii) untranscribed; the pyrimidine base is present on the coding strand; or (iii) untranscribed; the pyrimidine base is present in an intergenic region. Note that mutations in bidirectionally transcribed gene regions were evenly divided between the coding and template strand channels. Combined, this results in a classification consisting of 288 mutation channels that were used as input for de novo signature extraction of clustered substitutions. The ID83 mutation classification has been previously described. 55 .
[0246] Mutational signatures were extracted from the generated matrices using SigProfilerExtractor (v 1.1.0), a Python-based tool that uses nonnegative matrix factorization to decipher both the number of operational processes within a given cohort and the relative activity of each process within each sample. 56The algorithm was initialized by using random initialization and applying a multiplicative update using Kullback-Leibler divergence for 500 iterations. Each de novo extracted mutation signature was then matched against a signature ( https: / / cancer.sanger.ac.uk / signatures / All de novo extractions and subsequent digests were visually inspected and analyzed as previously performed. 1 Manual correction was performed on 2.2% of extractions (4 out of 180 extractions) to adjust the total number of operational signatures to ±1. 10 In this study, applicants included all cancer types within the PCAWG cohort, which may contain as few as one sample for a particular cancer type. Similarly, previous visualizations 1 Consistent with, the decomposed signature activity plot required each cancer type to have more than two samples, and mutation thresholds were used for each clustering category; double base substitutions, omikli events, and other clustered mutations required 25 mutations per sample; multi-base substitutions and kataegic events required 15 mutations per sample; and for clustered indels, 10 mutations per sample were required.
[0247] Experimental Verification
[0248] A subset of the clustered mutational signatures was validated using previously sequenced in vitro cell line models. As was done for the PCAWG samples, Applicants used SigProfilerSimulator 51 A background model was generated using SigProfilerMatrixGenerator to calculate the clustering IMD cutoff for each sample and split each substitution into the appropriate category of clustering events. 55Mutational spectra were generated for each subclass within each sample using and compared against de novo signatures extracted from human cancers. The cosine similarity between the in vitro mutational spectra and the de novo observed clustering signatures was calculated to assess the degree of similarity. Applicants note that the average cosine similarity between two random non-negative vectors is 0.75, and a cosine similarity greater than 0.81 reflects a p-value less than 0.01. 51 .
[0249] Association with cancer risk factors
[0250] Homologous recombination deficiency (HRD) was defined for breast cancer using BRCA1, BRCA2, RAD51C and PALB2 status 57 Samples with germline, somatic or epigenetic alterations in one of these genes were considered HR deficient, whereas samples without known alterations in these genes were considered HR competent. The number of clustered indels was compared between HR deficient and HR competent samples. Lung cancer smoking status was determined using clinical annotations from TCGA (https: / / portal.gdc.cancer.gov / repository). The number of clustered indels associated with smoking (ID6) was compared between samples annotated as lifetime non-smokers and those annotated as smokers and ex-smokers. Alcohol intake status was determined using annotations from the official PCAWG release (https: / / dcc.icgc.org / releases / PCAWG). The total number of clustered indels was compared between samples annotated as no alcohol intake and those annotated as daily and weekly drinkers.
[0251] Driver gene expression
[0252] All RNA-seq expression data were downloaded as part of the official PCAWG release (https: / / dcc.icgc.org / releases / PCAWG). Relative expression data found within this release was normalized using fragments per kilobase of exon per million mapped fragments (FPKM) normalization and upper quartile normalization. Relative expression of genes was compared between those with clustered or non-clustered events. Each distribution was then normalized to the average expression of wild-type genes. Only genes with at least 10 total events (i.e., clustered and non-clustered mutations) containing at least 5 clustered events were considered for testing.
[0253] Structural variants and clustering events
[0254] The distance to the nearest structural variation breakpoint was calculated for each mutation in each subclass using the minimum distance to the nearest adjacent upstream or downstream breakpoint. Each distribution was modeled using a Gaussian mixture with an automatic selection criterion for the number of components ranging from 1 to 5 components, using the minimum Bayesian information criterion (BIC) across all iterations. Modeling of kataegic events yielded an optimal fit of three components, which was used to separate kataegic substitutions into SV-associated and non-SV-associated mutations. Both double and multi-base substitutions were modeled using a single Gaussian distribution associated with non-SV-associated mutations, whereas omikli and other clustered mutations were modeled using a mixture of two components, which may reflect the leakage of smaller kataegic events that contribute to the weak SV-associated distribution. To account for the frequency of breakpoints across each sample, applicants normalized the minimum distance of each mutation to the nearest SV by calculating the expected distance between the mutation and the SV for each sample using the total number of breakpoints and the total length of a given chromosome (data not shown). After normalizing the kataegic events, Applicants observed a two-component optimal solution with one SV-associated distribution (on average, each mutation occurs within one thousandth of the expected distance to the nearest structural variation) and one non-SV-associated distribution (on average, each mutation occurs within one thousandth of the expected distance to the nearest structural variation). The normalized kyklonic events are the non-SV-associated distribution, which reflects kataegic events that typically occur on ecDNA of lengths 1-10 Mb. 35 Matches.
[0255] APOBEC3A / B enrichment analysis
[0256] The RTCA and YTCA penta-nucleotide enrichment scores quantify the frequency with which each TpCpA>TpKpA mutation occurs in either an RTCA or YTCA context. To account for motif availability, this score is calculated using + / - 20bp of sequence context around each mutation and normalized by the number of cytosine bases and C>N mutations within the set of 41-mers surrounding each mutation of interest7.
[0257] APOBEC3 gene expression and kyklonas
[0258] All RNA-seq expression data are published in the official PCAWG release ( https: / / dcc.icgc.org / releases / PCAWG ). Relative expression data found within this release were normalized using fragments per kilobase of exon per million mapped fragments (FPKM) normalization and upper quartile normalization. APOBEC3A / B normalized expression was compared between samples with ecDNA versus samples with no detectable ecDNA, and between samples with kyklonas versus samples without kyklonas. All p-values were generated using the Mann-Whitney U test and corrected for multiple hypothesis testing using the Benjamini-Hochberg false discovery rate procedure.
[0259] Circular ecDNA and kataegis
[0260] The collection of ecDNA coverage was intersected with a catalogue of clustered mutations, which was used to determine the overlapping mutational burden for each subclass of clustered events and the mutational spectrum of overlapping kataegic events. SigProfilerSimulator shuffled dominant mutations in each clustered event across the genome. 51 Event enrichment was calculated (i.e., the most frequent mutation type in a single event) using a statistical background model generated using. Decomposed kyklonic mutation spectra were generated using the Decomposition module in SigProfilerExtractor.56 Only mutational signatures that increased the overall cosine similarity by at least 0.01 were used. In both the original and validation cohorts, SBS2 and SBS13 were sufficient to explain the kyklonic mutational spectrum, with no other known mutational signatures increasing the cosine similarity above 0.01. Comparisons between ecDNA with and without cancer genes were performed using a set of cancer genes from the Cancer Gene Census (CGC). 58 Unless otherwise stated, all statistical comparisons and p-values were calculated using two-tailed Mann-Whitney U tests. For each test set, p-values were corrected for multiple hypothesis testing using the Benjamini-Hochberg false discovery rate procedure. The predicted effect of each overlapping variant was determined using the variant effect predictor tool in ENSEMBL by reporting only the most severe outcome. 59 .
[0261] Overall survival and clustered mutations
[0262] All survival analyses, including Kaplan-Meier curve construction, Cox regression, and log-rank tests, were performed using the Lifelines Python package (v 0.24.4). Across the 30 different whole-genome sequenced cancer types included in the PCAWG study, only six cancer types contained sufficient samples to examine the association between survival and total number of clustered mutations. Sufficient sample size criteria required more than 50 samples with a survival endpoint with at least 30 samples with observed clustering events. Each cancer type was analyzed separately by comparing survival times of samples with high clustered mutation burden (top 80th percentile across a given cancer type) to survival times of samples with low clustered mutation burden (bottom 20th percentile across a given cancer type).
[0263] The analysis of whole-exome sequencing samples from TCGA was modified to reflect the limiting resolution for identifying clustered mutations within the exome. Specifically, the SigProfilerSimulator (v1.0.2) 51 Using the , we obtained an IMD cutoff for each sample based on the tumor mutation burden in the exome and the mutation pattern of a given sample. Mutations were randomly shuffled while maintaining the mutation burden in the exome for each chromosome, the + / - 2bp sequence context of each mutation, and the transcription strand bias ratio across all mutations. Each sample was simulated 100 times and the IMD cutoff was calculated using the same method outlined for the detection of clustering events in PCAWG. Due to the limited number of events detected, 22 cancer types had sufficient data to perform survival analysis. Each cancer type was analyzed separately by comparing samples with at least a single clustering event to samples with no clustering events detected in the exome.
[0264] For both PCAWG and TCGA analyses, survival distributions within a given cancer type were compared using the log-rank test. Cox regression was performed to determine hazard ratios and adjust for age and total mutation burden. All p-values were also corrected for multiple hypothesis testing using the Benjamini-Hochberg false discovery rate procedure.
[0265] To examine survival differences associated with the detection of clustered events within cancer driver genes, Kaplan-Meier survival curves were compared between individuals with clustered versus non-clustered mutations within a given cancer driver gene. Distributions were compared using the log-rank test. Cox regression was performed to determine hazard ratios, adjusting for age, total mutation burden, and cancer type across TCGA. Cox regressions performed on the MSK-IMPACT cohort were adjusted for total mutation burden and cancer type. No adjustment was made for age, as these metadata were not available for the MSK-IMPACT cohort. All p-values were also adjusted for multiple hypothesis testing using the Benjamini-Hochberg false discovery rate procedure.
[0266] Validation of kyklonas in three cohorts
[0267] All three validation cohorts were analyzed similarly to the PCAWG cohort. Specifically, we classified clustered mutations by calculating a sample-dependent IMD threshold for clustered versus non-clustered mutations using the background model generated by SigProfilerSimulator. 51 All clustered mutations were subdivided into either DBS, MBS, omikli, kataegis, or other mutations. Regions of local amplification were determined using AmpliconArchitect (version 1.2). 60 This was exploited for subsequent validation of kyklonic events by matching all detected local amplifications with kataegic events. Decomposed kyklonic mutation spectra were generated using the Decomposition module within SigProfilerExtractor. 56 Only mutational signatures that increase the overall cosine similarity by at least 0.01 were used. In both the original and validation cohorts, SBS2 and SBS13 were sufficient to explain the kyklonic mutational spectrum, with no other known mutational signatures increasing the cosine similarity above 0.01.
[0268] Data availability
[0269] No data were generated specifically for this study. All data were downloaded from the appropriate links, repositories, and references and are available for download. Specifically, for the discovery cohort, all data and metadata were obtained from the official PCAWG release: https: / / dcc.icgc.org / releases / PCAWG. All data and metadata for TCGA samples were obtained from the GDC: https: / / gdc.cancer.gov / . Genomics data for clonally expanded cell lines were downloaded from the European Genome-phenome Archive: EGAD00001004201, EGAD00001004203, and EGAD00001004583. For the three validation cohorts, the datasets were downloaded as submitted by the original publications and the genomics data were downloaded from their respective repositories: 61 undifferentiated sarcomas. 44 EGAD00001004162 (European Genome-phenome Archive), 187 highly reliable esophageal squamous cell carcinomas 46 EGAD00001006868 (European Genome-phenome Archive) for pulmonary adenocarcinomas, and 280 lung adenocarcinomas. 45 phs001697.v1.p1 (dbGaP) for the MSK-IMPACT clinical sequencing cohort of 10,000 clinical cases 42 Somatic mutations and metadata of were downloaded from cBioPortal: https: / / www.cbioportal.org / study / summary?id=msk_impact_2017.
[0270] Code Availability The SigProfiler suite of tools has been developed as Python packages and is freely available for installation via PyPI or directly via GitHub (https: / / github.com / AlexandrovLab / ). For all tools, each package is fully functional, free, and open source distributed under the permissive 2-clause BSD license, accompanied by extensive documentation: (i) SigProfilerMatrixGenerator 55 (Version 1.2.0; https: / / github.com / AlexandrovLab / SigProfilerMatrixGenerator);(ii)SigProfilerSimulator 51 :(Version 1.0.2;https: / / github.com / AlexandrovLab / SigProfilerSimulator);(iii)SigProfilerExtractor 56 :(Version 1.1.0; https: / / github.com / AlexandrovLab / SigProfilerExtractor). Each SigProfiler tool also has an R wrapper available for installation via a GitHub repository. AmpliconArchitect 34 (version 1.2) is also freely available and can be downloaded from https: / / github.com / virajbdeshpande / AmpliconArchitect. The core computational pipeline used by the PCAWG Consortium for alignment, quality control, and variant calling is publicly available at https: / / dockstore.org / search?search=pcawg under the GNU General Public License v.3.0, which allows reuse and distribution.
[0271] Consideration
[0272] Clustered mutagenesis in cancer can occur through a variety of mutational processes, with AID / APOBEC3 deaminases playing the most prominent role. In addition to enzymatic deamination, other endogenous and exogenous sources imprint many of the observed clustered indels and substitutions. Importantly, numerous mutational processes can generate omikli events, including exposure to tobacco carcinogens and ultraviolet light. Clustered substitutions and clustered indels are highly enriched driver events, are associated with differential gene expression, and are involved in cancer development and cancer evolution. Several clustered mutation signatures are associated with known cancer risk factors or activity or dysfunction of DNA repair processes. Importantly, clustered mutations in TP53, EGFR and BRAF are associated with altered overall survival and can be detected with most types of sequencing data, including clinically actionable targeted panels such as MSK-IMPACT. Clinically significant clustered mutations were also detected in KIT, KMT2C, ELF3, APC and AIID1A.
[0273] The majority of kataegic events occur within 10 kb of the detected structural variant breakpoints with mutation patterns suggestive of APOBEC3 activity. Independent of the detected breakpoints, multiple distinct kataegic events were observed on circular ecDNA structures, termed kyklonas, that have been linked to recurrent APOBEC3 mutagenesis. Circular topology of ecDNA 47 and their rapid replication pattern is reminiscent of the structure and behavior of the circular genomes of several double-stranded DNA-based pathogens, including herpesviruses, papillomaviruses, and polyomaviruses. 32~35 Importantly, previous pan-viral flora studies have shown that these double-stranded DNA viral genomes often manifest mutations from APOBEC3 enzymes. 48~50Thus, recurrent APOBEC3 mutagenesis against ecDNA likely represents an antiviral response in which ecDNA virus-like structures are treated as infectious agents and attacked by APOBEC3 enzymes. ecDNA harbors numerous cancer-related genes and is responsible for many gene amplification events that may accelerate tumor evolution. These recurrent mutagenic attacks of ecDNA reveal functional effects within known cancer genes and imply further modes of tumorigenesis that may ultimately contribute to subclonal tumor evolution, subsequent evasion of therapy, and clinical outcomes. Further investigations with large-scale clinically annotated whole genome sequenced cancers are needed to fully understand the clinical significance of clustered mutations and kyklonas.
[0274] overview
[0275] The clinical utility of detecting clustered events in driver genes was evaluated by comparing survival between individuals with clustered mutations versus non-clustered mutations in each driver gene across all whole-exome sequencing samples in TCGA. For each of these comparisons, applicants performed Cox regression accounting for effects from age and TMB while correcting for cancer type and multiple hypothesis testing. These results were validated in targeted panel sequencing data from the Memorial Sloan Kettering-Integrated Mutation Profiling of Action Cancer Targets (MSK-IMPACT) cohort. 42、43These analyses revealed significant differences in survival between individuals with clustered and non-clustered mutations detected in TP53, EGFR, and BRAF. Specifically, individuals with clustered events in BRAF had better overall survival compared to individuals with non-clustered events (q-values <0.05; Figures 7F and 7G). Conversely, in both TCGA and MSK-IMPACT, individuals with clustered mutations in TP53 or EGFR had significantly worse outcomes compared to individuals with non-clustered mutations in each of these genes (q-values <0.05; Figures 3F and 3G).
[0276] Test Number 2:
[0277] Determining the number of clustered mutations in omikli and kataegic events
[0278] To determine the cutoff for the number of mutations in omikli versus kataegic events, applicants modeled the distribution of clustering event sizes (excluding DBS, MBS, and other clustering events with discordant variant allele frequencies) using a mixture of two Poisson distributions (Figure 9A). Modeling also excluded clustering mutations from cutaneous melanomas, which contribute a disproportionate number of DBS events, and from lymphomas, which contribute a large proportion of canonical and noncanonical AID kataegis. The first component, corresponding to omikli events (gold), had an average of 2.1 mutations per event, while the second component, corresponding to the larger kataegic events (turquoise), had an average of 4.4 mutations per event. Using the posterior probability of each distribution, applicants calculated the likelihood of a given clustering event belonging to a particular component. Events consisting of 4 or more mutations were attributed to the kataegic component with a probability of more than 95%. Furthermore, Applicants evaluated the IMD distribution of events of different sizes and revealed a roughly two-fold increase in the average IMD between events with three and four mutations, supporting the activity of two separate mutational processes (Figure 9B).
[0279] Analyzing the mapping scores of clustered indels. Applicants examined the mapping scores of the entire genome for clustered indels to ensure that the majority of events fell within high-confidence regions. For this analysis, Applicants used the consensus list of blacklisted genomic regions developed by ENCODE1 and the complete set of clustered indels identified from 2,583 PCAWG samples.
[0280] Experiment Number 3
[0281] Using the MSK-MET targeted panel sequencing cohort, we observed survival differences across four genes in primary disease, including EGFR in non-small cell lung cancer, KIT in gastrointestinal stromal tumors, KMT2C in bladder cancer, and ELF3 in bladder cancer, and across two genes in metastatic disease, including APC in colorectal cancer and ARID1A in bladder cancer. Both associations were associated with worse overall outcomes when clustered mutations were present within the gene of interest. See Figure 11A and Figure 11B.
[0282] Equivalent
[0283] Thus, while the present disclosure has been specifically disclosed by preferred embodiments and optional features, modifications, improvements, and variations of the disclosure disclosed and embodied herein may be resorted to by those skilled in the art, and such modifications, improvements, and variations are considered to be within the scope of the present disclosure. The materials, methods, and examples provided herein are representative of preferred embodiments, are illustrative, and are not intended as limitations on the scope of the present disclosure.
[0284] The present disclosure has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the present disclosure. This includes the generic description of the present disclosure with a condition or negative limitation that removes any subject matter from the genus, regardless of whether the deleted material is specifically described herein.
[0285] Furthermore, when features or aspects of the disclosure are described in terms of a Markush group, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual members or subgroups of members of the Markush group.
[0286] All publications, patent applications, patents, and other references mentioned herein are expressly incorporated by reference in their entirety to the same extent as if each was individually incorporated by reference. Some references are identified by Arabic numerals and full bibliographic citations or these references are provided below. In the case of conflict, the present specification, including definitions, will control. References [ka] [ka] [ka] [ka] [ka]
Claims
1. A composition for treating or inhibiting the growth of cancer cells in a subject in need thereof, or for treating cancer in a subject in need thereof, comprising active therapy, wherein the subject has clustered mutations in one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A, or lacks clustered mutations in the BRAF gene, in a sample isolated from the subject.
2. The cancer cells or cancer are selected from carcinoma, sarcoma or blood cancer, and optionally the cancer cells or cancer are located in the circulatory system, e.g., the heart (sarcomas [angiosarcoma, fibrosarcoma, rhabdomyosarcoma, liposarcoma], myxoma, rhabdomyoma, fibroma, lipoma and teratoma), the mediastinum and pleura, and other intrathoracic organs, vascular tumors and tumor-associated vascular tissue; the respiratory tract, e.g., the nasal cavity and middle ear, paranasal sinuses, larynx, trachea, bronchi and lungs ( small cell lung cancer (SCLC), non-small cell lung cancer (NSCLC), etc.), bronchogenic carcinoma (squamous, undifferentiated small cell, undifferentiated large cell, adenocarcinoma), alveolar (bronchiolar) carcinoma, bronchial adenoma, sarcoma, lymphoma, chondromatous hamartoma, mesothelioma; gastrointestinal system, e.g., esophagus (squamous cell carcinoma, adenocarcinoma, leiomyosarcoma, lymphoma), gastrointestinal (e.g., colon, colorectum, rectum), stomach (carcinoma, lymphoma, leiomyosarcoma), stomach, pancreas (pancreatic ductal adenocarcinoma, insulinoma, glucagonoma, gastrinoma, carcinoid tumor, vipoma), small intestine (adenocarcinoma, lymphoma, carcinoid tumor, Kaposi's sarcoma, leiomyoma, hemangioma, lipoma, neurofibroma, fibroma), large intestine (adenocarcinoma, tubular adenoma, villous adenoma, hamartoma, leiomyoma); gastrointestinal stromal tumors and neuroendocrine tumors occurring anywhere; genitourinary tract, e.g., kidney (adenocarcinoma, Wilms' tumor [nephroblastoma], lymphoma, leukemia), bladder and / or or urethra (squamous cell carcinoma, transitional cell carcinoma, adenocarcinoma), prostate (adenocarcinoma, sarcoma), testis (seminoma, teratoma, embryonal carcinoma, teratocarcinoma, choriocarcinoma, sarcoma, stromal cell carcinoma, fibroma, fibroadenoma, adenomatous tumor, lipoma); liver, e.g., liver cancer (hepatocellular carcinoma), cholangiocarcinoma, hepatoblastoma, angiosarcoma, hepatocellular adenoma, hemangioma, pancreatic endocrine tumors (pheochromocytoma, insulinoma, vasoactive intestinal peptide tumor, islet cell tumor, glucagonoma, etc.); Bone, e.g., osteogenic sarcoma (osteosarcoma), fibrosarcoma, malignant fibrous histiocytoma, chondrosarcoma, Ewing's sarcoma, malignant lymphoma (reticulum cell sarcoma), multiple myeloma, malignant giant cell tumor, chordoma, osteochondroma (osteochondroma exostosis), benign chondroma, chondroblastoma, chondromyxoid fibroma, osteoid osteoma and giant cell tumor; nervous system, e.g., neoplasms of the central nervous system (CNS), primary CNS lymphoma, skull cancer (osteoma, hemangioma, granuloma, xanthomas, osteitis deformans), meninges (meningiomas, meningeal sarcomas, gliomatosis), brain cancer (astrocytoma, medulloblastoma, glioma, ependymoma, germinoma [pinealoma], glioblastoma multiforme, oligodendroma, Dendroglioma, Schwannoma, retinoblastoma, congenital tumors), spinal neurofibroma, meningioma, glioma, sarcoma); reproductive system, e.g., gynecological system, uterus (endometrial carcinoma), cervix (cervical carcinoma, preneoplastic cervical dysplasia), ovary (ovarian carcinoma [serous cystadenocarcinoma, mucinous cystadenocarcinoma, unclassified carcinoma], granulosa theca cell tumor, Sertoli-Leydig cell tumor, dysgerminoma, malignant teratoma), vulva (squamous cell carcinoma, intraepithelial carcinoma) Cancer, adenocarcinoma, fibrosarcoma, melanoma), vagina (clear cell carcinoma, squamous cell carcinoma, botryoid sarcoma (embryonal rhabdomyosarcoma), fallopian tubes (cancer) and other areas related to the female reproductive organs; placenta, penis, prostate, testes, and other areas related to the male reproductive organs; blood system, e.g., blood (myeloid leukemia [acute and chronic], acute lymphoblastic leukemia, chronic lymphocytic leukemia, myeloproliferative disorders, multiple myeloma, myelodysplastic syndromes), Hodgkin's disease oral cavity, e.g., lips, tongue, gums, floor of the mouth, palate, and other parts of the mouth, parotid glands, and other parts of the salivary glands, tonsils, oropharynx, nasopharynx, pyriform sinuses, hypopharynx, and other parts of the lips, oral cavity, and pharynx; skin, e.g., malignant melanoma, cutaneous melanoma, basal cell carcinoma, squamous cell carcinoma, Kaposi's sarcoma, lentil dysplastic nevi, lipoma, hemangioma, dermatofibroma, and keloids;and secondary and unspecified malignant neoplasms of connective and soft tissues, retroperitoneum and peritoneum, eye, intraocular melanoma, and other tissues, including adnexa, breast, head and / or neck, anal region, thyroid, parathyroid, adrenal glands and other endocrine glands and associated structures, lymph nodes, respiratory and digestive system and other sites, optionally, wherein the cancer is a primary or metastatic cancer, and optionally, wherein the subject is a mammal, for example, wherein the mammal is selected from a human, cat, dog, or horse; The composition of claim 1.
3. The composition of claim 1 , wherein the sample is cancer cells isolated from a tumor, a peripheral blood sample, or a liquid biopsy.
4. 1. A method of analyzing the presence of clustered mutations for use as an indicator for selecting cancer patients for aggressive treatment, comprising detecting, in a sample isolated from said subject, at least one clustered mutation in a gene selected from one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A, or detecting the absence of said clustered mutation in the BRAF gene; The presence of one or more of the clustered mutations in ARID1A or the absence of the clustered mutations in the BRAF gene indicates the subject should be selected for the treatment, and optionally the cancer cells or cancer are selected from carcinoma, sarcoma, or blood cancer, and optionally the cancer cells or cancer are selected from cancers of the circulatory system, e.g., heart (sarcomas [angiosarcoma, fibrosarcoma, rhabdomyosarcoma, liposarcoma], myxoma, rhabdomyoma, fibroma, lipoma, and teratoma), mediastinum and pleura, and other intrathoracic organs, vascular tumors, and tumor-associated vasculature. tissues; respiratory tract, e.g., nasal cavity and middle ear, paranasal sinuses, larynx, trachea, bronchi and lungs (small cell lung cancer (SCLC), non-small cell lung cancer (NSCLC), etc.), bronchogenic carcinoma (squamous cell, undifferentiated small cell, undifferentiated large cell, adenocarcinoma), alveolar (bronchiolar) carcinoma, bronchial adenoma, sarcoma, lymphoma, chondromatous hamartoma, mesothelioma; gastrointestinal system, e.g., esophagus (squamous cell carcinoma, adenocarcinoma, leiomyosarcoma, lymphoma), gastrointestinal (e.g., colon, colorectum, rectum), stomach (carcinoma, lymphoma, leiomyosarcoma), stomach, pancreas (pancreatic ductal adenocarcinoma, insulinoma, glucagonoma, gastrinoma, carcinoid tumor, vipoma), small intestine (adenocarcinoma, lymphoma, carcinoid tumor, Kaposi's sarcoma, leiomyoma, hemangioma, lipoma, neurofibroma, fibroma), large intestine (adenocarcinoma, tubular adenoma, villous adenoma, hamartoma, leiomyoma); gastrointestinal stromal tumors and neuroendocrine tumors occurring anywhere; genitourinary tract, e.g., kidney (adenocarcinoma, Wilms' tumor [nephroblastoma], lymphoma, leukemia), bladder and / or urethra (squamous cell carcinoma, transitional cell carcinoma, adenocarcinoma), prostate (adenocarcinoma, sarcoma), testis (seminoma, teratoma, embryonal carcinoma, teratocarcinoma, choriocarcinoma, sarcoma, stromal cell carcinoma, fibroma, fibroadenoma, adenomatous tumor, lipoma);Liver, for example, liver cancer (hepatocellular carcinoma), cholangiocarcinoma, hepatoblastoma, angiosarcoma, hepatocellular adenoma, hemangioma, pancreatic endocrine tumors (pheochromocytoma, insulinoma, vasoactive intestinal peptide tumor, pancreatic islet cell tumor, and glucagonoma, etc.); Bone, e.g., osteogenic sarcoma (osteosarcoma), fibrosarcoma, malignant fibrous histiocytoma, chondrosarcoma, Ewing's sarcoma, malignant lymphoma (reticulum cell sarcoma), multiple myeloma, malignant giant cell tumor, chordoma, osteochondroma (osteochondroma exostosis), benign chondroma, chondroblastoma, chondromyxoid fibroma, osteoid osteoma and giant cell tumor; nervous system, e.g., neoplasms of the central nervous system (CNS), primary CNS lymphoma, skull cancer (osteoma, hemangioma, granuloma, xanthomas, osteitis deformans), meninges (meningiomas, meningeal sarcomas, gliomatosis), brain cancer (astrocytoma, medulloblastoma, glioma, ependymoma, germinoma [pinealoma], glioblastoma multiforme, oligodendroma, Dendroglioma, Schwannoma, retinoblastoma, congenital tumors), spinal neurofibroma, meningioma, glioma, sarcoma); reproductive system, e.g., gynecological system, uterus (endometrial carcinoma), cervix (cervical carcinoma, preneoplastic cervical dysplasia), ovary (ovarian carcinoma [serous cystadenocarcinoma, mucinous cystadenocarcinoma, unclassified carcinoma], granulosa theca cell tumor, Sertoli-Leydig cell tumor, dysgerminoma, malignant teratoma), vulva (squamous cell carcinoma, intraepithelial carcinoma) Cancer, adenocarcinoma, fibrosarcoma, melanoma), vagina (clear cell carcinoma, squamous cell carcinoma, botryoid sarcoma (embryonal rhabdomyosarcoma), fallopian tubes (cancer) and other areas related to the female reproductive organs; placenta, penis, prostate, testes, and other areas related to the male reproductive organs; blood system, e.g., blood (myeloid leukemia [acute and chronic], acute lymphoblastic leukemia, chronic lymphocytic leukemia, myeloproliferative disorders, multiple myeloma, myelodysplastic syndromes), Hodgkin's disease oral cavity, e.g., lips, tongue, gums, floor of the mouth, palate, and other parts of the mouth, parotid glands, and other parts of the salivary glands, tonsils, oropharynx, nasopharynx, pyriform sinuses, hypopharynx, and other parts of the lips, oral cavity, and pharynx; skin, e.g., malignant melanoma, cutaneous melanoma, basal cell carcinoma, squamous cell carcinoma, Kaposi's sarcoma, lentil dysplastic nevi, lipoma, hemangioma, dermatofibroma, and keloids;and secondary and unspecified malignant neoplasms of connective and soft tissues, retroperitoneum and peritoneum, eye, intraocular melanoma and other tissues including adnexa, breast, head and / or neck, anal region, thyroid, parathyroid, adrenal glands and other endocrine glands and associated structures, lymph nodes, respiratory and digestive system and other sites, optionally wherein the cancer is a primary or metastatic cancer, and further optionally wherein the subject is a mammal, for example wherein the mammal is selected from a human, cat, dog, or horse; method.
5. 5. The method of claim 4, wherein the sample is cancer cells isolated from a tumor, a peripheral blood sample, or a liquid biopsy.
6. 6. The method of claim 4 or claim 5, wherein the subject has clustered mutations in two or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A.
7. 6. The method of claim 4 or claim 5, wherein the subject further lacks clustered mutations in BRAF.
8. 1. A method of analyzing the presence of clustered mutations for use as an indicator to identify whether a cancer patient is likely to experience a relatively long or short overall survival, comprising assaying for at least one clustered mutation in a gene selected from one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC, ARID1A, or in a sample isolated from said patient, wherein if said clustered mutation is detected in BRAF, said patient will experience a longer overall survival. and if the clustered mutations are detected in one or more of TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A, the patient may experience shorter overall survival, and optionally the cancer cells or cancer are selected from carcinoma, sarcoma or blood cancer, and optionally the cancer cells or cancer are selected from cancers of the circulatory system, e.g., heart (sarcomas [angiosarcoma, fibrosarcoma, rhabdomyosarcoma, liposarcoma], myxoma, rhabdomyoma, fibroma, lipoma and teratoma), mediastinum and pleura, and other intrathoracic organs, vascular tumors. and tumor-associated vascular tissue; respiratory tract, e.g., nasal cavity and middle ear, paranasal sinuses, larynx, trachea, bronchi and lungs (such as small cell lung cancer (SCLC), non-small cell lung cancer (NSCLC)), bronchogenic carcinoma (squamous, undifferentiated small cell, undifferentiated large cell, adenocarcinoma), alveolar (bronchiolar) carcinoma, bronchial adenoma, sarcoma, lymphoma, chondromatous hamartoma, mesothelioma; gastrointestinal system, e.g., esophagus (squamous cell carcinoma, adenocarcinoma, leiomyosarcoma, lymphoma), gastrointestinal (e.g., colon, colorectum, rectum), stomach (carcinoma, lymphoma, leiomyosarcoma), stomach, pancreas (pancreatic ductal adenocarcinoma, insulinoma, glucagonoma, gastrinoma, carcinoma) id tumor, vipoma), small intestine (adenocarcinoma, lymphoma, carcinoid tumor, Kaposi's sarcoma, leiomyoma, hemangioma, lipoma, neurofibroma, fibroma), large intestine (adenocarcinoma, tubular adenoma, villous adenoma, hamartoma, leiomyoma); gastrointestinal stromal tumors and neuroendocrine tumors occurring anywhere; genitourinary tract, e.g., kidney (adenocarcinoma, Wilms' tumor [nephroblastoma], lymphoma, leukemia), bladder and / or urethra (squamous cell carcinoma, transitional cell carcinoma, adenocarcinoma), prostate (adenocarcinoma, sarcoma), testis (seminoma, teratoma, embryonal carcinoma, teratocarcinoma, choriocarcinoma, sarcoma, stromal cell carcinoma, fibroma, fibroadenoma, adenomatous tumor, lipoma);Liver, for example, liver cancer (hepatocellular carcinoma), cholangiocarcinoma, hepatoblastoma, angiosarcoma, hepatocellular adenoma, hemangioma, pancreatic endocrine tumors (pheochromocytoma, insulinoma, vasoactive intestinal peptide tumor, pancreatic islet cell tumor, and glucagonoma, etc.); Bone, e.g., osteogenic sarcoma (osteosarcoma), fibrosarcoma, malignant fibrous histiocytoma, chondrosarcoma, Ewing's sarcoma, malignant lymphoma (reticulum cell sarcoma), multiple myeloma, malignant giant cell tumor, chordoma, osteochondroma (osteochondroma exostosis), benign chondroma, chondroblastoma, chondromyxoid fibroma, osteoid osteoma and giant cell tumor; nervous system, e.g., neoplasms of the central nervous system (CNS), primary CNS lymphoma, skull cancer (osteoma, hemangioma, granuloma, xanthomas, osteitis deformans), meninges (meningiomas, meningeal sarcomas, gliomatosis), brain cancer (astrocytoma, medulloblastoma, glioma, ependymoma, germinoma [pinealoma], glioblastoma multiforme, oligodendroma, Dendroglioma, Schwannoma, retinoblastoma, congenital tumors), spinal neurofibroma, meningioma, glioma, sarcoma); reproductive system, e.g., gynecological system, uterus (endometrial carcinoma), cervix (cervical carcinoma, preneoplastic cervical dysplasia), ovary (ovarian carcinoma [serous cystadenocarcinoma, mucinous cystadenocarcinoma, unclassified carcinoma], granulosa theca cell tumor, Sertoli-Leydig cell tumor, dysgerminoma, malignant teratoma), vulva (squamous cell carcinoma, intraepithelial carcinoma) Cancer, adenocarcinoma, fibrosarcoma, melanoma), vagina (clear cell carcinoma, squamous cell carcinoma, botryoid sarcoma (embryonal rhabdomyosarcoma), fallopian tubes (cancer) and other areas related to the female reproductive organs; placenta, penis, prostate, testes, and other areas related to the male reproductive organs; blood system, e.g., blood (myeloid leukemia [acute and chronic], acute lymphoblastic leukemia, chronic lymphocytic leukemia, myeloproliferative disorders, multiple myeloma, myelodysplastic syndromes), Hodgkin's disease oral cavity, e.g., lips, tongue, gums, floor of the mouth, palate, and other parts of the mouth, parotid glands, and other parts of the salivary glands, tonsils, oropharynx, nasopharynx, pyriform sinuses, hypopharynx, and other parts of the lips, oral cavity, and pharynx; skin, e.g., malignant melanoma, cutaneous melanoma, basal cell carcinoma, squamous cell carcinoma, Kaposi's sarcoma, lentil dysplastic nevi, lipoma, hemangioma, dermatofibroma, and keloids;and secondary and unspecified malignant neoplasms of connective and soft tissues, retroperitoneum and peritoneum, eye, intraocular melanoma and other tissues including adnexa, breast, head and / or neck, anal region, thyroid, parathyroid, adrenal glands and other endocrine glands and associated structures, lymph nodes, respiratory and digestive system and other sites, and optionally the cancer is a primary or metastatic cancer, and optionally the subject is a mammal, for example the mammal is selected from a human, cat, dog, or horse; method.
9. 9. The method of claim 8, wherein the sample is cancer cells isolated from a tumor, a peripheral blood sample, or a liquid biopsy.
10. 10. The method of claim 8 or claim 9, wherein the subject has two or more clustered mutations in TP53, EGFR, KIT, KMT2C, ELF3, APC and ARID1A or lacks clustered mutations in the BRAF gene.
11. 10. The method of claim 8 or claim 9, wherein the subject further lacks clustered mutations in BRAF.