Detection of chromosomal instability
By associating chromosomal instability markers with single events, a new copy number feature space is created, solving the challenges of cancer evolution and tumor heterogeneity analysis in existing technologies and enabling precise treatment selection and prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CARLOS III CANCER RES CENT PUBLIC ASSETS FOUNDATION
- Filing Date
- 2024-11-15
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies struggle to accurately detect and analyze cancer evolution and tumor heterogeneity associated with chromosomal instability (CIN), leading to difficulties in treatment selection.
By designing a novel method, we can associate markers of chromosomal instability with single events, characterize copy number events for each individual, create a new copy number feature space, and use a computer-implemented method to quantify the association between copy number feature sets and markers of chromosomal instability.
It enables precise analysis of cancer evolution and tumor heterogeneity, improves treatment targeting, predicts treatment response and resistance, and provides personalized treatment plans.
Smart Images

Figure CN122439211A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to copy number signatures associated with different types of chromosomal instability, and more particularly to a method for identifying such copy number signatures associated with individual copy number change events in a sample. Background Technology
[0002] More than five million patients are diagnosed with extremely deadly cancers each year. These tumors are characterized by a disordered genome caused by chromosomal instability (CIN). This disorder exacerbates the occurrence and progression of cancer, making these tumors difficult to treat (McGranahan et al. 2021). Accurate detection of CIN is becoming an important biomarker for prognosis, diagnosis, and treatment selection (Macintyre et al. 2018; Drews et al. 2022; Steele et al. 2022). Building on the work of Macintyre et al. 2018, Drews et al. 2022 described a method for evaluating the origin and extent of CIN at the whole genome level. This technology represents a significant step-change in dealing with the most deadly cancer subtypes, which have until now been difficult for precision medicine approaches to overcome.
[0003] However, despite this progress, it remains difficult to dissect the cancer evolution and tumor heterogeneity associated with CIN, which are intrinsically linked to tumor progression and eventual treatment sensitivity and resistance. Summary of the Invention
[0004] The inventors have recognized that current methods for analyzing patterns of chromosomal instability cannot be directly applied to dissecting CIN from single copy number events that are genomically independent. This limits the ability to dissect genomic patterns resulting from historical CIN (sometimes occurring up to 20 years before diagnosis) and ongoing CIN (currently playing a role in the tumor and driving its progression). Mapping each individual copy number event to a specific CIN type is crucial for effectively evaluating CIN-related cancer evolution and tumor heterogeneity, as well as improving therapeutic targeting. Therefore, the inventors have devised a novel method that characterizes copy number instability by associating marker features of chromosomal instability with single events rather than the entire copy number spectrum. This requires designing new copy number features to characterize each copy number event that can be associated with a single event, thereby creating a new feature space characterizing copy number events.
[0005] Therefore, according to a first aspect, a computer-implemented method for characterizing chromosomal instability in a DNA sample obtained from a tumor is provided, the method comprising: obtaining a tumor copy number profile of the sample; quantifying a set of copy number features for each of one or more segments in the copy number profile, wherein the set of copy number features consists of copy number features associated with the corresponding segment (i.e., copy number features quantified relative to each individual segment); and for a segment of the one or more segments, using the features quantified for the segment to determine a measure indicating the association between one or more marker features of chromosomal instability and the segment, wherein the marker features of chromosomal instability include a set of weights for each of a plurality of components, each component being associated with a copy number feature in the set of copy number features, wherein the weights indicate the probability that a mutational process associated with the marker feature will produce a segment having a feature defined by the component.
[0006] The methods described herein may have any one or more of the following optional features.
[0007] The tumor copy number profile can identify one or more genomic segments, each associated with a copy number that differs from that of adjacent segments. The copy number feature set can be quantified for each of the segments, or for one or more segments identified in the copy number profile (e.g., all segments identified on one or more chromosomes). These methods are not limited in this respect and can even efficiently analyze single events (segments). The terms "segment" and "genomic segment" are used interchangeably. The copy number feature set may include or consist of the following: segment size, copy number variation points, and breakpoint counts in one or two flanking regions of a predetermined length. Determining a measure of the association between a marker of chromosomal instability and a segment may include obtaining a value for each of the plurality of components of the segment. Each component can define the segment's characteristics as a predetermined distribution of copy number feature values. Determining a measure of the association between a marker of chromosomal instability and a segment may include determining the probability that the copy number feature of the segment belongs to each of a predetermined distribution associated with a corresponding component of the plurality of components.
[0008] The copy number feature set of a segment may include a summarized value of the breakpoint count in the upstream and downstream flanking regions surrounding the segment over a predetermined length. The summarized value may be the maximum breakpoint count in the upstream and downstream flanking regions surrounding the segment over a predetermined length. The predetermined length may be 2 to 10 Mb, 3 to 8 Mb, or approximately 5 Mb.
[0009] Each copy number feature can be associated with multiple components, and each component can be characterized using a corresponding predetermined distribution. The predetermined distribution for quasi-continuous copy number features (e.g., segment size, copy number variation points) can be a set of Gaussian distributions. The predetermined distribution for copy number features of a counting feature (e.g., breakpoint counts in one or two flanking regions of a predetermined length) can be a set of Poisson distributions.
[0010] The predetermined distribution associated with the component can be a distribution defined by the parameters in Table 1, or a corresponding distribution obtained by fitting a mixture model to the values of a set of copy number features obtained from the copy number spectra of multiple tumor samples. The predetermined distribution may include 15 to 25 Gaussian distributions for segment size, 5 to 15 Gaussian distributions for copy number variation points, and / or 5 to 15 Poisson distributions for breakpoint counts in one or two flanking regions of a predetermined length.
[0011] The copy number feature group may include copy number change points, wherein the copy number change point of the current segment is determined as follows: when the copy number change point relative to the upstream segment is at or above a predetermined threshold, it is determined by referring to the upstream segment of the current segment; when the copy number change point relative to the upstream segment is below a predetermined threshold but the copy number change point relative to the downstream segment is at or above a predetermined threshold, it is determined by referring to the downstream segment of the current segment; and when the copy number change points relative to both the upstream and downstream segments are below the predetermined threshold, it is determined by referring to a reference copy value. The reference copy value may be 2. The predetermined threshold may be -3.
[0012] Each of the one or more segments can be a non-diploid segment. Quantifying the copy number feature set may include using unrounded copy number segments. Obtaining the copy number spectrum may include collapsing and merging near-diploid segments into a diploid state, wherein the near-diploid segments are segments with copy numbers within a predetermined distance of 2. The predetermined distance may be 0.1.
[0013] Determining the measure of association between a marker of chromosomal instability and a segment using a feature quantified for that segment may include: multiplying the probability associated with each of a plurality of components for that segment by a corresponding weight in that marker feature, and summing the results. The probability associated with a component for that segment may be the probability that the segment having the quantified feature was drawn from a distribution associated with that component. Determining the measure of association between a marker of chromosomal instability and a segment may include, for each of a plurality of marker features, multiplying the probability associated with each of the plurality of components for that segment by a corresponding weight in that marker feature, summing the results to obtain an exposure for that marker feature, and normalizing the resulting exposure by dividing each exposure by the sum of the exposures of the plurality of marker features. Determining the measure of association between a marker of chromosomal instability and a segment using a feature quantified for that segment may include: obtaining the probability associated with each of the plurality of components for that segment, and determining the distance between the probability and the weight of the marker feature. The distance can be a cosine distance or an Euclidean distance.
[0014] The method may include determining the sample's exposure to the marker feature by: for each component, obtaining the sum of probabilities associated with the corresponding component for each of a plurality of segments in the copy number spectrum; and determining a subset of components satisfying the equation. (Equation 1) is a linear combination of one or more characteristic features, where E is a vector of size n containing coefficients. E i The method represents the exposure of marker feature i; PbC is a vector of size c, where each element represents the sum of probabilities associated with the corresponding component; and SbC is a matrix of size c multiplied by n, where each value represents the weight of the component in marker feature i. The method may further include assigning marker features of chromosomal instability to segments of the one or more segments, where the marker feature is the highest-valued marker feature among the one or more chromosomal instability marker features for the segment containing the metric.
[0015] A metric for a marker feature can be the probability that the segment is associated with the marker feature.
[0016] The one or more markers of chromosomal instability may be selected from those defined in Table 3, or corresponding markers obtained by quantifying the copy number signatures in multiple tumor samples and identifying one or more markers that may lead to the copy number profiles of the multiple tumor samples through non-negative matrix factorization. The one or more markers may include one or more markers selected from: one or more markers associated with chromosomal missegregation, one or more markers associated with impaired homologous recombination, one or more markers associated with tolerance of whole genome duplication, one or more markers associated with impaired non-homologous end joining, one or more markers associated with replication stress, and / or one or more markers associated with impaired DNA damage perception, wherein the chromosomal missegregation is optionally caused by defective mitosis and / or telomere dysfunction.
[0017] The sample can be obtained from a subject diagnosed with cancer. Sequence data can be obtained using sequencing. Sequencing can be performed using WGS, sWGS, single-cell WGS, WES, or panel sequencing, or copy number arrays. The method may also include determining that the number of non-diploid regions contained in the copy number profile exceeds a predetermined threshold. The copy number profile may contain at least 15 non-diploid regions.
[0018] Obtaining the tumor copy number profile may include determining the tumor copy number profile from sequence data. The sequence data may be next-generation sequencing data. Determining the tumor copy number profile may include determining the copy number of each of a plurality of genomic bins of a predetermined size. The predetermined size may be approximately 100 kb, up to 100 kb, up to 50 kb, approximately 50 kb, or approximately 30 kb. The method may include obtaining the tumor copy number profile from DNA sequence data containing a plurality of sequence reads, wherein the tumor copy number profile includes copy number estimates for each of the plurality of genomic bins, the genomic bins containing a number of mapped reads exceeding a predetermined threshold. The predetermined threshold may be 15 reads.
[0019] The second aspect also describes a method for predicting whether a subject with cancer is likely to respond to treatment, the method comprising: using a method according to any embodiment of the first aspect, characterizing a DNA sample obtained from the subject's tumor as having at least one segment having a metric associated with one or more predetermined chromosomal instability markers that meets predetermined criteria, and / or characterizing a sample exposure to the markers that meets predetermined criteria, wherein the predetermined markers include markers associated with a response to inhibition of a specific gene, the treatment being a treatment targeting the gene, and the subject is likely to respond to the treatment when the sample is characterized as having at least one segment (the segment having a metric associated with the markers that meets predetermined criteria) and / or characterizing a sample exposure to the markers that meets predetermined criteria. For example, markers associated with genome-wide duplication tolerance or PI3K / AKT-mediated genome-wide duplication tolerance (e.g., CX4 or corresponding markers as shown in Table 3) have been shown to be associated with sensitivity to treatments that inhibit CCND1. As another example, biomarkers associated with impaired homologous recombination (e.g., CX5 or corresponding biomarkers as shown in Table 3) have been shown to be sensitive to treatment inhibiting PARP1. As another example, biomarkers associated with replication stress (optionally, where the biomarkers also indicate focal amplification; e.g., CX9 or corresponding biomarkers as shown in Table 3) have been shown to be associated with sensitivity to treatment inhibiting kinases in the mitotic pathway (e.g., EGFR, JAK1, MET, PRKCA, PI3KCA). As another example, biomarkers associated with replication stress (optionally, where the biomarkers also indicate clustered amplification; e.g., CX13 or corresponding biomarkers as shown in Table 3) have been shown to be associated with treatment inhibiting CDK4.
[0020] The third aspect also describes a method for predicting whether a subject with cancer is likely to respond to treatment, the method comprising: using a method according to any embodiment of the first aspect, characterizing a DNA sample obtained from the subject's tumor as having at least one segment having a metric associated with one or more predetermined chromosomal instability markers that meets predetermined criteria, and / or characterizing it as having sample exposure to the markers that meets predetermined criteria, wherein the treatment is a treatment related to resistance to the treatment and gene amplification (optionally, oncogene amplification), and the subject is likely to have gene amplification and is unlikely to respond to the treatment when the sample is characterized as having at least one segment (the segment having a metric associated with the one or more markers that meets predetermined criteria) and / or characterizing it as having sample exposure to the markers that meets predetermined criteria.
[0021] The methods of the second and third aspects may include recommending a treatment containing the treatment to subjects who have been identified as likely to respond to the treatment. The methods of the second and third aspects may also include recommending a treatment not containing the treatment to subjects who have been identified as unlikely to respond to the treatment.
[0022] The methods disclosed herein also cover treating subjects diagnosed with cancer, the methods including determining whether the subject is likely to respond to a treatment as described herein, administering the treatment to a subject determined to be likely to respond to the treatment, and / or administering a treatment not containing the treatment to a subject determined to be unlikely to respond to the treatment.
[0023] According to the fourth aspect, this document also describes a method for determining whether a subject with cancer is likely to undergo gene amplification, the method comprising: characterizing a DNA sample obtained from the subject's tumor using any embodiment of the first aspect as having at least one segment having a measure associated with one or more predetermined chromosomal instability markers that meets predetermined criteria, and / or characterizing it as having sample exposure to the markers that meets predetermined criteria; wherein the presence of at least one segment having a measure associated with one or more predetermined chromosomal instability markers that meets predetermined criteria, and / or having sample exposure to the markers that meets predetermined criteria, indicates that the subject may be likely to undergo gene amplification. The gene may be an oncogene.
[0024] According to the fifth aspect, this document also describes a method for providing a prognosis for a subject suffering from cancer, the method comprising: determining, using the method of the fourth aspect, whether the subject is likely to undergo gene amplification; and determining a prognosis for the subject, wherein a subject identified as likely to undergo gene amplification has a worse prognosis compared to a subject whose sample is characterized as not having at least one segment (the segment having a metric associated with the one or more marker features that meets a predetermined criterion) and / or not having sample exposure to the marker features that meets a predetermined criterion.
[0025] According to the sixth aspect, a method for providing a prognosis for a subject suffering from cancer is also described herein, the method comprising: characterizing a DNA sample obtained from the subject's tumor using any embodiment of the first aspect as having at least one segment having a metric associated with one or more predetermined chromosomal instability markers that meets predetermined criteria, and / or characterizing it as having a sample exposure to the markers that meets predetermined criteria; wherein the presence of at least one segment having a metric associated with one or more predetermined chromosomal instability markers that meets predetermined criteria and / or having a sample exposure to the markers that meets predetermined criteria indicates that the subject has a worse prognosis compared to a subject whose sample is characterized as not having at least one segment (the segment having a metric associated with the one or more markers that meets predetermined criteria) and / or not having a sample exposure to the markers that meets predetermined criteria.
[0026] According to the seventh aspect, this document also describes a method for determining whether a specific process leading to chromosomal instability is currently active in the cancer of the subject, the method comprising: characterizing a sample obtained from the subject using a method according to any embodiment of the first aspect, wherein the tumor copy number profile is obtained from single-cell sequencing data, and the characterization is performed individually using each of a plurality of single-cell tumor copy number profiles; identifying one or more common copy number events and / or one or more unique copy number events, the common copy number events being events present in a proportion above a predetermined threshold in the single-cell tumor copy number profile, the unique copy number events present in a proportion at or below the predetermined threshold in the single-cell tumor copy number profile; wherein a process leading to chromosomal instability associated with a marker feature identified as being associated with a common copy number event may have been active in the tumor prior to at least one clonal amplification in the tumor, and / or a process leading to chromosomal instability associated with a marker feature identified as being associated with a unique copy number event may currently be active in the tumor.
[0027] According to another aspect, one or more non-transitory computer-readable media are provided, which contain instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of any method described herein, such as the method according to any embodiment of any of the foregoing aspects.
[0028] According to another aspect, a computer program containing code is provided that, when executed on a computer, causes the computer to perform the steps of any method described herein, such as the method according to any embodiment of any of the first to seventh aspects.
[0029] According to another aspect, a system is provided comprising: a processor; and a computer-readable medium containing instructions that, when executed by the processor, cause the processor to perform the steps of any method described herein, such as a method according to any embodiment of any of the first to seventh aspects. Attached Figure Description
[0030] Figure 1 It is a flowchart that schematically illustrates a method for identifying one or more copy number markers associated with individual copy number change events in a sample.
[0031] Figure 2 It is a flowchart that schematically illustrates methods for characterizing tumor evolution history and / or heterogeneity, providing prognosis, and identifying treatment or treatment recipients.
[0032] Figure 3 An implementation of a system for performing the methods of this disclosure is shown.
[0033] Figure 4 A schematic diagram of copy number features and CIN marker feature identification according to some embodiments of this disclosure is shown. a) Schematic diagram of feature coding design and feature component derivation. The three feature distributions are segmented into smaller components by applying a variational Bayesian Gaussian mixture model (for segment size and change points) or a finite Poisson mixture model (for breakpoints in the 5Mb flanking). b) For all samples and features, the probability that each copy number event (non-diploid segment) belongs to each mixture component is calculated. The component probabilities of all events are then summed to obtain a 41-dimensional feature vector for each sample. In the illustrative example, the activity of the CIN marker features for all 6,335 samples in the TCGA cohort is already known according to Drews et al. 2022 (see Example 2 below). Therefore, a linear combinatorial decomposition is applied to derive the marker feature definition in a new feature space.
[0034] Figure 5Data on the impact of modifying the feature encoding of the variation points in the CIN marker features are presented. The heatmap shows the Pearson correlation coefficient between the CIN marker feature activity derived using the feature encoding (x-axis) described in Drews et al. (2022) and the CIN marker feature activity derived by extracting five original features but using a modified encoding version of the variation point features described in Example 1. Red indicates a highly positive correlation between activities, while blue indicates a highly negative correlation.
[0035] Figure 6 Histograms and fitted density curves are shown, illustrating the distribution of breakpoints per side feature in the TCGA cluster (see Example 1). For visualization purposes, the data is divided into four subsets based on the number of breakpoints within a 5 Mb window per segment side. Combined distributions are used to derive the different feature components of the signal capturing this new feature I.
[0036] Figure 7 A comparison of CIN marker features encoded using the original features (Drews et al. 2022) and the novel features of this disclosure is shown. A heatmap illustrates the Pearson r correlation coefficients between the activities of the original (x-axis) and novel (y-axis) CIN marker features, calculated in 6,335 TCGA samples. Red indicates a highly positive correlation between activities, while blue indicates a highly negative correlation.
[0037] Figure 8 Results for the metrics used to select the optimal number of marker features in a simulated genome are shown. Four evaluation metrics obtained for each number of marker features (x-axis) are compared. Circles and solid lines represent results for the simulated dataset, while triangles and dashed lines represent results for the input matrix with 1,000 random permutations (these can be considered null measures). Here, the basis refers to the marker feature definition matrix, the coefficients refer to the activity matrix (exposure), and consensus refers to the connectivity matrix of samples clustered by their dominant marker features in 1,000 runs. A value of 5 defines a stable point among the cophenetic coefficients, dispersion coefficients, and silhouette coefficients, and represents the maximum sparsity achievable on top of the null model for the basis matrix.
[0038] Figure 9 Data demonstrating the ability of the method of this disclosure to capture simulated CIN types are shown. The bar chart illustrates the accuracy in identifying the biomarker activity of samples with the dominant mutational process generated by the simulated parameters. The proportion (left) and absolute number (right) of samples correctly identified for each mutational process are shown.
[0039] Figure 10 A comparison of marker features extracted from simulated genomes using the original (Drews et al. 2022) and novel feature encodings (this disclosure) is shown. The heatmap illustrates the Pearson R correlation coefficients between the activities of marker features derived from the simulated dataset using the original (x-axis) and novel (y-axis) feature encodings. Red indicates a highly positive correlation between activities, while blue indicates a highly negative correlation.
[0040] Figure 11 Data on the stability of CIN marker features (markers as defined in this disclosure) across different copy number spectroscopy techniques are presented. A heatmap shows the Kendall's tau correlation coefficients between CIN marker activities derived from SNP6 (x-axis) and sWGS (y-axis) data from the same cohort of 478 tumors. Red indicates a high positive correlation between activities, while blue indicates a high negative correlation.
[0041] Figure 12 shows data on the stability of CIN marker features (marker features as defined in this disclosure) across different copy number spectral analysis resolutions. Figure 12A Density plots are shown, illustrating the cosine similarity between the CIN marker exposure in the spectrum sourced from SNP6 and the marker exposure obtained from the spectrum sourced from 478 sWGS using different block resolutions. Figure 12B A heatmap is shown illustrating the Kendall's tau correlation coefficients between CIN marker feature exposures derived from sWGS data segmented using block sizes of 30 kb (x-axis) and 50 kb (y-axis). Comparisons were made using the same set of 478 tumors. Red indicates a highly positive correlation between exposures, while blue indicates a highly negative correlation.
[0042] Figure 13 Data show the number of individual events required to quantify sample-level marker features using sample marker features defined by three dominant marker features for each sample.
[0043] Figure 14 shows the data on the number of individual events required to quantify the level of individual marker characteristics in the study sample. Figure 14A Data for the marking features CX1, CX4, and CX9 are shown. Figure 14B Data for the marking features CX2, CX5, and CX10 are shown. Figure 14C Data for the marking features CX3, CX8, and CX13 are shown.
[0044] Figure 15This diagram shows whole-genome copy number profiles of clones derived from two normal diploid cell lines (hTERT-HMEC and hTERT-RPTEC) after reversine treatment. The heatmap illustrates the whole-genome copy number profiles of different clones. Each row represents a clone, and the columns represent different genomic locations. Color indicates the copy number at a specific genomic location.
[0045] Figure 16 The labeling features weights in copy number events induced after resveratrol treatment are shown. Each point represents a copy number event, and each color represents a different clone.
[0046] Figure 17 This shows the TP53 from hTERT-RPE1. - / - BRCA2 - / - Whole genome copy number profiles of single cells from cell line models. The heatmap shows the whole genome copy number profiles of different single cells sequenced from the knockout model. Each row represents a cell, and the columns represent different genomic locations. The color indicates the copy number at a specific genomic location.
[0047] Figure 18 This is shown in hTERT-RPE1 TP53 - / - Marker feature weights in copy number events induced by BRCA2 gene knockout in cell lines. Each point represents a copy number event, and each color represents a different clone.
[0048] Figure 19 This shows the TP53 from hTERT-RPE1. - / - BRCA1 - / - Whole genome copy number profiles of single cells from cell line models. The heatmap shows the whole genome copy number profiles of different single cells sequenced from the knockout model. Each row represents a cell, and the columns represent different genomic locations. The color indicates the copy number at a specific genomic location.
[0049] Figure 20 This is shown in hTERT-RPE1 TP53 - / - Marker feature weights in copy number events induced by BRCA1 gene knockout in cell lines. Each point represents a copy number event, and each color represents a different clone. Detailed Implementation
[0050] In this disclosure, the following terms will be used and are intended to be defined as indicated below.
[0051] This disclosure broadly relates to the characterization of DNA samples, particularly tumor samples, in terms of their copy number profiles.
[0052] "Copy number profiling" refers to the quantification of the copy number of each of multiple parts of a genome sequence. In the context of this disclosure, copy number profiling is preferably whole-genome copy number profiling. Copy number profiling is obtained by analyzing sequencing data. For example, copy number profiling can be obtained by obtaining sequence data from a genomic DNA sample (or a DNA library derived from genomic DNA, as further explained below, including samples of DNA derived from genomic DNA obtained through fragmentation (e.g., cell-free DNA)) and quantifying the copy number of each part of the genome sequence (e.g., each block, where the block may be, for example, a 30kb region, a 50kb region, or a 100kb region), as known in the art. The methods of this disclosure are not limited in terms of the size of the blocks used. Smaller or larger blocks (i.e., higher or lower resolution) can be used (e.g., depending on the available data), and marker features and / or marker activity can be extracted at any such resolution. Furthermore, marker features extracted at a first (lower) resolution can be mapped to marker features extracted at a second (higher) resolution, for example, by identifying the marker feature at a second resolution that is most correlated with the marker feature at the first resolution in terms of marker feature activity across multiple samples (where the sample can be an individual cell—e.g., when using single-cell sequencing data, or it can be multiple cells or tissues) (e.g., using Spearman Rho or any other correlation metric known in the art). “Tumor copy number profile” refers to a copy number profile associated with a tumor genome. A tumor copy number profile can be obtained by sequencing a sample containing tumor genomic DNA, as further explained below. In some embodiments, a tumor copy number profile is a copy number profile obtained by sequencing a genomic DNA sample derived from tumor cells. Therefore, a copy number profile is generally, in the form of an estimate of the copy number of a segment in a sample or subset of samples for each of multiple genomic segments (e.g., in the case of a sample containing cells or genetic material derived from them with different genotypes, such as a sample containing a mixture of tumor and normal cells or tumor and normal DNA). Copy number profiles can be obtained from read count data by segmentation and copy number estimation using methods known in the art. For example, segmentation can be performed using methods such as cyclic binary segmentation. Copy number fitting can be performed as described in the following examples. Alternatively, segmentation and copy number estimation can be performed using methods such as Ascat (Raine et al. 2016), ichorCNA (Adalsteinsson et al. 2017), etc. (and in the cases of Ascat and ichorCNA, % tumor DNA estimation can also be performed). Copy number profiles can be whole-genome copy number profiles. Whole-genome copy number profiles can be obtained from samples using WGS, sWGS, or single-cell WGS.
[0053] In copy number profiling, a “segment” refers to a portion of the sequence represented in the copy number profile that is associated with a consistent absolute copy number. This consistent copy number differs from the copy number associated with the directly upstream sequence (if such a sequence exists and is associated with the copy number estimate) and the directly downstream sequence (if such a sequence exists and is associated with the copy number). In other words, a segment is a portion of the sequence with a copy number associated with it that differs from the copy number associated with its immediate neighboring segments. The copy number associated with a segment may differ from the copy number associated with its immediate neighboring segments because segments surrounding the segment are associated with different copy numbers, because segments surrounding the segment are not associated with the copy number (e.g., due to data loss or insufficient quality of the segment), or a combination of both (e.g., the segment may be surrounded on one side by segments associated with different copy numbers and on the other side by segments not associated with copy numbers). In other words, a segment is the longest contiguous portion of the copy number profile, each associated with a single copy number. A segment may be associated with a set of coordinates (e.g., genomic coordinates) that defines the segment’s boundaries. Each boundary can be associated with a copy number change point (also referred to as a "change point"). In some embodiments, the copy number profile contains up to 350 segments (copy number events), up to 300 segments, or up to 250 segments. Preferably, the copy number profile contains up to 250 segments (copy number events). These numbers can be particularly useful when looking at a human whole-genome copy number profile. It is not desirable to be bound by theory that a higher number of segments may indicate undesirable DNA degradation (e.g., formalin-mediated DNA degradation). In practice, the copy number of a segment can be a copy number estimate obtained using methods used to determine tumor copy number profiles (e.g., ASCAT [Van Loo et al., 2010] or Sequenza [Favero et al., 2015]). According to this disclosure, a "copy number event" refers to a single segment in the copy number profile. In particular, a copy number event can refer to a segment whose copy number deviates from the expected copy number in a normal genome (e.g., a non-diploid segment). A copy number event / segment is characterized by multiple copy number features, which will be further defined below. Therefore, the terms "copy number event" and "segment" are used interchangeably. Note that this is not necessarily the case in the prior art, for example in WO 2023 / 057392, where copy number features are determined for copy number events encompassing multiple segments. In this disclosure, all features are determined relative to a specific copy number event / segment.
[0054] Chromosomal instability (CIN) is the process of the accumulation of numerical and structural changes in DNA. A marker feature of CIN (also referred to herein as a copy number marker or “marker feature”) is a genome-wide imprint representing different putative mutational processes (where “mutational process” as used herein refers to any process capable of causing chromosomal instability). CIN marker features are described in Macintyre et al. 2018, Drews et al. (2022), and WO 2023 / 057392, which are incorporated herein by reference. The presence or absence of a CIN marker feature in a sample can be determined by calculating the sample’s exposure to (also referred to as “activity”). A “CIN marker feature” is a set of weights associated with each of several components. The sum of the weights may be 1, and each weight may represent the probability that the mutational process associated with the marker feature will produce a copy number event with a specific feature defined by the component. In other words, the multiple weights of a marker feature can each represent the probability that a mutation process associated with that marker feature will produce a specific type of copy number event. Each type of copy number event is called a "component." Each component is defined by a predetermined distribution of the possible values of the copy number feature.
[0055] The term "copy number (CN) feature" refers to the characteristics of observable copy number events in a copy number spectrum. According to this disclosure, copy number features include: segment size (also called "segment length," typically expressed in base count), point-of-change copy number (the absolute difference in copy number between a segment and its adjacent / nearby segments in the copy number spectrum, which can be defined relative to an upstream segment, a downstream neighboring segment, or a reference copy number state (typically 2 in the analysis of a normal diploid genome), and the number of breakpoints per flanking region (also called "breakpoint per side," the number of breakpoints being the change in copy number / connectivity between individual segments in a segment's flanking region, which can be upstream or downstream of the segment). The predetermined length of the flanking regions used to determine the "breakpoint per side" feature can be 2 to 10 megabases (Mb), preferably 3 to 8 Mb, or about 5 Mb. The number of breakpoints per flanking region can be determined for one or two flanking regions, and when two flanking regions are used, a generalized measure is obtained. A generalized metric may be the maximum or average number of breakpoints observed in the upstream and downstream flanking regions. In some embodiments, the breakpoint characteristic per flank of a segment is the maximum number of breakpoints observed in the upstream or downstream flanking regions of the segment. Copy number characteristics according to this disclosure do not include any characteristics that are not uniquely associated with an individual event. In other words, copy number characteristics according to this disclosure do not include any copy number characteristics quantified on genomic regions defined by a non-referenced segment. In particular, copy number characteristics according to this disclosure do not use any of the following: breakpoint count per x MB (the number of change points appearing in a sliding window on the copy number spectrum), breakpoint count per chromosome arm (the number of change points appearing in each chromosome arm), and the number of segments with oscillating copy numbers (sometimes referred to as "the length of segments with oscillating copy numbers"; the number of consecutive segments alternating between two copy number states, rounded to the nearest integer copy number state; also referred to as the length of the oscillating copy number state chain). In some embodiments, the copy number features used according to the invention do not include segment copy number (the observed absolute copy number state of each segment, also referred to herein as "copy number" or "absolute copy number"). Each CN feature is associated with one or more components. Each component is defined by a distribution. CN features are observed for individual copy number events. Typically, CN features for each event are observed on a genome-wide basis for a sample or sample set, and a generalized measure of each event in the sample is obtained for each component. These generalized measures, in turn, are used to determine the marker feature exposure of the sample. According to this disclosure, marker feature exposure is determined for individual events and optionally also for the sample (i.e., for the entire copy number spectrum).
[0056] Methods for determining exposure to marker features from copy number spectra (i.e., at the sample level) and for identifying marker features from a sample cohort are known in the art (see, for example, MacIntyre et al., 2018, Drews et al. (2022)). In particular, determining the exposure of a sample to one or more CN marker features (i.e., the activity of one or more marker features in the sample) can be achieved by identifying features that satisfy… Using matrix E, in In this matrix, C is the mutation class of one or more samples for which exposure is to be determined (also known as the patient-component, PbC matrix), P is the marker matrix containing one or more CN marker features for which exposure is to be determined (also known as the marker-component, SbC matrix), and E is the exposure matrix (also known as the patient-marker feature, PbS). For example, exposure to a marker feature can be determined by performing matrix factorization to identify values (or vectors of values) that satisfy the following equation: (Or, in practice) , where ε is the residual term to be minimized), where SbC is a row of the marker feature-component matrix corresponding to a specific marker feature in a set of marker features, and PbS is a vector of patient-marker feature values (or values). Exposure to multiple copy number marker features can be calculated using the corresponding rows of the marker feature-component matrix. In some embodiments, the exposure (E) to copy number marker feature i (from a set of n marker features, where n can be, for example, 16 or 17, and the marker features can include any or all of the marker features disclosed herein or their corresponding marker features) is a value E satisfying the following equation. i : (Equation 1), where: E is a vector of size n, containing coefficients E i PbC is the exposure of the signature feature i; PbC is a vector of size c, where each value represents the sum of the posterior probabilities of each copy number event in the copy number spectrum belonging to component C, where each component C is the distribution of values for the copy number feature; and SbC is a matrix of size c multiplied by n, where each value represents the weight of component C in the copy number signature feature i. This is equivalent to solving a minimization problem under the nonnegativity constraint of E, given SbC and PbC. For example, the linear combinatorial decomposition implemented in the R package YAPSA (Hübschmann, D. et al. 2021) can be used for this purpose. Similar methods can be used to identify marker features using sample groups and nonnegative matrix decomposition (see, for example, Drews et al. 2022) to identify both the exposure matrix E and the marker feature matrix SbC. In addition to determining the sample exposure to the marker feature, this disclosure also provides for determining the association between the marker feature and individual segments (particularly copy number events) in the copy number spectrum. This will be explained further below.
[0057] In the context of this disclosure, each component represents the type of copy number abnormality defined by an individual component of a mixture model fitted to a distribution of values for each of the three copy number features described herein, obtained from a sample cohort from the Cancer Genome Atlas (TCGA). Note that any other cancer sample cohort may be used for this purpose, and this disclosure does not in any way limit the sample cohort from which components (and corresponding marker features) are extracted. This follows the procedure described in Drews et al. 2022, except that the three features in Drews et al. 2022 (breakpoints per chromosome arm, breakpoints per 10 Mb, and the number of segments with oscillating copy numbers) are removed and replaced with new features as described herein (breakpoints per side). Components from Drews et al. 2022 are used for segment size and change point features, and new components are determined for the new features. Using these, copy number marker features (i.e., marker features of copy number change processes) are obtained, which capture the probability of a copy number change process leading to a copy number event distributed according to each copy number feature component. Therefore, in this context, the term "exposure" (also known as "marker activity" or "activity") reflects the strength of evidence for the existence of copy number change events attributable to the marker features.
[0058] The term "whole genome" refers to data that covers the entire genome or a large portion thereof (e.g., one or more complete chromosomes). As understood by those skilled in the art, sequencing data or copy number profiles derived from it can be referred to as "whole genome," even if it may not contain information about each and every individual location in the genome (i.e., reads or copy number estimates). In fact, even in the context of sWGS, there may be genomic regions without mapped reads due to factors such as low mappability, sequencing difficulties, or the sampling nature of next-generation sequencing. Therefore, data can be referred to as "whole genome" if it contains data about one or more chromosomes present in the reference genome, or at least 70%, at least 80%, or at least 90%, preferably at least 90%, of all chromosomes.
[0059] As used herein, “sample” can refer to cell or tissue samples, biofluids, or extracts (e.g., DNA extracts obtained from cell, tissue, or fluid samples) from which genomic material can be obtained for gene analysis, such as genome sequencing (e.g., whole-genome sequencing, or targeted sequencing such as whole-exome sequencing or targeted panel sequencing). A sample can be a cell, tissue, or biofluid sample obtained from an object (e.g., a biopsy). Such a sample may be referred to as an “object sample.” A sample can be any sample containing genomic DNA or cell-free DNA. In particular, a sample can be a tumor sample, a biofluid sample containing DNA or cells, a blood sample (including plasma or serum samples), a urine sample, a cervical smear, an ascites sample, or a sample derived therefrom (e.g., after DNA purification). Urine, ascites, and cervical smears have been found to contain cells and therefore can provide suitable samples for use according to the invention. Other sample types suitable for use according to the invention include fine-needle aspirates, lymph node samples (e.g., aspirates or biopsy), surgical margins, bone marrow, or other tissues from the tumor microenvironment from which trace amounts of tumor DNA can be found or are expected to be found. The sample may be a fresh sample obtained from the subject, or it may be a sample that has been processed and / or stored (e.g., frozen, fixed, or subjected to one or more purification, enrichment, or extraction steps) prior to genomic / transcriptome analysis. The sample may be a cell or tissue culture sample. Therefore, the sample as described herein can refer to any type of sample containing cells or genomic material derived therefrom, whether from a biological sample obtained from the subject or from, for example, a cell line. In some embodiments, the sample is a sample obtained from the subject (e.g., a human subject). The sample is preferably from a mammal (e.g., a mammalian cell sample or a sample from a mammalian subject (e.g., a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit, or guinea pig)), and preferably from a human (e.g., a human cell sample or a sample from a human subject). Furthermore, the sample may be transported and / or stored, and collection may occur at a location remote from the sequence data acquisition (e.g., sequencing) location, and / or any computer-implemented method steps described herein may occur at a location remote from the sample collection location and / or remote from the sequence data acquisition (e.g., sequencing) location (e.g., computer-implemented method steps may be performed via a networked computer, for example, through a "cloud" provider). The sample can be a tissue biopsy, such as a fresh frozen tissue sample or a formalin-fixed paraffin-embedded (FFPE) tissue sample, or a sample containing circulating tumor DNA (ctDNA), such as a biological fluid sample (sometimes called a liquid biopsy).The samples used in the methods of this disclosure are typically samples containing tumor cells (e.g., tumor samples or samples containing circulating tumor cells) or genetic material derived from tumor cells (e.g., cell-free DNA, or DNA extracted from samples containing tumor cells (e.g., from cell lines or tissue samples) or circulating tumor cells). Samples can be “mixed” samples containing cells or genetic material derived from them with different genotypes. For example, a sample can be a sample containing tumor cells and normal cells or DNA derived from them (e.g., in the context of a cell-free DNA (cfDNA) sample containing circulating tumor DNA (ctDNA)). Using methods known in the art, in samples containing both cells with genotypes exhibiting somatic copy number alterations (e.g., tumor cells) and cells without such alterations (e.g., normal cells), the copy number profile of cells with genotypes exhibiting somatic copy number alterations can be determined separately. For example, sequence analysis procedures for deconvolving tumor and germline genomes from copy number spectra include ASCAT (Raine et al., 2016), ABSOLUTE (Carter et al., 2012), or, in the context of cfDNA, ichorCNA (Adalsteinsson et al., 2017). A “tumor sample” refers to a sample containing tumor cells or genetic material derived from them. A tumor sample can be a cell or tissue sample obtained directly from a tumor (e.g., a biopsy). A tumor sample can also be a sample containing tumor cells or genetic material derived from tumor cells but not directly from the tumor. For example, a tumor sample can be a sample containing circulating tumor cells or circulating tumor DNA. Therefore, a tumor sample can also be a biological fluid (e.g., a liquid biopsy, such as a blood, urine, or cerebrospinal fluid biopsy). One or more processing steps can be performed on samples containing a mixture of tumor cells and other cells (or genetically derived material) either before or after sequence data acquisition to identify sequence data representing genetic material from the tumor. For example, a sample containing cells may undergo one or more cell purification steps, said steps selectively enriching tumor cells in the sample. As another example, a genetic material sample may undergo one or more capture and / or size selection steps to selectively enrich tumor-derived genetic material in the sample. Methods for doing so are known in the art. As another example, sequence data may undergo one or more filtering steps (e.g., based on segment length in the context of a cell-free DNA sample) to enrich data with information related to tumor-derived genetic material. Methods for doing so are known in the art. In some embodiments, the sample is a sample containing tumor cells.Preferably, the tumor purity of such a sample (where tumor purity can be quantified as the proportion of cells in the sample that are tumor cells) is at least 30%, at least 35%, at least 40%, at least 45%, or at least 50%. Advantageously, the tumor purity of the sample is at least 40%. It is not desirable to be theoretically constrained, but it is considered that copy number profiles generated from samples with lower tumor purity may be less suitable because signals corresponding to the tumor genome may be lost in signals from the genomes of other cells. A “normal sample” (also called a “germ sample”) refers to a sample containing non-tumor or non-modified cells or genetic material derived therefrom. A normal sample can be matched with a particular tumor or modified sample in the sense that it was obtained from the same biological source (subject or cell line) as the tumor or modified sample. A matched normal sample is a sample that is expected to represent the genome of the tumor sample in the absence of somatic cellular abnormalities. This is typically a sample obtained from the same subject as the tumor sample. In the context of this invention, normal samples can be used as controls for analyzing tumor samples. For example, tumor samples and (typically matched) normal samples can be analyzed together to obtain a copy number profile of the tumor. Methods for obtaining tumor copy number profiles from normal and tumor sample pairs are known in the art. Examples of such methods include ASCAT [Van Loo et al. 2010] and Sequenza [Favero et al. 2015]. Such methods provide tumor copy number profiles as well as an estimate of the purity (proportion of tumor cells) of the tumor sample. The purity estimate can be used to exclude tumor samples with low purity that may therefore lead to low-quality copy number estimates.
[0060] The term "sequence data" refers to information indicating the presence of genetic material with a specific sequence in a sample. Such information can be obtained using sequencing technologies such as next-generation sequencing (NGS), whole-exome sequencing (WES), whole-genome sequencing (WGS), or sequencing of captured genomic loci (targeted or panel sequencing) or array technologies such as SNP arrays or other molecular counting assays. In some embodiments, sequence data is obtained by DNA sequencing, and particularly next-generation sequencing. In such embodiments, sequence data includes sequencing reads or information derived therefrom, such as a count of the number of sequencing reads with a specific sequence. When using non-digital technologies such as array technologies, sequence data may include signals (e.g., intensity values) indicating the number of sequences with a specific sequence in the sample, for example, by comparison with appropriate controls. Using methods known in the art (e.g., Bowtie (Langmead et al., 2009) or BWA (Li and Durbin 2009)), information such as the count of sequencing reads with specific sequences can be derived from sequencing reads by mapping sequence data to a reference sequence, such as a reference genome. This yields aligned sequencing reads, for example, in the form of SAM or BAM files. Thus, sequence data can be associated with specific genomic locations (where “genomic location” refers to the location in the reference genome to which the sequence data is mapped). Sequence data can be filtered as part of the methods described herein, or can be filtered prior to the application of the methods described herein, for example, to remove low-quality reads or reads mapped to regions with “low mappability.” Methods for filtering low-quality reads are known in the art and include, for example, removing duplicate reads, removing supplementary alignments, removing reads with low sequencing quality scores (reads below Q20 or Q30), removing reads with low mapping quality scores (e.g., MAPQ < 37), etc. Methods for filtering reads mapped to regions with low mappability are known in the art and include, for example, Umap (Karimzadeh et al. 2018) and GenMap (Pockrandt et al. 2020).
[0061] The compositions described herein may be pharmaceutical compositions that additionally include pharmaceutically acceptable carriers, diluents, or excipients. Pharmaceutical compositions may optionally include one or more additional pharmaceutically active peptides and / or compounds. Such formulations may be in forms suitable for intravenous infusion, for example. As used herein, “treatment” means the reduction, relief, or elimination of one or more symptoms of the treated disease relative to symptoms prior to treatment. As used herein by any aspect of the invention, “patient” is intended to be equivalent to “object” and specifically includes both healthy individuals and individuals suffering from a disease or condition (e.g., a proliferative condition, such as cancer). Preferably, a patient is a human patient. In some cases, a patient is a person who has been diagnosed with cancer, is suspected of having cancer, or has been classified as being at risk of developing cancer.
[0062] As used herein, the term "computer system" includes hardware, software, and data storage devices for implementing a system or method according to the embodiments described above. For example, a computer system may include a central processing unit (CPU) and / or a graphics processing unit (GPU), input devices, output devices, and data storage, which may be embodied as one or more connected computing devices. A computer system may include a display, or a computing device having a display to provide visual output display. Data storage may include RAM, disk drives, or other non-transitory computer-readable media. A computer system may include multiple computing devices connected via a network and capable of communicating with each other through the network. It is explicitly contemplated that a computer system may consist of or include cloud computers. As used herein, the term "computer-readable media" includes, but is not limited to, any non-transitory media or media that can be directly read and accessed by a computer or computer system. Media includes, but is not limited to, magnetic storage media (e.g., floppy disks, hard disk storage media, and magnetic tape); optical storage media (e.g., optical discs or CD-ROMs); electrical storage media (e.g., memory, including RAM, ROM, and flash memory); and combinations thereof (e.g., magnetic / optical storage media).
[0063] Chromosomal instability markers at the single copy number event level
[0064] This disclosure provides a method for characterizing DNA samples by identifying chromosomal instability markers associated with individual copy number alterations in the sample.
[0065] Figure 1This is a flowchart illustrating a method for characterizing a DNA sample according to this disclosure. In optional step 10, a DNA sample is obtained from the tumor of the subject. Optionally, a matching normal sample may also be obtained from the subject. In optional step 12, sequence data is obtained from the tumor (and optionally the matching normal) DNA sample, for example, using sequencing or array techniques. In step 14, a tumor copy number profile of the sample is obtained, for example, by analyzing the sequence data obtained in step 12 or by receiving the copy number profile from a user, database, etc. Therefore, steps 10 to 12 are optional, as the method described herein can begin with a previously obtained tumor copy number profile of the sample. The data can be batch sequence data (from batch DNA sequencing technology or array technology) or single-cell sequencing data (from single-cell DNA sequencing technology). When using single-cell sequencing data, the process described with reference to steps 12 to 26 can be performed individually for each of multiple single cells. Step 14 of obtaining the tumor CN profile may include one or more of the following: aligning multiple sequence reads to a reference genome; determining that the number of non-diploid regions contained in the copy number profile is higher than a predetermined threshold (e.g., at least 15 or at least 20 events in the copy number profile); determining the copy number of each of multiple genomic blocks of predetermined sizes; excluding genomic blocks containing fewer mapped reads than a predetermined threshold; and / or performing absolute copy number fitting. The predetermined block size may be up to 100 kb, up to 50 kb, or approximately 30 kb. The predetermined threshold for the number of mapped reads may be 15 reads.
[0066] In step 16, for each of one or more segments in the copy number spectrum, a set of copy number features is quantified. This set of copy number features consists of copy number features associated with the corresponding segment. In other words, the set of copy number features may not include any copy number features defined for portions of the genome that are not segments or not defined with reference to a segment. Therefore, each copy number feature is defined for a specific segment and quantifies the characteristics of that segment and / or regions defined with reference to the segment (e.g., flanking regions of a predetermined length, regions of a predetermined length centered on that segment, etc.).
[0067] Step 16 includes quantifying the copy number feature set for each segment. This step may use uncrushed copy number segments. This step may include merging and combining near-diploid segments into a diploid state, where near-diploid segments are segments with copy numbers within a predetermined distance of 2. The predetermined distance may be 0.1. Thus, quantifying the copy number feature set may include allocating (i.e., merging) copy number 2 to any segment with copy numbers within a predetermined boundary (e.g., greater than 1.9 and less than 2.1), and merging any consecutive segments with the same copy number as a result of the allocation. This can advantageously avoid including signals from segments that may be normal diploid segments and which may act as noise in the process. Uncrushed copy number segments may be segments whose copy numbers have not yet been rounded to a near-integer. Therefore, the method preferably uses an uncrushed copy number spectrum, and / or the only rounding performed involves merging near-diploid segments. This utilizes more information present in the copy number spectrum (in terms of the number of segments that would otherwise be merged, and in terms of the copy number information for each segment contained in the spectrum). The copy number of a segment can be rounded at the cost of information loss to remove noise from the copy number spectrum. It has previously been shown that by merging segments that deviate slightly from the “normal” copy number state, it is possible to at least partially compensate for the additional noise associated with using segments with unequal copy numbers while minimizing the information loss related to CIN, and the inventors have shown that this also applies to the novel method described herein.
[0068] In step 16, the set of copy number features quantified for each segment may include segment size features, copy number variation point features, and per-flank breakpoint features (breakpoint counts in one or two flanking regions of a predetermined length). Features of diploid segments may not be quantified. In other words, the method of this disclosure may include quantifying multiple copy number features for one or more non-diploid segments, such as each non-diploid segment in a copy number spectrum. Copy number variation point features may be quantified for the first segment of any chromosome by a copy number variation relative to the diploid state (i.e., subtracting 2 from the copy number of the segment). In other words, if the first segment of any chromosome is not a normal segment, the variation point feature may be quantified by subtracting 2 from the absolute copy number of said segment. If the segment is not a normal segment, this assumes that the (non-existent) preceding segment is a normal segment (copy number 2). The copy number variation point feature of the current segment may be quantified relative to upstream or downstream segments. Specifically, when the change point value relative to the upstream segment is at or above a predetermined threshold (e.g., -3 or greater, such as -3, -2, -1, 1, 2, 3, etc.), the copy number change point can be quantified relative to the upstream (left) adjacent segment. When the change point value relative to the upstream segment is below a predetermined threshold (e.g., less than -3, such as -4, -5, etc.), the copy number change point can be quantified relative to the downstream (right) adjacent segment. When the copy number change point relative to the downstream adjacent segment is also below -3, the copy number change point can be calculated by subtracting two from the copy number of the current segment (i.e., calculating the difference between (the copy number of the current segment) and (the diploid state)). In any embodiment, the copy number change point can be quantified as an absolute value. In other words, the relative copy number change point between the current segment and the adjacent segment can be calculated, and its absolute value can then be used as the copy number change point of the current segment. This method differs from that of Drews et al. (2022), where the change point reflects the absolute copy number difference compared to the left adjacent segment, regardless of whether the relative change is positive or negative. The method proposed in this paper advantageously avoids associating events immediately adjacent to focal amplification with characteristics of replication stress, which is precisely the case in previous methods, where segments with low copy numbers immediately adjacent to focal amplification were defined as having high copy number changes. In other words, the left segment can be used by default unless it is determined that the left segment may be a segment with focal amplification. In this case, the right segment can be used unless it is determined that the right segment may also be a segment with focal amplification. In such cases, the copy number change point can be calculated relative to the diploid state.
[0069] In step 18, a measure indicative of the association between one or more chromosomal instability marker features and each segment is determined using features quantified for the segment. The chromosomal instability marker features comprise a set of weights for each of a plurality of components, each component associated with a copy number feature in a set of copy number features, where the weights indicate the probability that a mutation associated with the marker feature will produce a segment having the feature defined by that component. Therefore, step 18 may include, for each segment, obtaining a value for each of the plurality of components for that segment. Each component defines the feature of the segment as a predetermined distribution of copy number feature values. Determining the measure indicative of the association between the chromosomal instability marker features and the segment may include determining the probability that the copy number feature of the segment belongs to each of a predetermined distribution associated with the corresponding component among the plurality of components. Thus, the probability that a feature value from the segment belongs to the predetermined distribution associated with that component can be determined. This can be done for each component, thereby obtaining an “event-component” probability vector for each segment, or an event-component probability matrix comprising the corresponding vectors of each of the plurality of segments from the same copy number spectrum. These probabilities may be referred to as “posterior probabilities.” Therefore, step 18 may include quantifying the posterior probability that each feature value belongs to each of a predetermined set of distributions, also referred to as “components”. Components can be empirically identified by quantifying copy number features in multiple tumor samples. Each copy number feature may be associated with multiple components. Posterior probabilities for each component and for each event can be obtained. Furthermore, a generalized measure of the copy number spectrum can be obtained for each such component, for example, by summing the posterior probabilities obtained for each event for that component (step 20). Using empirically determined components implies that the features truly reflect biological processes, rather than arbitrary categories that may be biologically unrelated, even if convenient to manipulate. For example, a predetermined set of distributions can be identified by applying mixture modeling to a dataset containing groups of copy number features quantified for multiple tumor samples. This allows for the identification of the true state of copy number features present in tumors and thus the quantification of evidence of the presence of these states in new samples to be analyzed. Using the posterior probability that each feature value belongs to each of a predetermined set of distributions advantageously allows the method to handle the uncertainty of assigning segments to states (i.e., the uncertainty of whether the segment provides evidence of the presence of a particular class of CIN events characterized by one of the distributions of copy number features). This means that the method achieves more accurate characterization of samples by being designed to handle unavoidable noise in the data. For any quasi-continuous feature (e.g., segment size or change point features), the predetermined set of distributions for the copy number feature can be a Gaussian distribution. For any counting feature (e.g., breakpoint features), the predetermined set of distributions can be a Poisson distribution. A set of predetermined distributions can be obtained using hybrid modeling techniques. Quasi-continuous features can be any non-counting feature, such as segment size and copy number change points.For example, the segment size feature of an event can be quantified by obtaining the posterior probability for each of multiple Gaussian distributions (e.g., 20 to 25 Gaussian distributions, or, for a copy number spectrum, the sum of the posterior probabilities of copy number events in the copy number spectrum). As another example, the copy number variation point feature can be quantified by obtaining the posterior probability for each of multiple Gaussian distributions (e.g., 5 to 15 Gaussian distributions, or, for a copy number spectrum, the sum of the posterior probabilities of copy number events in the copy number spectrum). As yet another example, the per-sided breakpoint feature can be quantified by obtaining the posterior probability for each of multiple Poisson distributions (e.g., 5 to 15 Poisson distributions, or, for a copy number spectrum, the sum of the posterior probabilities of copy number events in the copy number spectrum). The exact number of distributions used can depend at least in part on the number and diversity of samples used to obtain the characteristic features. In the embodiments provided herein, the inventors used a very large number of samples from a wide variety of tumor types, enabling them to provide subtle differences in copy number signature behavior in cancer, resulting in a higher number of components than, for example, identified in Macintyre et al., 2018. Parameters for each predetermined distribution (e.g., the mean and variance of a Gaussian distribution, λ for a Poisson distribution) can be determined as part of the process of obtaining chromosomal instability marker signatures. Parameters for each predetermined distribution (e.g., the mean and variance of a Gaussian distribution, λ for a Poisson distribution) can be determined by fitting a mixture model to a set of quantified copy number signatures from multiple tumor samples for which chromosomal instability marker signatures have been obtained. In the embodiments below, these parameters were obtained by analyzing more than 6,000 samples from the Cancer Genome Atlas (TCGA). As those skilled in the art will understand, slightly different parameters can be obtained when different data are used. Parameters for each of a set of Gaussian distributions can be obtained using a variational Bayesian Gaussian mixture model. Specifically, the parameters of the Gaussian distribution (e.g., for features such as segment size and points of change) can be obtained by fitting a Dirichlet-Process Gaussian mixture model using variational inference. The parameters of each of a set of Poisson distributions (e.g., breakpoint features on each side) can be obtained using a Poisson mixture model and / or using predetermined parameters. Table 1 provides the parameters of the components used in the embodiments described herein. Thus, in some embodiments, step 16 uses the components defined in Table 1 or corresponding components defined using different sample groups. A corresponding component refers to a component defined by a distribution corresponding to the component distributions in Table 1, but the exact parameters of each distribution can vary. The component parameters for the points of change and segment size features in Table 1 are as in Drews et al. 2022 and WO 2023 / 057392. The breakpoint features on each side are new to this disclosure.In some embodiments, the group of marker features used in steps 18 and 20 includes one or more of the marker features in Drews et al. 2022 (e.g., marker features selected from CX1 to CX17 (optionally excluding CX6)). The etiologies associated with these marker features are provided in Table 2. Full definitions of these marker features are provided in Drews et al. 2022 and WO 2023 / 057392 (see Table 7 of WO 2023 / 057392). In some embodiments, the marker features corresponding to the marker features in WO 2023 / 057392 in the new feature space described herein are obtained by applying a linear combinatorial decomposition using a group of samples with known exposure to the marker features (see Table 7 of WO 2023 / 057392). In other words, the marker feature definition can be determined as a minimization problem under the non-negativity constraint of SbC, given E and PbC. Those. For example, the linear combinatorial decomposition implemented in the R package YAPSA (Hübschmann, D. et al. 2021) can be used for this purpose. The corresponding signature feature definitions according to this disclosure are provided in Table 3. As mentioned above, the process of defining the signature features in the new feature space in Drews et al. 2022 can be performed using one or more or all of CX1 to CX17 (excluding CX6).
[0070] Step 18 (determining a measure of the association between a marker of chromosomal instability and a segment using features quantified for the segment) may include multiplying the probability associated with each of the multiple components of the segment by the corresponding weight in the marker feature and summing the results. As explained above, the probability associated with a component of the segment is typically the probability that the segment with that quantitative feature is drawn from the distribution associated with that component. The resulting sum may be referred to as the “exposure” of the event to the corresponding marker feature. Step 18 may also include normalizing the marker feature exposure of the segment by dividing each exposure by the sum of the exposures of the multiple marker features for that segment. Alternatively, step 18 (determining a measure of the association between a marker of chromosomal instability and a segment using features quantified for the segment) may include obtaining the probability associated with each of the multiple components of the segment and determining the distance between said probabilities and the marker feature weights, optionally where the distance is a cosine distance or an Euclidean distance.
[0071] In step 20, sample level exposure to one or more marker features can be determined. This may include obtaining the sum of probabilities associated with the corresponding component of each of a plurality of segments in the copy number spectrum for each component; and determining the content that satisfies the equation (Equation 1) is a linear combination of one or more characteristic features, where E is a vector of size n containing coefficients. E i PbC is the exposure to marker feature i; PbC is a vector of size c, where each element represents the sum of probabilities associated with the corresponding component; and SbC is a matrix of size c multiplied by n, where each value represents the weight of the component in marker feature i. The resulting exposure can also be normalized per sample so that its sum is 1 (i.e., the exposure for each marker feature in the sample is divided by the sum of the exposures for all marker features in the sample).
[0072] In step 22, the results of steps 18 and optionally 20 are used to determine the presence / absence of a marker feature (and correspondingly, the presence / absence of a process leading to CIN associated with the marker feature). This may include the step of assigning a chromosomal instability marker feature to a segment of one or more segments, the marker feature being the one with the highest metric value for that segment among one or more chromosomal instability marker features. For example, the metric for a marker feature may be the exposure of a segment to the marker feature, and the segment may be associated with the marker feature with the highest exposure for that segment. Alternatively, the metric for a marker feature may be the distance between a vector of probabilities associated with the plurality of components for that segment and the corresponding weights in the marker feature, and the segment may be associated with the marker feature with the lowest metric value for that segment (when the distance is a true distance metric, such as Euclidean distance) or the highest metric value for that segment (when the distance is a similarity metric, such as cosine similarity). Alternatively, when the metric for a marker feature is the probability (exposure) of a segment being associated with a marker feature, an event may be associated with multiple marker features, each with its corresponding probability. As explained above, one or more markers of chromosomal instability may be selected from those defined in Table 3 or obtained by: (i) quantifying a set of copy number features in multiple tumor samples, and (ii) identifying one or more markers that may contribute to the copy number profiles of the multiple tumor samples by nonnegative matrix factorization. Some examples of markers and their associations are provided in Table 2. In some embodiments, a subset of the markers in Table 3 is used. For example, a subset of markers discovered across multiple technologies (i.e., multiple methods for obtaining sequence data suitable for generating copy number profiles) may be used. These may be referred to as “universal markers.” For example, the inventors found that CX1-5, CX8-10, and CX13 exhibited robust signals across technologies (including at least array data such as SNP6 arrays and shallow whole-genome sequencing data sWGS) and classified them as universal markers. Other markers in Table 3 are classified as SNP6-specific markers (CX7, CX11-12, and CX14-17). CX14 is not considered universal because its signal is captured by CX1 in the sWGS data. Therefore, determining the presence or absence of a marker feature is equivalent to determining whether the corresponding CIN process is currently active or has been active in the sample, or in other words, determining whether the copy number change in the copy number spectrum has been at least partially caused by the CIN process.Therefore, the one or more marker features comprise one or more marker features selected from: one or more marker features associated with chromosomal missegregation, one or more marker features associated with impaired homologous recombination, one or more marker features associated with genome-wide duplication tolerance, one or more marker features associated with impaired non-homologous end joining, one or more marker features associated with replication stress, and / or one or more marker features associated with impaired DNA damage perception, wherein the chromosomal missegregation is optionally caused by defective mitosis and / or telomere dysfunction. Furthermore, the presence / absence of a marker feature or corresponding process can be determined at the sample level by comparing the exposure obtained in step 20 with a corresponding predetermined threshold. The predetermined threshold can be obtained through simulation, such that a predetermined proportion (e.g., 95%) of simulated samples have a truly absent marker feature (but the marker feature exposure may still be non-zero due to noise in the simulated data) below this threshold. Examples of thresholds for the marker features in Table 3 are provided in Table 4. Other thresholds may be used depending on the desired stringency in identifying the presence of a marker feature.
[0073] In optional step 24, the results of one or more previous steps (e.g., one or more of the following: tumor copy number profile, segment-quantified characteristics, segment-quantified measures, parameters of the components used, sample exposure, identification of markers / processes associated with the event or present in the sample, etc.), and any information derived therefrom, such as references Figure 2 The explained treatment recommendations or prognosis are provided to the user (e.g., through a user interface) or to the memory or database of a computing system.
[0074] Table 1. Definitions of the 41 mixed components used. Characteristic = one of the five basic characteristics of copy number distortion; Mean = if SD exists, this is the mean of the Gaussian component, otherwise it is the mean of the Poisson component; SD = if present, this is the standard deviation of the Gaussian component; Number = the number of the component used as an identifier.
[0075]
[0076]
[0077] Table 2. Etiology of CIN landmark features in Drews et al. (2022). Sign = landmark feature. Note: CX6 was not used in this work.
[0078]
[0079] Table 3. CIN Flag Features - Component Definition Matrix. CIN flag features are defined in a new feature space consisting of 41 components from 3 copy number features. Component definitions are provided in Table 1 (in the same order as below, i.e., segsize1 = the first listed segsize component, etc.).
[0080]
[0081]
[0082] Table 3 (continued). CIN signature features - component definition matrix. CIN signature features are defined in a new feature space consisting of 41 components from 3 copy number features. Component definitions are provided in Table 1 (in the same order as below, i.e., segment size 1 = the first segment size component listed, etc.).
[0083]
[0084]
[0085] Table 3 (continued). CIN signature features - component definition matrix. CIN signature features are defined in a new feature space consisting of 41 components from 3 copy number features. Component definitions are provided in Table 1 (in the same order as below, i.e., segment size 1 = the first segment size component listed, etc.).
[0086]
[0087]
[0088] application
[0089] The above methods have applications in identifying markers of chromosomal instability associated with a sample and in individual events within the sample, such as characterizing historical or ongoing chromosomal instability and tumor heterogeneity in CIN.
[0090] Generally, characterizing DNA samples from tumors in terms of active chromosomal instability markers in the sample can be used to determine treatments for subjects with cancer, as well as to identify drug targets for treating cancer. All these applications remain feasible using the methods described herein. Therefore, the present invention also provides a method for treating cancer in a subject, wherein the method includes administering or recommending the administration of a specific treatment to the subject, depending on the chromosomal instability markers found to be active in the sample.
[0091] Figure 2A method for providing prognosis, identifying drug targets, determining treatment, or treating a subject according to one embodiment of this disclosure is illustrated. The method may include optional step 30: obtaining a DNA sample from the subject's tumor or from multiple tumor samples. Optionally, a matching normal sample may also be obtained from the subject. This can be used to provide tumor copy number profiles, as is known in the art (see, for example, Van Loo et al., 2010). The step of obtaining a sample from the subject may include physically obtaining the sample from the subject. Alternatively, the sample may be previously obtained and may not require interaction with the subject. In other words, obtaining a DNA sample may include receiving a previously obtained DNA sample. Optionally, sequence data may be obtained from tumor (and optionally matching normal) DNA samples. The step of obtaining sequence data from a DNA sample may include sequencing the DNA sample or analyzing the sample using a genome array. The step of obtaining sequence data from a DNA sample may include performing single-cell or batch DNA sequencing. Alternatively, the sequence data may be previously obtained. The sequence data may include batch DNA sequencing data, single-cell DNA sequencing data, or pseudo-batch sequencing data. Pseudo-batch sequencing data refers to sequencing data obtained by combining single-cell DNA sequencing data from multiple cells in a sample. Therefore, obtaining sequence data may include receiving data from one or more databases or from a user via a user interface. The sequence data may contain sequence data for each of multiple single cells. In step 32, the methods described herein (e.g., as referenced) are used. Figure 1 A) Characterizing the sample. This can be done using batch sequencing data or pseudo-batch sequencing data on the sample as a whole. Alternatively or otherwise, this can be done using corresponding single-cell sequencing data on each of the multiple single cells in the sample individually. Thus, as a result of step 32, one or more copy number events can each be associated with one or more marker features, and / or the sample as a whole or each of one or more cells in the sample can be associated with one or more marker features. In other words, the exposure of individual events and / or the sample to one or more marker features can be determined (in this case, the sample can refer to the entire sample or individual cells within the sample). This, in turn, can be used to determine which marker features are associated with individual events and / or are present in the sample.
[0092] Based on this determination, in step 40, the object can be classified as potentially acquiring a specific event, which has previously been shown to be associated with a specific marker feature identified as being present in the sample. For example, a specific oncogene amplification event may be associated with one or more marker features, and the object can be classified as potentially acquiring that specific oncogene amplification event when the sample's exposure to that one or more marker features is above a predetermined threshold. The predetermined threshold may include multiple corresponding predetermined thresholds for each of the multiple marker features, or a single threshold applied to a metric obtained as a weighted combination of the exposures of multiple marker features (e.g., a single predetermined threshold may be applied to a linear combination of the sample's exposures to multiple marker features, where each exposure may be weighted by a corresponding predetermined weight). Samples with exposures to that one or more marker features above the predetermined threshold can be identified as potentially acquiring that oncogene amplification. Oncogene amplification events may be associated with treatment resistance and / or prognosis. Therefore, this, in turn, can be used to provide treatment recommendations and / or prognosis (steps 34 to 38 and / or 35).
[0093] Based on the determinations made in steps 32 and / or 40, in step 34, subjects may be categorized as likely to respond to or unlikely to respond to a particular treatment process, where the known responder / non-responder status is associated with exposure to one or more marker features. This particular treatment process may target a specific gene. If exposure to a marker feature of chromosomal instability is significantly correlated with the effects of perturbing a specific gene (e.g., through genetic perturbation (e.g., in CRISPR necessity screening, also known as CRISPR knockout screening, or in RNAi necessity screening) or through pharmacological perturbation of that gene (e.g., in drug response screening)) (e.g., cell proliferation, growth inhibition, toxicity, etc.), then the marker feature may be known to be associated with a response to inhibiting that gene. In an optional step 36, a particular treatment process (which may comprise one or more different individual treatments) may be determined based on the results of step 34. For example, subjects identified in step 34 as unlikely to respond to the particular treatment process may be identified as potentially benefiting from a treatment different from that particular treatment process. Alternatively, subjects identified in step 34 as likely to respond to the particular treatment process may be identified as potentially benefiting from a treatment that includes the particular treatment process. In optional step 38, the subject can be treated with the treatment determined in step 36.
[0094] Alternatively or otherwise, the methods described herein can be used to provide a prognosis for subjects. Thus, based on the determinations made in steps 32 and / or 40, subjects can be classified in step 35 as having a good or poor prognosis, where the known prognosis is associated with exposure to one or more biomarker features. For example, if a sample with high or low exposure to one or more biomarker features is associated with different prognoses in a patient cohort, then that biomarker feature is known to be associated with prognosis. In other words, if the expected exposure to a biomarker feature is significantly different between samples with a good prognosis and samples with a poor prognosis in a patient cohort, then that biomarker feature is known to be associated with prognosis. The prognosis may be specific to a particular tumor type, such as ovarian cancer.
[0095] Alternatively or otherwise, the methods described herein can be used to identify drug targets for treating cancer. This may include obtaining tumor copy number profiles of multiple samples and characterizing them using the methods described herein in step 32. In this case, the multiple DNA samples comprise samples obtained from tumors or tumor cell lines, wherein the drug target is already being inhibited and the response to inhibiting the drug target has been quantified. This may also include determining in step 37 whether one or more chromosomal instability markers are associated with a response to inhibiting the drug target, wherein the presence of chromosomal instability markers associated with a response to inhibiting the drug target indicates that the drug target is available for treating cancers in which the marker is active. This may also include an optional step 39: identifying a drug targeting the drug target, for example, through a drug database or using any drug design method known in the art. Alternatively or otherwise, the methods described herein can be used to identify genomic features associated with a specific marker. This may include obtaining tumor copy number profiles of multiple samples and characterizing them using the methods described herein in step 32. Step 37 of these implementations may include determining whether one or more chromosomal instability marker features are associated with one or more colocalized genomic features, such as epigenetic states or chromatin structures. For example, the presence of a copy number event associated with a particular marker feature may be associated with the presence of a particular genomic feature encompassing that copy number event or a region defined by distance relative to that copy number event. Genomic features may be, for example, epigenetic states or chromatin states. For example, the methods described herein can be used to identify a link between chromatin states (e.g., the presence of heterochromatin or euchromatin, a particular type of DNA and / or histone modifications or combinations thereof) and the activity of a particular mutational process associated with a CIN marker feature described herein. This can be used, for example, in the context of drug screening to characterize drugs affecting epigenetic states in terms of potentially associated copy number instability processes, or in the context of characterizing patients or patient groups with cancer (e.g., associating a patient's genomic characteristics with a drug response).
[0096] Alternatively or otherwise, the results of step 32 can be used in step 41 to determine which biomarkers are currently active in the sample and which biomarkers have been historically active in the sample. Specifically, when multiple single cells are identified, biomarkers identified as being associated with events common to multiple single cells in the sample (e.g., events present in a proportion of single cells above a predetermined threshold (e.g., more than one cell)) can be identified as biomarkers that were already active in the sample prior to the clonal expansion event. In contrast, biomarkers identified as being associated with events not common to multiple single cells in the sample (e.g., events present in a proportion of single cells below a predetermined threshold (e.g., in an individual cell)) can be identified as biomarkers currently active in the sample. This, in turn, can be used to determine the treatment response and / or prognosis of the subject, for example, based on known prognoses or treatment responses associated with currently active mutational processes in the sample.
[0097] Any treatment described herein may be used alone or in combination with other treatments. For example, any treatment using a drug may be used in combination with one or more chemotherapy treatments, one or more courses of radiation therapy, and / or one or more surgical interventions. In particular, any treatment described herein may be used in combination with treatments that the subject is identified as likely to respond to.
[0098] system
[0099] Figure 3An embodiment of a system for implementing any method according to this disclosure is shown. The system includes a computing device 1 comprising a processor 101 and a computer-readable storage device 102. In the illustrated embodiment, the computing device 1 also includes a user interface 103, shown as a screen, but may include any other means of conveying information to the user, such as through auditory or visual signals. The computing device 1 is communicatively connected, for example, via a network 6 to a sequence data acquisition device 3 (e.g., a sequencer) and / or one or more databases 2 storing sequence data. The one or more databases may additionally store other types of information usable by the computing device 1, such as reference sequences, parameters, etc. The computing device may be a smartphone, tablet, personal computer, or other computing device. The computing device is configured to implement the methods described herein. In some alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which itself is configured to implement the methods described herein. In such cases, the remote computing device may also be configured to send the results of the method to the computing device. Communication between computing device 1 and remote computing devices can be made via wired or wireless connections, and via local or public networks, such as the public internet or WiFi. Sequence data acquisition device 3 can be wired to computing device 1 and / or database 2, or can communicate wirelessly, for example, via network 6 as shown. The connection between computing device 1 and sequence data acquisition device 3 can be direct or indirect (e.g., via a remote computer). Sequence data acquisition device 3 is configured to acquire sequence data from nucleic acid samples (e.g., genomic DNA samples extracted from cell and / or tissue samples). The sequence data acquisition device is preferably a next-generation sequencer, but can also be an array reader. Sequence data acquisition device 3 can be directly or indirectly connected to one or more databases 2, and sequence data (raw or partially processed) can be stored on the databases 2.
[0100] The following content is presented by way of example and should not be construed as limiting the scope of the claims.
[0101] Example
[0102] Example 1 - Mapping Flag Features to Individual Copy Number Events
[0103] Drews, RM, et al. (2022) previously described a method for evaluating the origin and extent of chromosomal instability at the whole-genome level. This method was used to determine a compendium of 17 copy number markers. As part of this method, copy number variations observed in tumors were embedded into a feature space consisting of five basic copy number features: segment size, copy number difference between adjacent segments, length of oscillating copy number segments, breakpoint count per 10 Mb, and breakpoint count per chromosome arm. These five basic features capture well-known copy number patterns and thus enable differentiation of different CIN types at the whole-genome level.
[0104] However, this feature encoding cannot be directly applied to analyze individual copy number variations in the genome. In the original feature encoding, the three copy number features used to model the distribution of copy number changes in the genome (length of the oscillating copy number segment chain, breakpoint count per 10 Mb, and breakpoint count per chromosome arm) are encoded by genomic region rather than by individual copy number event. Therefore, they cannot be used to map CIN marker features to each individual copy number event in the tumor.
[0105] To overcome this limitation, in this embodiment, the inventors designed a single new feature to replace these three features ( Figure 4 a) This new feature, referred to as the "per-side breakpoint" feature, is described below. This feature, along with the segment size and change point features, can be used to recapitulate the original CIN signature features and facilitates mapping the signature features to individual copy number events. In addition to this new feature, the inventors have made minor modifications to the change point feature to achieve single-event mapping.
[0106] New "breakpoint per side" feature
[0107] In defining this new feature, the inventors focused on encoding that not only captures the key characteristics of the three features it replaces, but also helps to recover the CIN marker features that have high weights on these features. Furthermore, the inventors aimed to maintain the distinction between mutational processes that lead to clustered (i.e., replication stress) and segregated (i.e., mitotic errors) copy number changes. To construct an event-based encoding, it is necessary to start with a specific copy number event in the genome. For ease of interpretation, this event is referred to as the current CN (currentCN).
[0108] To encode clustered copy number variations previously represented by breakpoint counts per 10 Mb and per chromosome arm, the inventors devised a coding scheme that takes a certain window around the current CN and counts the number of breaks occurring on either side. The side with the highest number of breaks is considered the count for that feature. The inventors chose the maximum value to account for cases where the current CN is located at the telomere or centromere. For example, for a segment located within, say, 5 Mb of the telomere or centromere, it is anticipated that the count in the flanking regions near the centromere / telomere will not be adequately represented. The code has a high count if other copy number events are clustered with the current CN, and a 0 if there are no adjacent copy number events. The inventors selected a window of 5 Mb flanking each side of the current CN. While this does not perfectly represent the breakpoint feature per chromosome arm, the inventors infer that it is a good compromise between the two original breakpoint features (breakpoint count per 10 Mb and breakpoint count per chromosome arm). This feature captures more information than the feature that only counts breakpoints every 10Mb (although it also captures information captured by the latter) because if the current CN is large, the count effectively spans a region much larger than 10Mb. For example, if the current CN is 40MB, taking a +5MB flanking region means that the total region being examined is 50MB (close to the size of a chromosome arm).
[0109] Fortunately, this encoding also manages to capture the information encoded by the previous oscillation chain length feature: adjacent copy number events of the same size that are part of the oscillation chain will have the same number of breaks in the 5Mb flanking region, and will therefore lead to overrepresentation of a particular count value in the new feature. Thus, the encoding of the breakpoint feature on each side also captures the oscillation chain.
[0110] As done in Drews et al. 2022, diploid segments were not considered, and therefore breakpoint features were extracted only for non-diploid copy number events. To test the robustness of this new feature encoding, the inventors re-derived the marker features and compared their activity with those obtained by using the original encoding (see below).
[0111] Modification of features at the point of change
[0112] This feature is also included in the feature space of Drews et al. 2022, but the feature encoding has been modified to avoid associating events immediately adjacent to focal amplification with signature features of replication stress. In the original feature encoding, the change point captured the difference in absolute copy number compared to the left neighboring segment, regardless of whether the relative change was positive or negative. Therefore, segments immediately adjacent to focal amplification with low copy numbers were defined as having high copy number changes and were thus assigned to one of the signature features associated with amplification.
[0113] For non-diploid regions, relative change point values are now calculated relative to the adjacent regions (even if the adjacent regions are diploid). When the change point value is equal to or greater than -3, the value is assigned relative to the left adjacent region. If this condition is not met, the change point value is assigned relative to the right adjacent region. Then, for those regions with relative change point values less than -3 (i.e., regions with focal amplification on both sides), the change point value is recalculated by subtracting 2 (the normal copy number) from the copy number of that region. In other words, this subtracts the normal state (as if the region's flanks were normal) from the current copy number, thus avoiding the assignment of regions between two focal amplifications to replication stress marker features. Finally, the relative change point values are converted to absolute change point values. Performing these changes improves the accurate mapping of each event to a specific marker feature. This is less important when evaluating marker features at the whole-genome level, but it improves accuracy at the single-event level.
[0114] As Drews et al. 2022 have done, change point values were not extracted for diploid regions. For non-diploid regions located at chromosome beginnings, change point values were calculated relative to the normal copy number (i.e., subtracted by 2 from the copy number).
[0115] This change has almost no overall impact on the marker activity quantified for TCGA samples using the original feature encoding. Figure 5 (See Example 2 for the TCGA data analysis method). The main difference lies in the merging of mitotic markers CX1 and CX6 (the activities of these markers are highly correlated, and therefore these two markers may actually be the same marker, which was artificially segmented during the NMF process due to limited data input). Therefore, the CX6 marker was excluded from any downstream analysis.
[0116] Absolute copy number characteristics
[0117] As explained in Drews et al. 2022, the inventors made a deliberate decision not to use the absolute copy number of a segment as a feature for decoding chromosomal instability. This choice was motivated by the need to ensure robustness in the context of whole-genome doubling (WGD). Through extensive examination of ovarian copy number markers (Macintyre et al. 2018), the inventors found that even with the same underlying mutational process, including the copy number feature artificially segments the marker based on ploidy status. For example, a single copy loss on a whole-genome duplication (WGD) background would result in a copy number status of 3, while the same deletion on a non-WGD background would result in a copy number status of 1. Therefore, a segment loss could be encoded by two different markers, even though they originate from the same mutational process that produces the deletion pattern. As Drews et al. 2022 explained, the novel feature encoding described herein incorporates only the change point feature, which quantifies the copy number change relative to adjacent segments. By using a simulated genome, the inventors have clearly demonstrated that the novel feature encoding remains resistant to the effects of WGD (see the section "Evaluating Feature Encoding Performance"). Furthermore, the method of assigning CIN marker features to individual copy number events is also robust to WGD effects.
[0118] Feature distribution and modeling
[0119] By applying mixture modeling to the feature distribution, each feature is divided into categories or components (see [link to model]). Figure 4 a). For segment size and change point features, the same mixture components used in Drews et al. 2022 were employed (see Table 1). For the new features (bp window), the inventors estimated the Poisson mixture model using the flexmix package in R (Grün, B. & Leisch, F. 2008) and used the Bayesian Information Criterion (BIC) to determine the number of components. The algorithm was initially run on all data with 1 to 10 components. This only produced 2 components, insufficient to represent the range of observed values ( Figure 6Therefore, given that 96% of the data presented fewer than 6 breakpoints within a 5Mb window on each side of the segment, the inventors decided to manually define 7 components representing 0 to 6 counts (see Table 1). The inventors then applied mixture modeling to the remaining 4% of the data and derived 2 additional components, resulting in a total of 9 components for the breakpoint count characteristics within a 5Mb window on each side of the segment (see Table 1). The λ / mean of each of these distributions is provided in Table 1. The bp window components with λ values of 0 to 6 were manually predetermined. For the last two components, the inventors ran flemix to obtain the λ values using the following parameters: MINCOMP=1, MAXCOMP=10, FITDIST="pois", PRIOR=0.001, NITER=1000, DECISION="BIC", SEED=77777. Note that for events presenting more than 6 breakpoints per 5Mb and for different sample groups, other parameters can be used to identify mixtures of Poisson distributions, resulting in Poisson distributions with slightly different λ values.
[0120] In summary, this yields a feature space defined by 41 mixed components (22 from segment size features, 10 from change point features, and 9 from breakpoint features on each side).
[0121] Event-component probability matrix
[0122] Once the components were derived, the inventors calculated the probability that each copy number event belonged to each component for each feature, resulting in a 1×41 dimensional posterior probability vector. For a given sample, all probability vectors were combined into an event-component probability matrix, which served as the input matrix for mapping each individual copy number event to a specific CIN type.
[0123] Event-signature feature score matrix
[0124] For each marker feature and copy number event, the inventors multiply the component probability vector by the marker feature definition (the marker feature definition in the new feature space, see Table 3—obtained through linear combination decomposition, based on the sample-level input matrix in the new feature space and the marker feature definition in Drews et al. 2022 in the original feature space). Then, the inventors sum the resulting vectors to calculate the marker feature score for each event. In some cases (not used in this embodiment), the resulting event-marker feature matrix is multiplied by the sample-level marker feature activity (see the section "Quantification of Sample-Level Marker Features") to calculate the activity-corrected marker feature score for each copy number event.
[0125] The feature scores for each event are ultimately normalized based on the feature. This is done for each event such that the sum of the feature scores is 1, that is, by dividing the score of each feature for a particular event by the sum of the scores of all features for that event.
[0126] Derivation of new symbolic features
[0127] The inventors aim to define an outline of CIN marker features in a new feature space (see [link to original text]). Figure 4 (b) In this embodiment, the inventors used the marker features from Drews et al. 2022 in the new feature encoding. However, the marker features can also be derived from scratch using the new feature encoding (e.g., using NMF), as described in Drews et al. 2022. To derive the marker feature definition, the inventors used 6,335 TCGA samples with sufficient evidence of chromosomal instability (≥20 CNAs) based on the threshold estimated in Drews et al. 2022. With the activity of the original copy number marker features of all 6,335 samples already known, the inventors used the LCD function of YAPSA (Hübschmann, D. et al. 2021) to derive the marker feature definition based on the new feature space input matrix (a sum-of-posterior matrix of 6,335 patients multiplied by 41 components).
[0128] However, the new feature encoding introduces variations in the input matrix that need to be taken into account when following the procedure. As previously described (see the section "Copy Number Feature Encoding: Modification of Variation Point Features"), the inventors modified the original encoding of the variation point features, which limited their ability to distinguish between the mitotic marker features CX1 and CX6. This makes it impossible to use the original marker feature activity from Drews et al. 2022 as the gold standard activity for deriving event-level marker features.
[0129] After excluding CX6 from the marker feature definition matrix, the final marker feature activities for all 6,335 TCGA samples with detectable chromosomal instability were calculated on the input matrix using YAPSA's linear combination decomposition (LCD) function. This input matrix contains the sum of posterior probabilities for the components of the five original features, but with a novel change-point encoding applied.
[0130] As described in Drews et al. 2022, the inventors calculated and applied a marker-specific threshold to clarify small marker activities. The activity matrix for each sample was then normalized to obtain the new gold standard activity for all TCGA samples. As explained above, a linear combinatorial decomposition was used, employing the gold standard sample-activity matrix to derive the definition of the new marker. The new marker (defined in Table 3) captures signals of different CIN types, which are identical to those in the original outline of copy number markers in Drews et al. 2022. Figure 7 ).
[0131] Quantitative analysis of sample level marker characteristics
[0132] After defining CIN marker features that can be mapped to individual events, the final marker activity for all 6,335 TCGA samples was computed using YAPSA's linear combinatorial decomposition function on the input matrix (sample-component posterior sum matrix). This input matrix was derived by summing the event probability vector for each sample (i.e., the matrix contains a vector for each sample whose activity is the sum of the rows of the event-marker matrix (i.e., the sum of all event-specific activities of a single marker feature) before any adjustments based on the sample-level marker features and normalization by marker features). In fact, because linear decomposition is imprecise, sample-level activities were recomputed based on the marker features defined in a new feature space. For new samples, component vectors (new encodings) were computed, and then activities were estimated using linear value decomposition (e.g., implemented in YAPSA) with the new marker feature definitions (Table 3 or any equivalent marker feature definitions using the new feature encodings).
[0133] Following the same procedure as in Drews et al. 2022, the inventors determined a marker feature specificity threshold for the confidence level of small marker feature activity. In short, the inventors performed 1000 simulations, thereby adding noise to the input components. For marker features with true zero values, non-zero simulation values were fitted to a single Gaussian distribution using Mclust from the mclust R package v5.4.6. The threshold was estimated to be 95% of this distribution, and if the activity was below the marker feature specificity threshold, it was set to zero (Table 4). Finally, the sample-activity matrix for each sample was normalized (thus obtaining the sample-level activity matrix) by dividing by the sum of all marker features for that sample, so that the sum of the marker feature activity vectors was 1.
[0134] Table 4. Thresholds for setting the specificity of marker features to zero activity value
[0135]
[0136] Example 2 - Robustness Analysis of the New CIN Feature Encoding
[0137] In this embodiment, the inventors evaluated the performance of the novel marker feature mapping method described in Embodiment 1, specifically in terms of its ability to capture CIN sample types in simulated data, correctly assign these types of CINs to individual events, and perform well for a variety of techniques and CIN degrees.
[0138] method
[0139] See Example 1.
[0140] Data. All data used in this embodiment, except for the raw PCAWG data (WGS), are publicly available. Access to the TCGA portion of the PCAWG data is available via a data access request through the genotype and phenotype databases (www.cancergenomicscloud.org / controlled-access-data). The TCGA copy number profiles used here are from Drews et al., 2022 (hosted at github.com / VanLoo-lab / ascat). TCGA clinical data are available on the TCGA Research Network: www.cancer.gov / tcga. The PCAWG copy number profiles are from Gerstung et al., 2020; Dentro et al., 2021 (hosted at dcc.icgc.org / releases / PCAWG). Clinical annotations of the PCAWG data are available from the Pan-Cancer Whole Genome Analysis Consortium (2020). Hosted at dcc.icgc.org / releases / PCAWG / .
[0141] Software. Unless otherwise specified, all analyses were performed using the free statistical software R (v>4.0). Functions used outside of standard packages are mentioned and cited throughout the text. The following software packages were used in this example and Example 1: YAPSA v. 1.22.0 (Hübschmann, D. et al. 2021), flexmix v. 2.3-19 (Grün, B. & Leisch, F., 2008).
[0142] Sequence read alignment. Reads were aligned as single ends to the human genome assembly GRCh37 using BWA-MEM (v0.7.17) (Li and Durbin, 2009). Repeated reads were then identified and marked using samtools-markdup (v1.15) (Li et al., 2009).
[0143] Copy number profiles were generated. After alignment, absolute copy number profiles were fitted from sequencing data generated using different high-throughput technologies (the methods applied to each dataset are detailed below). After absolute copy number fitting, samples were rated using a star rating system, as previously described (Macintyre, G. et al. 2018; Madrid, L. et al. 2023). In short, 1 star was given to samples with noisy or flat copy number profiles where the absolute copy number could not be estimated, 2 stars to samples with a fit to the absolute copy number but some noise, and 3 stars to samples with a good fit without noise. Samples with 1 star were discarded from downstream analysis. In addition, samples without an appropriate read count (15 reads per block per chromosome copy) to support the fit were also discarded.
[0144] Absolute copy numbers were fitted from sWGS. The inventors used the QDNAseq R package (Scheinin, I. et al. 2014) to count reads within 30kb blocks (using the resolution thresholds determined by Drews et al. 2022 for accurate call of copy number marker features). Blocks mapped to undefined sequence regions in the centromere and reference genome hg19 were excluded. Read counts were then corrected for sequence mappability and GC content before copy number segmentation. After segmenting the spectrum, absolute copy numbers were inferred based on a combination of purity and ploidy values (0.05 < purity < 1, and 1.5 < ploidy < 8), as follows:
[0145]
[0146] Where purity is the fraction of tumor cells in the sample, r is the read count, and d is a constant proportional to the read depth and the mean absolute copy number (ploidy) of tumor cells in the sample: .
[0147] The optimal mathematical solution is considered to be the one with the lowest root mean squared deviation (RMSD) between the unrounded and rounded copy numbers for all blocks. All solutions are manually checked to confirm or correct for a more suitable fit. For pure cancer models (i.e., cell lines, single cells), a purity range of 0.9 to 1 can be used alternatively to fit the absolute copy number.
[0148] Fitting absolute copy number from single-cell WGS. Although single-cell data is not used in this embodiment, the copy number profile can be inferred from such data, and the methods described herein can be applied in the same manner as those described for WGS. For example, to infer a copy number profile from single-cell DNA sequencing data, a 100kb window can be defined along the genome, and the same methods described in the "Inferring Absolute Copy Number from sWGS" section can be applied.
[0149] Fitting Absolute Copy Numbers from Pseudo-Batch WGS. As mentioned above, single-cell sequencing data can be used to obtain single-cell and pseudo-batch CN profiles. For the pseudo-batch profile, the batch copy number profile can be reconstructed by concatenating reads from all single cells sequenced from a specific sample. The pseudo-batch BAM can then be downsampled to the appropriate read count based on block size (30kb), purity, and ploidy. Absolute copy number fitting can then be performed as described in the "Inferring Absolute Copy Numbers from sWGS" section.
[0150] result
[0151] Evaluate feature encoding performance
[0152] To evaluate the robustness of the novel feature encoding described in Example 1, the inventors used a dataset of 240 simulated genomes generated by Drews et al. 2022. Briefly, the activity of five CIN types was simulated by introducing copy number variations into diploid genomes. These CIN types include: Chromosomal missegregation due to mitotic errors (CHR), Large-scale state transition (LST) due to homologous recombination defects, Focal amplification (ecDNA) due to ecDNA circularization and amplification, Early genome-wide duplication (early WGD) due to cytoplasmic division failure, and Late genome-wide duplication (late WGD) due to nuclear replication. For any given sample, three of the five mutational processes were active. Half of the samples had one dominant marker feature and the other half had two dominant marker features. Each mutational process was designed to be dominant in a specific subset of samples (N=100 for CHR, ecDNA, and LST; N=96 for early WGD; N=24 for late WGD).
[0153] The inventors applied the novel feature encoding described in Example 1 to a simulated genome to generate a sample-component posterior sum matrix. The sample-component posterior sum matrix was then deconvolved into a sample-level activity matrix and a marker feature definition matrix using the NMF package in R (with the Brunet algorithm specification). A marker feature search range of 3 to 6 was applied, and the NMF algorithm was run 1,000 times with different random seeds. The optimal number of marker features was based on the stability achieved in terms of co-phenotypic coefficients, dispersion coefficients, and silhouette coefficients. These evaluation metrics were computed against the definition matrix (foundation), activity matrix (coefficients), and connectivity matrix (consensus). This consensus matrix represents the sample cluster based on its dominant marker features in 1,000 runs. The optimal solution was defined as the minimum number of marker features that achieves stability in co-phenotypic coefficients, dispersion coefficients, and silhouette coefficients; and maximizes sparsity without exceeding the sparsity observed in the random permutation matrix. This process produced a total of 5 marker features (…). Figure 8 ). Figure 8 This figure illustrates a set of metrics commonly used to determine the optimal number of signature features (decomposition rank). Specifically, the figure shows a comparison of the number of signature features (x-axis) across four measures used to determine the optimal number of signature features. Circles and solid lines represent the results of sample runs, while triangles and dashed lines represent the results of 1,000 random permutation matrices (these can be considered zero measures). Here, the basis refers to the signature-component matrix, the coefficients to the patient-signature matrix, and the consensus to the connectivity matrix of patients clustered by their dominant signature features over 1,000 runs. The best fit is the run that shows the lowest target score over 1,000 runs. A value of 5 defines the stationary point for the cophenotypic coefficient, dispersion coefficient, and silhouette coefficient, and is the maximum sparsity achievable on top of the zero model for the basis matrix.
[0154] Next, the inventors evaluated whether the identified marker features effectively captured the simulated CIN type. The inventors observed a strong consistency between the definitions of the identified marker features in the feature space and the expected definitions of the simulated CIN type (mean cosine similarity of 0.66; Table 5). The expected definitions were generated by adding weights to those components reflecting the mutation process (adding 1 to those components defining the simulated CIN type and 0 to those components not defining the simulated CIN type, then normalizing the definition vector to a sum of 1). The marker activity largely corresponded to the simulated distribution among samples, resulting in an overall accuracy of 70.5% (296 / 420) for identifying samples with dominant marker features. Figure 9 The inventors also compared the activity of the identified marker features by using original and new feature codes. This analysis showed that both feature spaces were equally proficient in extracting the CIN signal, as indicated by the average Pearson r of 0.71. Figure 10 Overall, these results demonstrate that the novel feature encoding described in this paper effectively accomplishes both the identification of marker features of mutation processes and the quantification of their corresponding activities.
[0155] Table 5. Similarity between the expected and observed definitions of the identified marker features.
[0156]
[0157] Assignment of indicative features to assess individual events
[0158] The novel method for mapping CIN marker features to individual copy number events described in Example 1 outputs a score reflecting the probability that the event is attributable to each marker feature. This flexibility accommodates cases where copy number events exhibit characteristics similar to more than one marker feature, and also avoids assigning a specific CIN type to a copy number event resulting from an uncertain or unrepresented mutation process. Note that there are several options for mapping marker features to events, including: 1) assigning the probability caused by the marker feature to the event (as described in Example 1), and 2) assigning a specific marker feature to the event by selecting the marker feature with the highest probability. The process of calculating the marker feature probability or score of an event is performed by multiplying the "event-component" vector by the "marker feature-component" vector. These scores may, but do not necessarily, be corrected for by sample-level activity. Another possibility for mapping marker features to events is to assign the marker feature with the most similar definition to the "event-component" vector to the event (e.g., using cosine similarity, Euclidean distance, etc.).
[0159] To evaluate the reliability of the proposed mapping method, the inventors assigned each copy number event to the marker feature with the highest association probability. The accuracy of these assignments was then evaluated by calculating the fractions of events exhibiting features similar to the definition of the assigned marker features. The marker feature with the highest cosine similarity in the comparison of the event-component vector and the marker feature-component vector was considered the most similar. This analysis revealed an overall accuracy of 72.2% (479,565 out of 647,630 events), indicating a significant consistency between the mapped marker features and the features defining individual copy number events. Table 6 summarizes the mapping accuracy for each CIN marker feature.
[0160] Table 6. Scores of events correctly assigned to the CIN signature features with the most similar definitions.
[0161]
[0162]
[0163] The number of individual events required to quantify a sample-level biomarker.
[0164] The inventors conducted an analysis to determine the minimum number of copy number events required within a sample to accurately reproduce its signature components. To achieve this, the number of events in the sample spectrum was progressively increased by random selection. The sample-level signature activity was then quantified by applying a linear combination decomposition function of YAPSA to the input matrix (sample-component posterior sum matrix). In each iteration of this decomposition, the information within the input matrix was expanded by progressively introducing an increasing number of events from 1 to 100. The signature components of the sample (defined by its three dominant signatures) were then evaluated based on the number of events. This procedure was repeated 1000 times, with 1000 TCGA samples randomly selected for each run. Only 15 copy number events were sufficient to correctly decompose >90% of the cases (…). Figure 13 ).
[0165] This procedure was also extended to cover individual biomarkers. For each specific biomarker, the inventors randomly selected copy number events from a subset of 200 TCGA samples exhibiting the highest biomarker activity. The sample-level biomarker activity was then quantified, and the success rate of detecting that biomarker was evaluated. The procedure was performed 100 times for each biomarker. The results showed that an average of 7 copy number events were sufficient to detect individual biomarkers with >90% accuracy (Figure 14).
[0166] Cross-technology sample level marker stability
[0167] CIN marker features were derived using TCGA SNP 6.0 array data, as this represents the largest pan-cancer set of tumors analyzed to date using high-resolution copy number profiling. In Drews et al. (2022), the inventors evaluated the robustness of this method across different high-throughput sequencing technologies by comparing marker feature definitions and activity. As expected, a decrease in activity and similarity scores for marker feature definitions was found between SNP6 and sWGS data (Drews et al., 2022).
[0168] Here, the inventors extended this analysis by testing the ability to capture marker feature outlines in sWGS. For this analysis, the inventors used 478 TCGA samples with detectable levels of CIN, which were also part of the PCAWG project. WGS data from the 478 TCGA / PCAWG samples were downsampled to 15 reads per block per chromosome copy (based on ploidy, purity, and block size) to simulate sWGS. The inventors fitted an absolute copy number spectrum (see the section “Fitting Absolute Copy Numbers from sWGS”), extracted five original copy number features, but used a new change point encoding (see the section “Copy Number Feature Encoding: Change Points” in Example 1), quantified the activity of 16 outline marker features (excluding CX6) using YAPSA, and compared it with the activity of marker features derived from SNP6 using Kendall tau. This examined the impact of changes in the change point features.
[0169] This analysis shows that some marker features deteriorate across technologies. Figure 11 These deteriorations are not due to methodological reasons, but rather to differences in how these techniques measure copy number. With a sufficient number of training samples, all marker features are detectable. Fortunately, this deterioration was observed only in tumor type-enriched marker features; the remaining marker features showed robust signals across techniques. Marker features showing robust signals across techniques were classified as general marker features (CX1-5, CX8-10, and CX13), while other marker features were classified as SNP6-specific marker features (CX7, CX11-12, and CX14-17). CX14 is not considered general because its signal is captured by CX1 in sWGS data. It is important to note here that SNP6-specific marker features can also be detected in sWGS, but the copy number variations captured by this marker feature behave differently in the two techniques (e.g., ...). Figure 10 (As shown in Supplementary Figures 46 and 47 of Drews et al., 2022).
[0170] Stability of sample level marker features across block resolution
[0171] Given that the gold standard block size used in Drews et al. (2022) for quantifying the CIN marker feature outline is 30kb, while the block size used to derive copy number spectra from WES data is 50kb, the inventors explored the robustness of the new marker features at different block resolutions by comparing activity.
[0172] Here, copy number spectra from 478 TCGA / PCAWG samples were segmented using eight different block sizes (30, 40, 50, 60, 70, 80, 90, and 100 kb). Once the absolute copy number spectra were fitted, the inventors quantified the novel biomarker developed in this work for each sample. The cosine similarity between the biomarker activity derived from SNP6 and the biomarker activity obtained at each block resolution was then calculated (Fig. 12a). Additionally, Kendall's tau was used to compare the biomarker activities derived at 30 kb (gold standard) and 50 kb (resolution for WES data) (Fig. 12b). Overall, these comparisons show that the novel biomarker is stable across different block resolutions.
[0173] Example 3 - Application of the New Marker Feature Method
[0174] The novel marker-based approach can be applied in a variety of contexts. In particular, it is beneficial in the context of analyzing which CIN types are active in the tumor genome at the current time, and thus identifying the processes leading to the acquisition of new copy numbers. When analyzing batch sequencing data, one can focus on the “average” tumor genome and dissect ongoing CINs by comparing copy number changes acquired in a sample with previous evolutionary stages (e.g., metastatic vs. primary vs. pre-malignant samples). When examining single cells, individual events within individual cells can be analyzed and compared. Therefore, at the single-cell level, ongoing CIN processes can be inferred by identifying cell-specific events (and thus events caused by processes active in that tumor). However, previous CIN markers do not fully utilize this information because they cannot assign markers to individual events. By assigning markers to individual events, the CIN types active in the tumor at the current time can be detected, and thus the processes leading to the acquisition of new copy numbers can be identified. These “ongoing CIN markers” can be used to predict responses to drugs.
[0175] In this embodiment, the inventors used a novel marker feature method to analyze the CIN types that lead to changes in the number of new copies obtained in different cell line models.
[0176] Examples of CIN types that cause mitotic errors in vitro:
[0177] The inventors downloaded batch whole-genome sequencing data from Watson et al., 2024. In this work, the authors treated normal diploid human telomerase reverse transcriptase (hTERT)-immortified human mammary epithelial cells (hTERT-HMEC) and renal proximal tubular epithelial cells (hTERT-RPTEC) with the spindle assembly checkpoint inhibitor resveratrol to generate aneuploid cell pools with different copy number alterations. Single-cell clones were then obtained, expanded in culture, and batch-sequencing was performed on the clones. The sequencing data were processed as described in Example 2 to generate whole-genome copy number profiles at 100 kb resolution. Clones with fewer than 15 reads per block per chromosome copy were excluded from downstream analysis (i.e., 15 copies per block). Ploidy is calculated using the average copy number of the entire genome as the ploidy value, and requires a minimum of 15 reads. ploidy Number of blocks, where the number of blocks equals the genome length divided by the block length. Figure 15 The whole genome copy number profiles of different clones generated after resveratrol treatment are shown.
[0178] Here, the inventors aim to estimate the mutational processes underlying copy number events induced by resveratrol. To this end, the novel copy number encoding feature space described herein is used to compute an event-component vector for each copy number change observed in the cloned genome. The event-component vector is then multiplied by a marker-component vector (defined as CIN marker features) to estimate the CIN type (event-marker weight, also known as the event-level marker score) leading to each copy number change. At 100kb resolution, the marker feature set is grouped into a total of five CIN marker features, which can be mapped to the marker features extracted at 30kb in Example 1, representing the major mutational processes leading to CIN: CX1, mitotic error; CX8 / CX13, amplification process; CX2 / CX5, impaired homologous recombination; CX3, impaired homologous recombination; and CX9 / CX11, replication stress. Specifically, the above procedure is used, but the marker features are extracted again at 100kb, resulting in five marker features. This means that some of the marker features in Example 1 (extracted at 30kb) are “merged” to create a lower number of marker features at 100kb. The expected result of performing NMF at a lower resolution is that some marker features are “merged” together or “grouped” into a single marker feature. The inventors then mapped the 100kb resolution marker features to the original marker features at 30kb to understand what they actually represented. To map the marker features, the activity of the 100kb resolution marker features and the 30kb resolution marker features were correlated, and the 100kb resolution marker features were identified as CX1, CX8 / CX13, CX2 / CX5, CX3, and CX9 / CX11 based on high Spearman Rho values, corresponding to each of the five 100kb marker features. As expected, CX1 is the marker feature with the highest weight among most copy number events induced by resveratrol (…). Figure 16 This demonstrates that the new marker-based approach can identify mitotic errors as CIN types that lead to the accumulation of copy number alterations in these cell line models.
[0179] Examples of CIN types that lead to impaired homologous recombination in vitro:
[0180] The inventors have created a novel in vitro cell line model of homologous recombination-impaired CIN by individually knocking out BRCA1 and BRCA2 in the normal human immortalized cell line hTERT RPE1 using CRISPR / Cas9. Knockout cells were arrested in the G1 phase using 1 μM palbociclib for 16 hours, and then analyzed using the high-precision imaging platform CellenONE. TMSingle-cell sorting was performed at (www.cellenion.com / products / cellenone / ). Then, the inventors performed single-cell whole-genome sequencing to investigate the copy number alterations induced by the knockout of these two key genes in the homologous recombination (HR) pathway. The aim was to verify that the novel marker-based method described herein can reproduce the induced CIN types that thus produce these new copy number alterations.
[0181] Based on single-cell whole-genome sequencing data, the inventors obtained single-cell copy number profiles using the same procedure as for processing batch sequencing data (see Example 2). A block resolution of 100 kb was used, and cells with at least 15 reads per block per chromosome copy were selected (as described above). Cycling cells and / or over-fragmented copy number profiles were excluded from downstream analysis. Figure 17 and Figure 19 The following are examples of the components from hTERT-RPE1 TP53. - / - BRCA2 - / - and hTERT-RPE1 TP53 - / - BRCA1 - / - Whole genome copy number profiles of different single cells in cell line models were sequenced.
[0182] The inventors then used the new copy number encoding feature space to compute the event-component vector for each copy number alteration observed in the knockout genome. The event-component vector was then multiplied by the marker-component vector (defined as a CIN marker) to estimate the CIN type leading to each copy number alteration (event-marker weight, also known as the event-level marker score). The inventors grouped this marker set into a total of five CIN markers, representing the major mutational processes leading to CIN: CX1, mitotic error; CX8 / CX13, amplification process; CX2 / CX5, impaired homologous recombination; CX3, impaired homologous recombination; and CX9 / CX11, replication stress. As expected, a large number of copy number events exhibited a high weight for one of the two markers associated with impaired homologous recombination (see [link to relevant documentation]). Figure 18 and 20 The results demonstrate that the novel marker-based approach described in this paper identifies mutational processes induced by CRISPR editing methods and thus actively lead to the accumulation of copy number alterations in these cell line models.
[0183] References
[0184] Raine KM, et al. ascatNgs: Identifying Somatically Acquired Copy-Number Alterations from Whole-Genome Sequencing Data. Curr ProtocBioinformatics. 2016 Dec 8;56:15.9.1-15.9.17.
[0185] Van Loo, P. et al. Allele-specific copy number analysis of tumors.Proc. Natl. Acad. Sci. U. S. A. 107, 16910–16915 (2010).
[0186] Langmead, B., Trapnell, C., Pop, M. et al. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol10, R25 (2009).
[0187] Carter, S., Cibulskis, K., Helman, E. et al. Absolute quantificationof somatic DNA alterations in human cancer. Nat Biotechnol 30, 413–421(2012).
[0188] Favero F, Joshi T, Marquard AM, Birkbak NJ, Krzystanek M, Li Q,Szallasi Z, Eklund AC. Sequenza: allele-specific copy number and mutationprofiles from tumor sequencing data. Ann Oncol. 2015 Jan;26(1):64-70.
[0189] Adalsteinsson, V.A., Ha, G., Freeman, S.S. et al. Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatictumors. Nat Commun 8, 1324 (2017).
[0190] Drews, R.M., Hernando, B., Tarabichi, M. et al. A pan-cancercompendium of chromosomal instability. Nature 606, 976–983 (2022).
[0191] Macintyre, G. et al. Copy number signatures and mutational processesin ovarian carcinoma. Nat. Genet. 50, 1262–1270 (2018).
[0192] Christopher Pockrandt and others, GenMap: ultra-fast computation ofgenome mappability, Bioinformatics, Volume 36, Issue 12, June 2020, Pages3687–3692.
[0193] Mehran Karimzadeh and others, Umap and Bismap: quantifying genome andmethylome mappability, Nucleic Acids Research, Volume 46, Issue 20, 16November 2018, Page e120.
[0194] Li H, Durbin R. 2009. Fast and accurate short read alignment withBurrows-Wheeler transform. Bioinformatics 25: 1754–1760.
[0195] Li H. (2013) Aligning sequence reads, clone sequences and assemblycontigs with BWA-MEM. arXiv:1303.3997v2
[0196] Seshan VE, Olshen A (2023). DNAcopy: DNA Copy Number Data Analysis.doi:10.18129 / B9.bioc.DNAcopy, R package version 1.76.0, bioconductor.org / packages / DNAcopy.
[0197] Scheinin, I. et al. DNA copy number analysis of fresh and formalin-fixed specimens by shallow whole-genome sequencing with identification andexclusion of problematic regions in the genome assembly. Genome Res. 24,2022–2032 (2014).
[0198] Madrid, L. et al. Predicting response to cytotoxic chemotherapy.bioRxiv 2023.01.28.525988 (2023) doi:10.1101 / 2023.01.28.525988.
[0199] Alexandrov, L. B. et al. The repertoire of mutational signatures inhuman cancer. Nature 578, 94–101 (2020).
[0200] Grün, B. & Leisch, F. FlexMix Version 2: Finite Mixtures withConcomitant Variables and Varying and Constant Parameters. J. Stat. Softw.028, (2008).
[0201] Gerstung, M. et al. The evolutionary history of 2,658 cancers. Nature578, 122–128 (2020).
[0202] Dentro, S. C. et al. Characterizing genetic intra-tumor heterogeneityacross 2,658 human cancer genomes. Cell 184, 2239–2254.e39 (2021).
[0203] ICGC / TCGA Pan-Cancer Analysis of Whole Genomes Consortium. Pan-canceranalysis of whole genomes. Nature 578, 82–93 (2020).
[0204] Hübschmann, D. et al. Analysis of mutational signatures with yetanother package for signature analysis. Genes Chromosomes Cancer 60, 314–331(2021).
[0205] Li, H. et al. The Sequence Alignment / Map format and SAMtools.Bioinformatics 25, 2078–2079 (2009).
[0206] McGranahan, N., Burrell, RA, Endesfelder, D., Novelli, MR &Swanton, C. Cancer chromosomal instability: therapeutic and diagnostic challenges. EMBO Rep. 13,528–538 (2012).
[0207] Steele, CD et al. Signatures of copy number alterations in humancancer. Nature 606, 984–991 (2022).
[0208] Watson, EV, Lee, J.JK., Gulhan, DC et al. Chromosome evolutionscreens recapitulate tissue-specific tumor aneuploidy patterns. Nat Genet 56,900–912 (2024).
[0209] All references cited in this article are incorporated herein by reference in their entirety and for all purposes, to the same extent that each individual publication, patent, or patent application is specifically and individually indicated to be incorporated herein by reference in its entirety.
[0210] The specific embodiments described herein are provided by way of example and not limitation. Various modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described technology. Any subheadings included herein are for convenience only and should not be construed as limiting the scope of this disclosure in any way. Unless the context otherwise indicates, the description and limitations of the features listed above are not limited to any particular aspect or embodiment of the invention and are equally applicable to all aspects and embodiments described. The methods of any embodiment described herein may be provided as a computer program or as a computer program product or a computer-readable medium carrying a computer program configured to perform the methods described above when run on a computer. Throughout the specification and claims, unless the context clearly indicates otherwise, the following terms shall have the meaning explicitly relevant herein. The phrase “in one embodiment” as used herein does not necessarily refer to the same embodiment, but may refer to the same embodiment. Furthermore, the phrase “in another embodiment” as used herein does not necessarily refer to different embodiments, but may refer to different embodiments. Therefore, multiple embodiments of the invention can be readily combined without departing from the scope or spirit of the invention. It is important to note that, unless the context explicitly indicates otherwise, nouns without quantifiers as used in the specification and appended claims include plural pronouns. A range may be expressed herein as from “about” a particular value and / or to “about” another particular value. When such a range is expressed, another embodiment includes from one particular value and / or to another particular value. Similarly, when a value is expressed as an approximation using the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to numerical values is optional and means, for example, + / - 10%. Throughout the specification (including the appended claims), unless the context requires otherwise, the words “comprising” and “including” and their variations will be understood to imply inclusion of the indicated integer or step or group of integers or steps, but not to exclude any other integer or step or group of integers or steps. Unless the context otherwise indicates, other aspects and embodiments of the invention provide for the foregoing aspects and embodiments in which the terms “consisting of” or “substantially consisting of” replace the terms “comprising / including / containing”. As used herein, “and / or” should be considered as specifically disclosing each of the two indicated features or components, with or without the other. For example, “A and / or B” should be considered as specifically disclosing each of the following cases: (i) A, (ii) B, and (iii) A and B, as if each case were listed separately herein.Features expressed in their particular form, or in the foregoing description, or in the appended claims or drawings, or in terms of the manner in which the disclosed function is performed, or in the method or process for obtaining the disclosed result, may suitably be used alone or in any combination of such features to implement the invention in its various forms.
Claims
1. A computer-implemented method for characterizing chromosomal instability in DNA samples obtained from tumors, the method comprising: Obtain a tumor copy number profile of the sample that identifies one or more genomic segments, wherein each segment is associated with a copy number that is different from the copy number of adjacent segments; For each of one or more segments in the copy number spectrum, a set of copy number features is quantified, wherein the set of copy number features consists of copy number features quantified relative to each individual segment; and For a segment of the one or more segments, a metric indicating the association between one or more chromosomal instability marker features and the segment is determined using features quantified for the segment, wherein the chromosomal instability marker features comprise a set of weights for each of a plurality of components, each component being associated with a copy number feature in the copy number feature set, wherein the weights indicate the probability that a mutation process associated with the marker feature will produce a segment having features defined by the components.
2. The method of claim 1, wherein the copy number feature set comprises or consists of the following: The segment size, copy number variation points, and breakpoint counts in one or two flanking regions of a predetermined length, and / or wherein determining the measure of the association between a marker of chromosomal instability and the segment includes obtaining a value for each of a plurality of components of the segment, wherein each component defines the segment’s characteristics as a predetermined distribution of copy number feature values.
3. The method of any of the preceding claims, wherein determining the measure of the association between a marker of chromosomal instability and a segment comprises determining the probability that a copy number feature of the segment belongs to each of a predetermined distribution associated with a corresponding component of the plurality of components.
4. The method of any of the preceding claims, wherein the copy number feature set of the segment comprises a summary value of the breakpoint count in the upstream and downstream flanking regions of a predetermined length surrounding the segment, optionally wherein the summary value is a maximum value and / or wherein the predetermined length is 2 to 10 Mb, 3 to 8 Mb, or about 5 Mb.
5. The method of any preceding claim, wherein each copy number feature is associated with a plurality of components, and each component uses a corresponding predetermined distribution to define the feature of a copy number event, wherein the predetermined distribution of quasi-continuous copy number features (e.g., segment size, copy number change points) is a set of Gaussian distributions, and / or wherein the predetermined distribution of any counting feature (e.g., breakpoint counts in one or two flanking regions of a predetermined length) is a set of Poisson distributions, optionally, the predetermined distribution associated with the component is a distribution defined by parameters in Table 1 or a corresponding distribution obtained by fitting a mixture model to the values of the set of copy number features obtained from the copy number spectra of a plurality of tumor samples, and / or wherein the predetermined distribution comprises 15 to 25 Gaussian distributions for the segment size, 5 to 15 Gaussian distributions for the copy number change points, and 5 to 15 Poisson distributions for breakpoint counts in one or two flanking regions of the predetermined length.
6. The method of any of the preceding claims, wherein the copy number feature group includes copy number change points, wherein the copy number change point of the current segment is determined as follows: when the copy number change point relative to the upstream segment is at or above a predetermined threshold, it is determined by referring to the upstream segment of the current segment; when the copy number change point relative to the upstream segment is below a predetermined threshold and the copy number change point relative to the downstream segment is at or above a predetermined threshold, it is determined by referring to the downstream segment of the current segment; and when the copy number change point relative to both the upstream and downstream segments is below a predetermined threshold, it is determined by referring to a reference copy value, wherein the reference copy value is optionally 2, and wherein the predetermined threshold is optionally -3.
7. The method of any preceding claim, wherein using features quantified for a segment to determine a measure of the association between a marker of chromosomal instability and said segment comprises: The probability associated with each of the multiple components for the segment is multiplied by the corresponding weight in the label feature, and the resulting values are summed, optionally wherein the probability associated with a component for the segment is the probability that the segment having the quantitative feature is drawn from the distribution associated with the component.
8. The method of claim 10, wherein determining the measure of the association between markers indicating chromosomal instability and segments comprises: For each of the multiple marker features, the probability associated with each of the multiple components for the segment is multiplied by the corresponding weight in the marker feature, and the resulting values are summed to obtain the exposure for the marker feature. The resulting exposure is then normalized by dividing each exposure by the sum of the exposures of the multiple marker features.
9. The method of any one of claims 1 to 8, wherein determining the measure of the association between a marker of chromosomal instability and said segment using features quantified for the segment comprises: Obtain the probability associated with each of the multiple components for the segment, and determine the distance between the probability and the weight of the flag feature, optionally wherein the distance is a cosine distance or an Euclidean distance.
10. The method of any preceding claim, further comprising assigning a chromosomal instability marker feature to a segment of the one or more segments, the marker feature being the highest value of the metric among the one or more chromosomal instability marker features for the segment, and / or wherein the metric for the marker feature is the probability that the segment is associated with the marker feature.
11. The method of any of the preceding claims, wherein the one or more chromosomal instability marker features are selected from those defined in Table 3, or corresponding marker features obtained by quantifying the copy number feature set in a plurality of tumor samples and identifying one or more marker features that may lead to the copy number profile of the plurality of tumor samples by nonnegative matrix factorization, and / or wherein the one or more marker features comprise one or more marker features selected from: one or more marker features associated with chromosomal missegregation, one or more marker features associated with impaired homologous recombination, one or more marker features associated with genome-wide duplication tolerance, one or more marker features associated with impaired non-homologous end joining, one or more marker features associated with replication stress, and / or one or more marker features associated with impaired DNA damage perception, wherein the chromosomal missegregation is optionally caused by defective mitosis and / or telomere dysfunction.
12. The method of any preceding claim, further comprising determining that the number of non-diploid segments contained in the copy number spectrum is higher than a predetermined threshold, and / or wherein the copy number spectrum contains at least 15 non-diploid segments, and / or Obtaining the tumor copy number profile includes determining the tumor copy number profile from sequence data, optionally wherein the sequence data is next-generation sequencing data, and / or wherein determining the tumor copy number profile includes determining the copy number of each of a plurality of genomic blocks of a predetermined size, and / or The method includes obtaining a tumor copy number profile from DNA sequence data containing multiple sequence reads, wherein the tumor copy number profile contains copy number estimates for each of multiple genomic blocks, the genomic blocks containing a number of mapped reads exceeding a predetermined threshold.
13. A method for predicting whether a subject with cancer is likely to respond to treatment, the method comprising: Using the method of any one of claims 1 to 12, a DNA sample obtained from the tumor of the subject is characterized as having at least one segment having a metric associated with one or more predetermined chromosomal instability markers that meets predetermined criteria, and / or characterized as having sample exposure to the markers that meets predetermined criteria, wherein: (a) the predetermined markers include markers associated with a response to the inhibition of a specific gene, the treatment being a treatment targeting the gene, and the subject may respond to the treatment when the sample meets the following: it is characterized as having at least one segment having There are metrics associated with the marker feature that meet predetermined criteria, and / or characterized as sample exposure to the marker feature that meets predetermined criteria; or (b) the treatment is a treatment for resistance to the treatment related to gene amplification, optionally the gene amplification is oncogene amplification, and the subject is likely to have gene amplification and is unlikely to respond to the treatment when the sample meets the following: it is characterized by having at least one segment having metrics associated with the one or more marker features that meet predetermined criteria, and / or characterized as sample exposure to the marker feature that meets predetermined criteria.
14. A method for determining whether a person with cancer is susceptible to gene amplification, the method comprising: Using the method of any one of claims 1 to 12, a DNA sample obtained from the tumor of the object is characterized as having at least one segment having a metric that meets a predetermined criterion associated with one or more predetermined chromosomal instability marker features, and / or characterized as having a sample exposure to the marker features that meets a predetermined criterion. If at least one segment has a metric that meets a predetermined standard and is associated with one or more predetermined chromosomal instability markers, and / or has a sample exposure to the marker that meets a predetermined standard, then the object may have undergone gene amplification. Optionally, the gene mentioned therein is an oncogene.
15. A method for determining whether a specific process leading to chromosomal instability is currently active in a cancer of a subject, the method comprising: A sample obtained from the object is characterized using the method of any one of claims 1 to 12, wherein the tumor copy number profile is obtained from single-cell sequencing data, and the characterization is performed individually using each of a plurality of single-cell tumor copy number profiles; Identify one or more shared copy number events and / or one or more unique copy number events, wherein the shared copy number events are those that exist in a proportion higher than a predetermined threshold in the single-cell tumor copy number spectrum, and the unique copy number events exist in a proportion at or below the predetermined threshold in the single-cell tumor copy number spectrum. Specifically, the process leading to chromosomal instability associated with a marker feature identified as being associated with a shared copy number event may have been active in the tumor prior to at least one clonal expansion in the tumor, and / or the process leading to chromosomal instability associated with a marker feature identified as being associated with a unique copy number event may currently be active in the tumor.