Methods and systems for detecting homologous recombination deficiency in cancer therapy
HRProfiler, a machine learning method using specific genomic features, addresses the limitations of existing HRD detection by accurately predicting HRD in both whole-genome and whole-exome sequencing data, enhancing clinical utility for cancer treatment stratification.
Patent Information
- Application Number
- JP2024573155
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-14
- Filing Date
- 2023-06-14
- Publication Date
- 2025-07-08
AI Technical Summary
Current methods for detecting homologous recombination deficiency (HRD) in cancer are limited by the requirement for whole-genome sequencing data, which is not commonly available in clinical settings, and existing tools for whole-exome sequencing data have reduced accuracy and sensitivity.
A machine learning approach called HRProfiler uses a set of six genomic features to predict HRD in both whole-genome and whole-exome sequencing data, employing a linear kernel support vector machine with L1 regularization to identify specific somatic mutations and genomic patterns.
HRProfiler achieves high accuracy and sensitivity in detecting HRD, outperforming existing methods on whole-exome sequencing data and providing clinical applicability by reliably stratifying patients sensitive to PARP inhibitors and platinum therapy.
Smart Images

Figure 2025521263000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 366,392, filed on June 14, 2022, and incorporates its entire content by reference.
[0002] The present technology relates to a method for generating a homologous recombination feature set, a method for training a prediction model for predicting the presence of homologous recombination deficiency, and a system configured to output a homologous recombination classification. The present technology also relates to a method of administering a cancer therapeutic agent.
Background Art
[0003] The repair of DNA double - strand breaks by homologous recombination (HR) is an essential cellular mechanism for maintaining genomic stability and preventing tumorigenesis. In previous studies, genes important in the HR pathway, such as BRCA1, BRCA2, RAD51, and PALB2, have been elucidated. These genes generally exhibit germline or somatic mutations observed in breast cancer, ovarian cancer, and pancreatic cancer.
[0004] Deficiencies in HR genes inactivate the HR repair pathway, thereby rendering cells unstable with respect to double - strand breaks, and as a result, providing a therapeutic opportunity. In particular, cancer patients in whom incomplete HR repair is likely to occur may be sensitive to poly(ADP - ribose) polymerase (PARP) inhibitors and / or platinum therapy. PARP inhibitors induce double - strand breaks by stalling replication forks during DNA replication, thereby increasing the certainty in alternative repair pathways that are error - prone in homologous recombination deficiency (HRD) mutant cells, accumulating mutations in the cells and ultimately causing apoptosis. Similarly, platinum therapy causes inter - strand cross - links and induces apoptosis triggered by p53 in HRD cells.
[0005] Conventional stratification of homologous recombination deficiency (HRD) patients involves screening for pathogenic germline variants and somatic copy number changes in the HR genes and reference genomic markers. Two commercial HRD companion diagnostic (CDx) tests, Myriad myChoice® CDx and FoundationOne® CDx, are approved by the U.S. Food and Drug Administration (FDA) for patients with ovarian cancer. Both Myriad myChoice® CDx and FoundationOne® CDx determine HRD by combining the status of BRCA1 and BRCA2 to quantify overall genomic instability.
[0006] In addition, at least three academic approaches, SigMA, HRDetect, and CHORD, have been developed to capture HR-deficient cancers by applying machine learning approaches that learn the patterns of somatic mutations found in cancer sequencing data. For example, SigMA was specifically developed to detect SBS3, a signature of mutations that were previously thought to be due to HRD. Unfortunately, SigMA is applicable only to targeted panel and whole-exome sequencing data from only a subset of highly mutated cancers (less than 15% of all breast, ovarian, and pancreatic cancers). HRDetect is a machine learning tool that detects HR-deficient cancers from whole-genome sequencing (WGS) data by leveraging a comprehensive profile of signatures of mutations associated with homologous recombination deficiency. Specifically, HRDetect uses deletions in microhomology reflected by HRD-related substitution signatures SBS3 and SBS8, HRD-related rearrangement signatures RS3 and RS5, and HRD-related indel signatures ID6 and ID8. CHORD is an alternative WGS-based HRD prediction tool that uses the mutation patterns directly observed in cancer genomes. CHORD has performance comparable to HRDetect and is computationally efficient because it does not require deriving signatures of mutations from the observed mutation patterns. Both CHORD and HRDetect outperform SigMA. Since all of these utilize the phenotypic footprints of all deficiencies regardless of the mechanism causing the deficiency, they can serve as good alternatives to conventional screening methods. Furthermore, CHORD and HRDetect capture approximately 50% more responders to PARP inhibitors when compared to companion diagnostic (CDx) tests. However, neither CHORD nor HRDetect is widely used because both require whole-genome sequencing data, which is generally not available in most clinical settings. Notably, CHORD cannot be applied to cancers that are whole-exome sequenced (WES), and the performance of HRDetect on WES data is comparable to unfounded speculation.
[0007] In recent years, whole-exome sequencing of cancer has become very common among multiple cancer centers and external providers that routinely generate WES for clinical decision-making. Therefore, there is a need for new approaches applicable to whole-exome sequencing data that improve accuracy and sensitivity.
[0008] In view of the above, one object of the present disclosure is to provide a highly accurate and sensitive artificial intelligence approach for detecting homologous recombination deficiency applicable to both whole-exome and whole-genome sequencing data.
Summary of the Invention
[0009] The present technology relates to methods, systems, and devices for detecting homologous recombination deficiency in cancer. Accordingly, one object of the present invention is to provide a method for generating a homologous recombination feature set. Another object of the present invention is to provide a method for training a prediction model configured to predict the presence of homologous recombination deficiency in a subject. Another object of the present invention is to provide a method for administering a cancer therapeutic agent to a subject. Yet another object of the present invention is to provide a computer system configured to output a homologous recombination classification of a subject.
[0010] In some aspects, a method for generating a homologous recombination feature set is provided. The method includes (a) receiving sequencing data of a subject and a corresponding homologous recombination classification, and (b) generating a homologous recombination feature set, the homologous recombination feature set including a plurality of genomic features of the sequencing data of the subject and the corresponding homologous recombination classification.
[0011] In other aspects, a method of training a prediction model configured to predict the presence of homologous recombination deficiency in a subject is provided. The method includes: (a) receiving sequencing data of the subject and a corresponding homologous recombination classification; (b) generating a homologous recombination feature set, the homologous recombination feature set including a plurality of genomic features of the sequencing data of the subject and the corresponding homologous recombination classification; and (c) training a prediction model using the homologous recombination feature set to generate a trained prediction model configured to predict the presence of homologous recombination deficiency in the subject.
[0012] In other aspects, a method of administering a cancer therapeutic agent to a subject is provided. The method includes: (a) receiving sequencing data of the subject; (b) determining a homologous recombination classification of the subject as an output of a trained prediction model, the trained prediction model being provided with the sequencing data of the subject as an input and being trained with a homologous recombination feature set; and (c) administering a cancer therapeutic agent to the subject at least according to the homologous recombination classification of the subject.
[0013] In a further aspect, a computer system configured to output a homologous recombination classification of a subject is provided. The computer system includes: (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software, the software including executable instructions that, as a result of execution, cause one or more processors of the computer system to: (i) receive sequencing data of the subject; and (ii) output a homologous recombination classification of the subject as an output of a trained prediction model when the trained prediction model is provided with the sequencing data of the subject as an input, the trained prediction model being trained with a homologous recombination feature set.
[0014] The foregoing paragraphs are provided for general introduction and are not intended to limit the scope of the following claims. The described embodiments will be best understood by reference to the following detailed description in conjunction with the accompanying examples, which will afford further advantages.
[0015] This application contains at least one drawing finished in color. Copies of this application including one or more color drawings will be provided by the Patent Office upon request and payment of the necessary fees.
Brief Description of the Drawings
[0016]
Figure 1a(i)
Figure 1a(ii)
Figure 1a(iii)
Figure 1b(i)
Figure 1b(ii)
Figure 1b(iii)
Figure 1c
Figure 1d
Figure 1e
Figure 2
Figure 2a
Figure 2b
Figure 2c
Figure 2d
Figure 2e
Figure 2f
Figure 2g
Figure 3a
Figure 3b(i)
Figure 3b(ii)
Figure 3b(iii)
Figure 3b(iv)
Figure 3c
Figure 3d(i)
Figure 3d(ii)
Figure 3d(iii)
Figure 4a
Figure 4b
Figure 4c
Figure 4d
Figure 4e(i)
Figure 4e(ii)
Figure 4e(iii)
Figure 5a(i)
Figure 5a(ii)
Figure 5b(i)
Figure 5b(ii)
Figure 6
Figure 7
Figure 8
Figure 9
DETAILED DESCRIPTION OF THE INVENTION
[0017] The present disclosure can be embodied in various forms, but the following description of some embodiments is made on the understanding that the present disclosure should be considered as an exemplification of the invention and is not intended to limit the invention to the specific embodiments illustrated. The headings are provided for convenience only and are not to be construed as limiting the invention in any way. Embodiments illustrated under any heading can be combined with embodiments illustrated under other headings.
[0018] Terms such as "comprise(s)", "include(s)", "have", "has", "contain", and variations thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional acts or structures. The present disclosure also contemplates other embodiments that "comprise", "consist of", and "consist essentially of" the embodiments or elements presented herein, whether or not explicitly recited.
[0019] As used herein, words such as "a" and "an" have the meaning of "one or more".
[0020] As used herein, the term "about" may be used to indicate that a described value is within the range of reasonable expectations when describing a magnitude. For example, a numerical value may have a value such as ±0.1% of the recited value (or range of values), ±1% of the recited value (or range of values), ±2% of the recited value (or range of values), ±5% of the recited value (or range of values), or ±10% of the recited value (or range of values).
[0021] The use of numerical values in the various quantitative values described in this application, unless otherwise explicitly indicated, is described as approximate values as if both the minimum and maximum values within the described range were preceded by the word "about". Although not always explicitly stated, it should be understood that the term "about" precedes every numerical representation. It should be understood that such range formats are used for convenience and brevity and should be interpreted flexibly to include not only the numerically explicit limits of the range but also every individual numerical value or sub-range subsumed within that range as if each numerical value and sub-range were explicitly stated. For example, a ratio in the range of about 1 to about 200 should be understood to include the explicitly enumerated limits of about 1 and about 200, as well as individual ratios such as about 2, about 3, and about 4, and sub-ranges such as about 10 - about 50, about 20 - about 100, etc. Also, although not always explicitly stated, it should be understood that the reagents described herein are merely exemplary and that equivalents of such are known in the art.
[0022] When numerical limits or ranges are described herein, the endpoints are included. Also, all values and sub-ranges within the numerical boundaries or ranges are specifically included as if explicitly written out.
[0023] The terms "subject" and "patient" are used interchangeably. As used herein, they refer to any subject for which a treatment method, including the methods described in this disclosure, is desired. In most embodiments, the subject is a mammal, including but not limited to humans, non-human primates such as chimpanzees, livestock such as cows, horses, pigs, companion animals such as dogs, cats, and rabbits, and experimental subjects such as rodents such as rats, mice, and guinea pigs. In some embodiments, the subject is a human.
[0024] Repair of DNA double-strand breaks by homologous recombination (HR) is an essential cellular mechanism for maintaining genomic stability and preventing tumorigenesis 1In prior studies, genes important in the HR pathway, including BRCA1, BRCA2, RAD51, and PALB2, have been elucidated, and these genes generally exhibit pathogenic germline variants and / or somatic mutations in breast cancer, ovarian cancer, prostate cancer, and pancreatic cancer. 2~5 Defects in HR genes disable the HR repair pathway, thereby rendering cells unstable for double-strand breaks, and as a result, treatment opportunities are provided. Specifically, patients with cancers bearing incomplete HR repair are highly sensitive to both poly(ADP-ribose) polymerase (PARP) inhibitors and / or platinum therapy. 6,7 PARP inhibitors induce double-strand breaks by stalling replication forks during DNA replication, thereby enhancing the certainty in alternative repair pathways that are error-prone in HR-deficient (HRD) cancer cells, accumulating mutations in the cells, and ultimately causing apoptosis. 8 Similarly, platinum therapy causes interstrand crosslinks and induces apoptosis induced by p53 in HRD cells. 9
[0025] The conventional stratification of HRD cancers and HR-proficient (HRP) cancers involves screening for reference genomic markers including pathogenic germline variants and somatic copy number changes in HR genes. 10~12 Currently, there are multiple Clinical Laboratory Improvement Amendments (CLIA)-certified tests and at least two U.S. Food and Drug Administration (FDA)-approved commercial companion diagnostic (CDx) tests available to cancer patients. 13 FDA-approved tests include Myriad myChoice® CDx and FoundationOne® CDx, which determine HRD by combining the status of BRCA1 and BRCA2 to quantify overall genomic instability. 13 For example, Myriad myChoice® CDx measures telomeric allelic imbalance (TAI). 12 , long state transition (LST) events 14 , and loss of heterozygosity (LOH) 10 It relies on the use of a genomic instability score (GIS) or HRD score, which is a composite score of three specific copy number changes, including 11 . HRD score cut-offs of 33 and 63 have been applied to ovarian cancer 15,16 .
[0026] At least three research approaches have also developed machine learning algorithms: HRDetect 17 , CHORD 18 and SigMA 19 to capture HR-deficient cancers by applying them. HRDetect is a machine learning tool that detects HR-deficient cancers from whole-genome sequencing (WGS) data by utilizing a subset of the footprints of mutations associated with homologous recombination deficiency 17 . Specifically, HRDetect uses insertions and deletions in microhomology reflected by HR-related single-base substitution (SBS) footprints 20 SBS3 and SBS8, HRD-related rearrangement footprints 21 RS3 and RS5, and HRD-related indel footprints 22 ID6 and ID8.
[0027] CHORD is an alternative WGS-based HRD prediction tool that does not rely on mutation footprints. Instead, it uses 145 types of mutations directly observed in whole-genome sequenced cancers 18 . CHORD is computationally more efficient and has been shown to have the same performance as one of HRDetect in previous studies 17Both CHORD and HRDetect can serve as good alternatives to conventional screening methods because they utilize the footprint of phenotypic mutations of deficiencies regardless of the mechanism causing the deficiencies. 17,18 Furthermore, prior studies have shown that the predictions from these tools outperform the performance of conventional stratification of HRD patients. 23 However, both CHORD and HRDetect rely on the use of HRD-specific patterns of structural variations that can only be reliably detected from WGS data alone. 17,18 By excluding structural variations, HRDetect can also be applied to whole exome sequencing (WES) data, although it will significantly reduce its performance. 17 Conversely, the implementation of CHORD cannot utilize WES cancer. Both CHORD and HRDetect are limited to clinical use because they require whole genome sequencing data, which is generally not available in most clinical settings.
[0028] In contrast to CHORD and HRDetect, it has been developed to detect HRD from whole genome sequencing data, whole exome sequencing data, and targeted panel sequencing data, and mainly focuses on panel sequencing data. 19 This tool utilizes a machine learning approach to exclusively identify SBS3 and requires a total of at least 5 single nucleotide variants from panel sequencing. 19 MSK-IMPACT data 24 Based on this, since these panel-sequenced samples have at least 5 mutations, this limits the applicability of SigMA to 35% of breast samples and 33% of ovarian samples.
[0029] In principle, two distinct approaches have been utilized to evaluate the performance of methods for detecting HRD. In the original publications, CHORD and HRDetect rely on the concordance between their predictions and previous HRD / HRP annotations based on germline or somatic genomic alterations in HR pathway genes, including BRCA1 and BRCA2. 17、18 This concordance can be quantified by the area under the receiver operating characteristic curve (AUC) using both CHORD and HRDetect, and AUCs exceeding 0.90 have been reported. 17、18 Unfortunately, this type of comparison requires ground truth for HRD and HRP cancers, which in most cases is not straightforward to derive. The second approach relies on comparing clinical endpoints of HRD-predictive and HRP-predictive cancers, including overall progression-free survival and / or disease-free survival in patients treated with either platinum therapy or PARP inhibitors. The advantage of this approach is that it provides immediate clinical relevance. Unfortunately, such comparisons require the availability of well-annotated clinical genomics datasets, which are currently limited, especially at whole-genome resolution.
[0030] In recent years, whole-exome sequencing has begun to be incorporated into clinical oncology workflows. 25, there is a lack of approaches for detecting HRD samples from exome-sequenced cancers. Here, we present a highly accurate and sensitive artificial intelligence approach called Homologous Recombination Proficiency Profiler (HRProfiler) for distinguishing between homologous recombination proficient (HRP) and homologous recombination deficient (HRD) breast and ovarian cancers. HRProfiler utilizes six distinct types of somatic mutations detectable from whole-exome and whole-genome sequencing data. Based on the agreement between tool predictions and previous HRD / HRP annotations, HRProfiler performs at the same level as CHORD, HRDetect, and SigMA in whole-genome sequencing data and outperforms these tools in whole-exome sequencing data. Based on clinical endpoints, HRProfiler outperforms the performance of all existing approaches in detecting patients who respond to platinum therapy. Overall, HRProfiler can use the footprint of exome-derived mutations in failed DNA repair processes to detect clinical biomarkers for reliably stratifying patients sensitive to PARP inhibitors or platinum therapy.
[0031] Exemplary method for generating homologous recombination deficient (HRD) positive and HRD negative feature sets In some embodiments, the disclosure provides a method for generating a homologous recombination feature set. The method includes (a) receiving sequencing data of a subject and a corresponding homologous recombination classification, and (b) generating a homologous recombination feature set, the homologous recombination feature set including a plurality of genomic features of the sequencing data of the subject and the corresponding homologous recombination classification.
[0032] In some embodiments, the sequencing data includes whole genome sequencing data, whole exome sequencing data, portions thereof, or any combination thereof. In some embodiments, the sequencing data includes whole genome and whole exome sequencing data.
[0033] In some embodiments, the homologous recombination feature set includes the total number and proportion of deletions in microhomology features of the sequencing data, the total number and proportion of genomic segments having a heterozygosity loss feature in the sequencing data, the total number and proportion of heterozygous genomic segment features in the sequencing data, the total number and proportion of C:G>T:A single nucleotide substitutions in the 5’-NpCpG-3’ context feature of the sequencing data, the total number and proportion of C:G>G:C single nucleotide substitutions in the 5’-NpCpT-3’ context feature of the sequencing data, or any combination thereof.
[0034] In some embodiments, the total number and proportion of genomic segments having heterozygosity loss includes a size of about 1 to about 40 megabases having at least 1 copy of the genomic segments having heterozygosity loss. For example, the total number and proportion of genomic segments having heterozygosity loss includes a size of about 1 to about 40 megabases, about 2 to about 36 megabases, about 4 to about 32 megabases, about 8 to about 28 megabases, about 12 to about 24 megabases, or about 16 to about 20 megabases having at least 1 copy of the genomic segments having heterozygosity loss.
[0035] In some embodiments, the total number and proportion of heterozygous genomic segments include a size of about 3 to about 40 megabases, or about 10 to about 40 megabases, having 3 to 9 copies of each heterozygous genomic segment of heterozygosity loss. For example, the total number and proportion of heterozygous genomic segments include a size of about 10 to about 40 megabases, about 5 to about 35 megabases, about 7 to about 30 megabases, about 9 to about 25 megabases, about 11 to about 20 megabases, or about 13 to about 15 megabases, having 3 to 9 copies, 4 to 7 copies, or 5 to 6 copies of each heterozygous genomic segment of heterozygous genomic segments.
[0036] In other embodiments, the total number and proportion of heterozygous genomic segments include a size of at least 40 megabases, having 2 to 4 copies of each heterozygous genomic segment of heterozygous genomic segments. For example, the total number and proportion of heterozygous genomic segments include a size of at least 40 megabases, at least 50 megabases, at least 75 megabases, or at least 100 megabases, having 2 to 4 copies, or 3 copies of each heterozygous genomic segment of heterozygous genomic segments.
[0037] In some embodiments, the total number and proportion of deletions in microhomology include a size of at least 5 base pairs. For example, the total number and proportion of deletions in microhomology include a size of at least 5 base pairs, at least 7 base pairs, at least 9 base pairs, at least 15 base pairs, or at least 17 base pairs.
[0038] In some embodiments, the homologous recombination classification includes homologous recombination deficiency positive or homologous recombination deficiency negative.
[0039] In some embodiments, the sequencing data of the subject includes retrospective study sequencing data of patients who participated in a clinical trial, and the patients are the same as or different from the subject.
[0040] Exemplary method for training a homologous recombination deficiency prediction model In other aspects, the present disclosure provides a method of training a prediction model configured to predict the presence of homologous recombination deficiency in a subject. The method includes: (a) receiving sequencing data of the subject and corresponding homologous recombination classification; (b) generating a homologous recombination feature set, the homologous recombination feature set including a plurality of genomic features of the subject's sequencing data and the corresponding homologous recombination classification; (c) training a prediction model using the homologous recombination feature set to generate a trained prediction model configured to predict the presence of homologous recombination deficiency in the subject.
[0041] In some embodiments, the training includes a linear kernel support vector machine (SVM) with L1 regularization.
[0042] In some embodiments, the prediction model includes a random forest prediction model, a naive Bayes classifier prediction model, a support vector machine prediction model, a logistic regression prediction model, or any combination thereof.
[0043] In some embodiments, the prediction model is configured to predict the presence of homologous recombination deficiency in the subject with an accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%. For example, the accuracy rate can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0044] In some embodiments, the prediction model is configured to predict the presence of homologous recombination deficiency in a subject, with a precision rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the precision rate can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0045] In some embodiments, the prediction model is configured to predict the presence of homologous recombination deficiency in a subject, with an F1 of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the F1 can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0046] In some embodiments, the prediction model is configured to predict the presence of homologous recombination deficiency in a patient, with a sensitivity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the sensitivity can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0047] In some embodiments, the prediction model is configured to predict the presence of homologous recombination deficiency in a patient, with a specificity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the specificity can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0048] In some embodiments, the prediction model is configured to predict the presence of homologous recombination deficiency in a patient, with a balanced accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the balanced accuracy rate can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0049] In some embodiments, the sequencing data includes whole genome sequencing data, whole exome sequencing data, a part of these, or any combination thereof.
[0050] In some embodiments, the homologous recombination feature set includes the total number and ratio of deletions in microhomology features of the sequencing data, the total number and ratio of genomic segments with loss of heterozygosity features of the sequencing data, the total number and ratio of heterozygous genomic segment features of the sequencing data, the total number and ratio of C:G>T:A single nucleotide substitutions in the 5'-NpCpG-3' context features of the sequencing data, the total number and ratio of C:G>G:C single nucleotide substitutions in the 5'-NpCpT-3' context features of the sequencing data, or any combination thereof.
[0051] In some embodiments, the total number and ratio of genomic segments with loss of heterozygosity include a size of about 1 to about 40 megabases having at least 1 copy of the genomic segments with loss of heterozygosity. For example, the total number and ratio of genomic segments with loss of heterozygosity include a size of about 1 to about 40 megabases, about 2 to about 36 megabases, about 4 to about 32 megabases, about 8 to about 28 megabases, about 12 to about 24 megabases, or about 16 to about 20 megabases having at least 1 copy of the genomic segments with loss of heterozygosity.
[0052] In some embodiments, the total number and percentage of heterozygous genomic segments include a size of about 3 to about 40 megabases, or about 10 to about 40 megabases, having 3 to 9 copies of each heterozygous genomic segment of loss of heterozygosity. For example, the total number and percentage of heterozygous genomic segments include a size of about 10 to about 40 megabases, about 5 to about 35 megabases, about 7 to about 30 megabases, about 9 to about 25 megabases, about 11 to about 20 megabases, or about 13 to about 15 megabases, having 3 to 9 copies, 4 to 7 copies, or 5 to 6 copies of each heterozygous genomic segment of heterozygous genomic segments.
[0053] In other embodiments, the total number and percentage of heterozygous genomic segments include a size of at least 40 megabases having 2 to 4 copies of each heterozygous genomic segment of heterozygous genomic segments. For example, the total number and percentage of heterozygous genomic segments include a size of at least 40 megabases, at least 50 megabases, at least 75 megabases, or at least 100 megabases having 2 to 4 copies, or 3 copies of each heterozygous genomic segment of heterozygous genomic segments.
[0054] In some embodiments, the total number and percentage of deletions at microhomology include a size of at least 5 base pairs. For example, the total number and percentage of deletions at microhomology include a size of at least 5 base pairs, at least 7 base pairs, at least 9 base pairs, at least 15 base pairs, or at least 17 base pairs.
[0055] In some embodiments, the homologous recombination classification includes homologous recombination deficiency positive or homologous recombination deficiency negative.
[0056] In some embodiments, the sequencing data of the subject includes retrospective clinical trial sequencing data of patients who participated in a clinical trial, and the patients are the same as or different from the subject.
[0057] Exemplary method of stratifying cancer therapeutic administration based on HRD classification In other aspects, the present disclosure provides a method of administering a cancer therapeutic agent to a subject. The method includes (a) receiving sequencing data of the subject, and (b) determining a homologous recombination classification of the subject as an output of a trained prediction model, wherein the trained prediction model is provided with the sequencing data of the subject as an input, and the trained prediction model is trained with a homologous recombination feature set, and (c) administering the cancer therapeutic agent to the subject at least according to the homologous recombination classification of the subject.
[0058] In some embodiments, the cancer therapeutic agent includes at least one selected from the group consisting of platinum therapy and poly(ADP-ribose) polymerase (PARP) inhibitors. Examples of PARP inhibitors include, but are not limited to, veliparib, pamiparib, talazoparib, olaparib, niraparib, rucaparib, iniparib, and 3-aminobenzamide. In some embodiments, the PARP inhibitor includes talazoparib, olaparib, niraparib, rucaparib, veliparib, or any combination thereof. Examples of platinum therapy include, but are not limited to, cisplatin, oxaliplatin, carboplatin, and nedaplatin. In some embodiments, the platinum therapy includes cisplatin, oxaliplatin, carboplatin, or any combination thereof.
[0059] In some embodiments, the cancer therapeutic agent causes double-strand breaks in genomic molecules of the subject's cells and causes apoptosis induced by p53.
[0060] The subject can be any subject who already has cancer, a subject who has not yet experienced or presented cancer symptoms, or a subject who is susceptible to cancer. In some embodiments, the subject is a human who is susceptible to cancer, such as a human having a family history of cancer. For example, (i) having a certain inherited gene (e.g., mutant BRCA1 and / or mutant BRCA2), (ii) having taken estrogen alone (without progesterone) for several years (at least 5 years, at least 7 years, or at least 10 years) after menopause, and / or (iii) having taken the ovulation inducer clomiphene citrate, a female has a higher risk of developing breast cancer.
[0061] In some embodiments, the subject is suspected of having cancer. Examples of cancer include, but are not limited to, bone cancer, testicular cancer, gastric cancer, sarcoma, lymphoma, Hodgkin's lymphoma, lymphoma, head and neck cancer, squamous cell head and neck cancer, thymic cancer, epithelial cancer, salivary gland cancer, liver cancer, stomach cancer, thyroid cancer, lung cancer, ovarian cancer, breast cancer, prostate cancer, esophageal cancer, pancreatic cancer, glioma, leukemia, multiple myeloma, renal cell cancer, bladder cancer, cervical cancer, choriocarcinoma, colorectal cancer, oral cancer, skin cancer, and melanoma.
[0062] In some embodiments, the cancer is at least one selected from the group consisting of breast cancer, ovarian cancer, prostate cancer, pancreatic cancer, and sarcoma.
[0063] In some embodiments, the trained prediction model includes a prediction model trained by a linear kernel support vector machine (SVM) with L1 regularization.
[0064] In some embodiments, the trained prediction model is configured to determine the homologous recombination classification of a subject with an accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the accuracy rate can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0065] In some embodiments, the trained prediction model is configured to determine the homologous recombination classification of a subject with a precision rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the precision rate can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0066] In some embodiments, the trained prediction model is configured to determine the homologous recombination classification of a subject with an F1 of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the F1 can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0067] In some embodiments, the trained prediction model is configured to determine the homologous recombination classification of a subject with a sensitivity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the sensitivity can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0068] In some embodiments, the trained prediction model is configured to determine the homologous recombination classification of a subject with a specificity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the specificity can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0069] In some embodiments, the trained prediction model is configured to determine the homologous recombination classification of a subject with a balanced accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the balanced accuracy rate can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0070] In some embodiments, the sequencing data includes whole genome sequencing data, whole exome sequencing data, a portion of these, or any combination thereof. In some embodiments, the sequencing data includes whole genome and whole exome sequencing.
[0071] In some embodiments, the homologous recombination feature set includes genomic features.
[0072] In some embodiments, the genomic features include the total number and proportion of deletions in the microhomology features of the sequencing data, the total number and proportion of genomic segments with heterozygosity loss features in the sequencing data, the total number and proportion of heterozygous genomic segment features in the sequencing data, the total number and proportion of C:G>T:A single nucleotide substitutions in the 5'-NpCpG-3' context features of the sequencing data, the total number and proportion of C:G>G:C single nucleotide substitutions in the 5'-NpCpT-3' context features of the sequencing data, or any combination thereof.
[0073] In some embodiments, the total number and percentage of genomic segments having loss of heterozygosity include sizes of about 1 to about 40 megabases that include at least one copy of the genomic segments having loss of heterozygosity. For example, the total number and percentage of genomic segments having loss of heterozygosity include sizes of about 1 to about 40 megabases, about 2 to about 36 megabases, about 4 to about 32 megabases, about 8 to about 28 megabases, about 12 to about 24 megabases, or about 16 to about 20 megabases that include at least one copy of the genomic segments having loss of heterozygosity.
[0074] In some embodiments, the total number and percentage of heterozygous genomic segments include sizes of about 3 to about 40 megabases or about 10 to about 40 megabases that have 3 to 9 copies of each heterozygous genomic segment having loss of heterozygosity. For example, the total number and percentage of heterozygous genomic segments include sizes of about 10 to about 40 megabases, about 5 to about 35 megabases, about 7 to about 30 megabases, about 9 to about 25 megabases, about 11 to about 20 megabases, or about 13 to about 15 megabases that have 3 to 9 copies, 4 to 7 copies, or 5 to 6 copies of each heterozygous genomic segment.
[0075] In other embodiments, the total number and percentage of heterozygous genomic segments include sizes of at least 40 megabases that have 2 to 4 copies of each heterozygous genomic segment. For example, the total number and percentage of heterozygous genomic segments include sizes of at least 40 megabases, at least 50 megabases, at least 75 megabases, or at least 100 megabases that have 2 to 4 copies or 3 copies of each heterozygous genomic segment.
[0076] In some embodiments, the total number and percentage of deletions in microhomology include a size of at least 5 base pairs. For example, the total number and percentage of deletions in microhomology include a size of at least 5 base pairs, at least 7 base pairs, at least 9 base pairs, at least 15 base pairs, or at least 17 base pairs.
[0077] In some embodiments, the homologous recombination classification includes homologous recombination deficiency positive or homologous recombination deficiency negative.
[0078] In some embodiments, the sequencing data of the subject includes retrospective clinical trial sequencing data of patients who participated in a clinical trial, and the patients are the same as or different from the subject.
[0079] An exemplary system configured to determine the HRD classification of a subject using a trained HRD model In a further aspect, the present disclosure provides a computer system configured to output a homologous recombination classification of a subject. The computer system includes (a) one or more processors and (b) a non-transitory computer-readable storage medium including software, the software, as a result of execution, causing one or more processors of the computer system to (i) receive sequencing data of the subject and (ii) output, as an output of the trained prediction model, the homologous recombination classification of the subject when the trained prediction model is provided with the sequencing data of the subject as an input, wherein the trained prediction model is trained with a homologous recombination feature set.
[0080] In some embodiments, the software includes determining a cancer therapeutic agent at least according to the homologous recombination classification of the subject.
[0081] In some embodiments, the cancer therapeutic agent comprises at least one selected from the group consisting of platinum therapy and poly(ADP-ribose) polymerase (PARP) inhibitors. In some embodiments, the PARP inhibitor comprises talazoparib, olaparib, niraparib, rucaparib, veliparib, or any combination thereof. In some embodiments, the platinum therapy comprises cisplatin, oxaliplatin, carboplatin, or any combination thereof.
[0082] In some embodiments, the cancer therapeutic agent causes double-strand breaks in the genomic molecules of the target cells and causes apoptosis induced by p53.
[0083] The subject can be any subject who already has cancer, a subject who has not yet experienced or presented symptoms of cancer, or a subject who is susceptible to cancer. In some embodiments, the subject is a human who is susceptible to cancer, such as a human with a family history of cancer. For example, (i) having a specific inherited gene (e.g., mutated BRCA1 and / or mutated BRCA2), (ii) taking estrogen alone (without progesterone) during the postmenopausal period for several years (at least 5 years, at least 7 years, or at least 10 years), and / or (iii) taking the ovulation inducer clomiphene citrate, a woman has a higher risk of developing breast cancer.
[0084] In some embodiments, the subject is suspected of having cancer. Examples of cancer include, but are not limited to, bone cancer, testicular cancer, gastric cancer, sarcoma, lymphoma, Hodgkin lymphoma, lymphoma, head and neck cancer, squamous cell head and neck cancer, thymic cancer, epithelial cancer, salivary gland cancer, liver cancer, stomach cancer, thyroid cancer, lung cancer, ovarian cancer, breast cancer, prostate cancer, esophageal cancer, pancreatic cancer, glioma, leukemia, multiple myeloma, renal cell cancer, bladder cancer, cervical cancer, choriocarcinoma, colorectal cancer, oral cancer, skin cancer, and melanoma.
[0085] In some embodiments, the cancer is at least one selected from the group consisting of breast cancer, ovarian cancer, prostate cancer, pancreatic cancer, and sarcoma.
[0086] In some embodiments, the trained prediction model includes a prediction model trained by a linear kernel support vector machine (SVM) with L1 regularization.
[0087] In some embodiments, the sequencing data includes whole-genome sequencing data, whole-exome sequencing data, a portion thereof, or any combination thereof. In some embodiments, the sequencing data includes whole-genome and whole-exome sequencing.
[0088] In some embodiments, the homologous recombination feature set includes genomic features.
[0089] In some embodiments, the genomic features include the total number and percentage of deletions in the microhomology features of the sequencing data, the total number and percentage of genomic segments having heterozygosity loss features in the sequencing data, the total number and percentage of heterozygous genomic segment features in the sequencing data, the total number and percentage of C:G>T:A single nucleotide substitutions in the 5'-NpCpG-3' context features of the sequencing data, the total number and percentage of C:G>G:C single nucleotide substitutions in the 5'-NpCpT-3' context features of the sequencing data, or any combination thereof.
[0090] In some embodiments, the total number and percentage of genomic segments having heterozygosity loss include a size of about 1 to about 40 megabases having at least one copy of the genomic segments having heterozygosity loss. For example, the total number and percentage of genomic segments having heterozygosity loss include a size of about 1 to about 40 megabases, about 2 to about 36 megabases, about 4 to about 32 megabases, about 8 to about 28 megabases, about 12 to about 24 megabases, or about 16 to about 20 megabases having at least one copy of the genomic segments having heterozygosity loss.
[0091] In some embodiments, the total number and proportion of heterozygous genomic segments include a size of about 3 to about 40 megabases, or about 10 to about 40 megabases, having 3 to 9 copies of each heterozygous genomic segment of the heterozygous genomic segments. For example, the total number and proportion of heterozygous genomic segments include a size of about 10 to about 40 megabases, about 5 to about 35 megabases, about 7 to about 30 megabases, about 9 to about 25 megabases, about 11 to about 20 megabases or about 13 to about 15 megabases, having 3 to 9 copies, 4 to 7 copies, or 5 to 6 copies of each heterozygous genomic segment of the heterozygous genomic segments.
[0092] In other embodiments, the total number and proportion of heterozygous genomic segments include a size of at least 40 megabases, having 2 to 4 copies of each heterozygous genomic segment of the heterozygous genomic segments. For example, the total number and proportion of heterozygous genomic segments include a size of at least 40 megabases, at least 50 megabases, at least 75 megabases or at least 100 megabases, having 2 to 4 copies, or 3 copies of each heterozygous genomic segment of the heterozygous genomic segments.
[0093] In some embodiments, the total number and proportion of deletions at microhomology include a size of at least 5 base pairs. For example, the total number and proportion of deletions at microhomology include a size of at least 5 base pairs, at least 7 base pairs, at least 9 base pairs, at least 15 base pairs or at least 17 base pairs.
[0094] In some embodiments, the homologous recombination classification includes homologous recombination deficiency positive or homologous recombination deficiency negative.
[0095] In some embodiments, the trained prediction model is configured to output a homologous recombination classification of an object with an accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the accuracy rate can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0096] In some embodiments, the trained prediction model is configured to output a homologous recombination classification of an object with a precision rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the precision rate can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0097] In some embodiments, the trained prediction model is configured to output a homologous recombination classification of an object with an F1 of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the F1 can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0098] In some embodiments, the trained prediction model is configured to output a homologous recombination classification of an object with a sensitivity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. For example, the sensitivity can be in the range of about 60% to about 100%, about 65% to about 99%, about 70% to about 95%, about 75% to about 90%, about 80% to about 85%.
[0099] In some embodiments, the trained prediction model is configured to output a homologous recombination classification for a subject, with at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% specificity. For example, the specificity can range from about 60% to about 100%, from about 65% to about 99%, from about 70% to about 95%, from about 75% to about 90%, from about 80% to about 85%.
[0100] In some embodiments, the trained prediction model is configured to output a homologous recombination classification for a subject, with at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% balanced accuracy. For example, the balanced accuracy can range from about 60% to about 100%, from about 65% to about 99%, from about 70% to about 95%, from about 75% to about 90%, from about 80% to about 85%.
[0101] In some embodiments, the sequencing data of the subject includes retrospective study sequencing data of patients who participated in a clinical trial, and the patients are the same as or different from the subject.
[0102] Obviously, in light of the above teachings, many modifications and variations of the present invention are possible. Therefore, within the scope of the appended claims, it should be understood that the present invention can be practiced in ways other than those specifically described herein.
[0103] The following examples are intended to further illustrate protocols for creating, characterizing, and using composite forms of the present disclosure, and are not intended to limit the scope of the claims.
Example
[0104] Example 1. HRD genomic biomarker Repair of DNA double-strand breaks by homologous recombination (HR) is an essential cellular mechanism for maintaining genomic stability and preventing tumorigenesis. In previous studies, important genes in the HR pathway, including BRCA1, BRCA2, RAD51, and PALB2, have been elucidated. These genes commonly exhibit germline or somatic mutations in breast cancer, ovarian cancer, and pancreatic cancer. Defects in HR genes inactivate the HR repair pathway, thereby rendering cells unstable to double-strand breaks, which in turn provides a therapeutic opportunity. In particular, patients with cancers harboring incomplete HR repair are highly sensitive to both poly(ADP-ribose) polymerase (PARP) inhibitors and / or platinum therapy. PARP inhibitors induce double-strand breaks by stalling replication forks during DNA replication, thereby enhancing the certainty of alternative repair pathways that are error-prone in HR-deficient (HRD) cells, accumulating mutations in cells and ultimately causing apoptosis. Similarly, platinum therapy causes interstrand crosslinks and induces apoptosis triggered by p53 in HRD cells.
[0105] Conventional stratification of HRD patients involves screening for pathogenic germline variants and somatic copy number changes in genomic markers, including in the HR gene. Two commercial HRD companion diagnostic (CDx) tests have been approved by the US Food and Drug Administration for patients with ovarian cancer. Both Myriad myChoice® CDx and FoundationOne® CDx determine HRD by combining the status of BRCA1 and BRCA2 to quantify overall genomic instability. At least three academic approaches, namely SigMA, HRDetect, and CHORD, have also been developed to capture HR-deficient cancers by applying machine learning approaches that learn the patterns of somatic mutations found in cancer sequencing data. SigMA was specifically developed to detect SBS3, a footprint of mutations of single base substitution (SBS) previously attributed to HRD, from targeted panel and whole-exome sequencing data. Unfortunately, SigMA is applicable only to targeted panel and whole-exome sequencing data from only a minority (less than 15% of all breast, ovarian, and pancreatic cancers) of highly mutated cancers. HRDetect is a machine learning tool that detects HR-deficient cancers from whole-genome sequencing (WGS) data by leveraging a comprehensive profile of footprints of mutations associated with homologous recombination deficiency. Specifically, HRDetect uses indels in microhomology reflected by HRD-related substitution footprints SBS3 and SBS8, HRD-related rearrangement footprints RS3 and RS5, and HRD-related indel footprints ID6 and ID8. CHORD is an alternative WGS-based HRD prediction tool that uses the mutation patterns directly observed in cancer genomes. CHORD is computationally more efficient as it does not require deriving footprints of mutations from the observed mutation patterns and has performance comparable to HRDetect. Both CHORD and HRDetect outperform SigMA and can serve as good alternatives to conventional screening methods as they leverage the mutational footprint of the loss-of-function phenotype regardless of the mechanism causing the loss.Furthermore, CHORD and HRDetect capture approximately 50% more responders to PARP inhibitors when compared to companion diagnostic (CDx) tests. However, CHORD and HRDetect are limited to clinical use only because they require whole-genome sequencing data, which are generally not available at most clinical sites. Importantly, CHORD cannot be applied to cancers that are whole-exome sequenced (WES), while the performance of HRDetect on WES data is comparable to unfounded speculation. In recent years, whole-exome sequencing of cancer has become very common among multiple cancer centers and external providers that routinely generate WES for clinical decision-making.
[0106] The present disclosure presents a highly accurate and sensitive artificial intelligence approach for detecting homologous recombination deficiency applicable to both whole exome and whole genome sequencing data. The approach disclosed herein includes (i) the total number and percentage of deletions spanning at least 5 base pairs (bp) with microhomology, (ii) the total number and percentage of genomic segments with loss of heterozygosity (LOH) having a size of 1 to 40 megabases, (iii) the total number and percentage of heterozygous genomic segments having a total copy number (TCN) of 3 to 9 and a size of 10 to 40 megabases, (iv) the total number and percentage of heterozygous genomic segments having a TCN of 2 to 4 and a size exceeding 40 megabases, (v) the total number and percentage of C:G>T:A single nucleotide substitutions in the 5’-Np[C]pG-3’ context (the mutated base is within brackets [ ], and N reflects any base 5’ of the mutated cytosine), and (vi) the total number and percentage of C:G>G:C single nucleotide substitutions in the 5’-Np[C]pT-3’ context. By applying a linear kernel support vector machine (SVM) with L1 regularization to these features, an AI approach for predicting homologous recombination deficiency was trained. The training and prediction of the model were applied to both whole genome sequencing data and whole exome sequencing data. The trained model outperforms the performance of SigMA, CHORD, and HRDetect for both whole genome sequencing data and whole exome sequencing data. Notably, the trained model provides the same resolution for detecting homologous recombination deficiency from whole exome-sequenced samples, making it immediately applicable in the clinical setting. Overall, the developed AI approach fills the gap in using the molecular phenotypic footprint of a failed DNA repair process as a clinical biomarker to reliably stratify patients sensitive to PARP inhibitors and / or platinum therapies.
[0107] Example 2. Generation, training, and application of the HRD model The following features were used: (i) the total number and proportion of deletions spanning at least 5 base pairs (bp) by microhomology, (ii) the total number and proportion of genomic segments with loss of heterozygosity (LOH) having a size of 1 to 40 megabases, (iii) the total number and proportion of heterozygous genomic segments with a total copy number (TCN) of 3 to 9 and a size of 10 to 40 megabases, (iv) the total number and proportion of heterozygous genomic segments with a TCN of 2 to 4 and a size exceeding 40 megabases, (v) the total number and proportion of C:G>T:A single nucleotide substitutions in the 5’-Np[C]pG-3’ context (the mutated base is in brackets [ ], and N reflects any base 5’ of the mutated cytosine), and (vi) the total number and proportion of C:G>G:C single nucleotide substitutions in the 5’-Np[C]pT-3’ context, using a minimal set of six genomic features.
[0108] An AI approach for predicting homologous recombination deficiency was trained by applying a linear kernel support vector machine (SVM) with L1 regularization to the above features. The training and prediction of the model were applied to both whole-genome sequencing data and whole-exome sequencing data. The trained model outperformed the performance of SigMA, CHORD, and HRDetect for whole-genome sequencing data and whole-exome sequencing data.
[0109] Notably, the trained model provides the same resolution for detecting homologous recombination deficiency from whole-exome sequenced samples, demonstrating that it is immediately applicable in the clinical setting. Overall, the developed AI approach was continued when using the molecular phenotypic footprint of the failed DNA repair process as a clinical biomarker to reliably stratify patients sensitive to PARP inhibitors and / or platinum therapy.
[0110] The approach described herein is readily applicable to any exome sequencing data. In essence, the present invention enables the detection of HRD status from such sequencing data and can be applied to identify favorable treatments for multiple cancer types, including but not limited to breast cancer, ovarian cancer, pancreatic cancer, prostate cancer, and sarcoma. Potential commercial uses of the present invention include, for example, precision oncology, such as the identification of cancer patients who respond to platinum therapy and / or PARP therapy.
[0111] Example 3. Feature engineering of variants enriched in HRD samples To determine the genomic footprint of homologous recombination deficiency (HRD) across patients profiled using WGS and WES, single nucleotide substitutions (SBS) 26 , insertions and deletions (ID) 27 and copy number alterations (CN) 28 were identified as significantly enriched variants specific to them. In particular, the types of somatic mutations enriched in either HRD cancers or HR proficient (HRP) cancers were compared using a previously developed scheme for classifying SBS, DBS, and CN 27、29 . The comparison was performed for whole-genome sequenced breast cancers using a subset of the Sanger Institute's 560 breast cancer genome cohort 21 (Sanger-WGS-Breast; Figures 1a(i)-(iii)) and for whole-exome sequenced breast cancers using a subset of the TCGA breast cancer cohort 30 (TCGA-WES-Breast; Figures 1b(i)-(iii)). For feature engineering and training purposes, patients were classified as HRD based on at least a 42 HRD score or based on the presence of pathogenic germline variants, somatic mutations, or methylation of BRCA1 and BRCA2 (Figures 5a(i)-(ii)).
[0112] At the SBS resolution, significant enrichment of C:G>T:A single-base substitutions was observed in the 5’-Np[C]pG-3’ context (mutated base in brackets [ ], N reflects any base 5’ of the mutated cytosine) in the HRP samples (Figure 1a(i)–(iii) and Figure 1b(i)–(iii)). This suggested that a relatively large proportion of the mutations in the HRP samples were C:G>T:A transitions at CpG sites when compared with the HRD samples. Conversely, the HRD samples were enriched for C:G>G:C single-base substitutions in the 5’-Np[C]pT-3’ context. At the indel resolution, enrichment of deletions was observed over at least 5 base pairs (bp) with adjacent microhomology sequences across the HRD samples (Figure 1a(i)–(iii) and Figure 1b(i)–(iii)). These mutations could arise from incorrect activation of the microhomology-mediated end joining (MMEJ) or single-strand annealing (SSA) DNA repair pathways in the absence of a functional HR pathway 31 At the copy number resolution, loss of heterozygosity (LOH) events spanning 1–40 Mb and heterozygosity events spanning 10–40 Mb with total copy number (TCN) states of 3–9 were enriched in the HRD samples (Figure 1a(i)–(iii) and Figure 1b(i)–(iii)). In contrast, very large (over 40 Mb) heterozygous segments with TCNs of 2–4 were enriched in the HRP samples (Figure 1a(i)–(iii) and Figure 1b(i)–(iii)). This finding suggests that very large diploid segments or regions that have undergone genome doubling were enriched in the HRP samples, consistent with the finding that the HRP samples were genomically clearly stable, had relatively low copy number aberrations, and thus had a low HRD score compared with the HRD samples 32 .
[0113] Based on these findings, significant mutation channels (methods) were combined with the following six genomic features: (i) the total number and percentage of deletions spanning at least 5 bp of microhomology (abbreviated as DEL.5.MH), (ii) the total number and percentage of genomic segments with loss of heterozygosity (LOH) having a size of 1 to 40 megabases (LOH: 1-40Mb), (iii) the total number and percentage of heterozygous genomic segments with a TCN of 3 to 9 and a size of 10 to 40 megabases (3-9: HET: 10-40Mb), (iv) the total number and percentage of heterozygous genomic segments with a TCN of 2 to 4 and a size greater than 40 megabases (2-4: Het: >40Mb), (v) the total number and percentage of C:G>T:A single-base substitutions in the 5’-NpCpT-3’ context (N[C>T]G), and (vi) the total number and percentage of C:G>G:C single-base substitutions in the 5’-Np[C]pT-3’ context (N[C>G]T). To determine whether these genomic features could accurately separate HRD samples and HRP samples, principal component analysis (PCA) was performed. This showed that these six features could distinguish HRD from HRP samples across two principal components for both WGS (Figure 1c) and WES (Figure 1d) samples.
[0114] Next, using the TCGA-WES-Breast cohort, the associations of the six genomic features were compared with previously developed HRD annotations, including (i) germline or somatic changes in BRCA1 / 2, (ii) different thresholds for HRD scores already reported in the literature 11,15,16 , (iii) the trace of copy number HRD CN17 29 , (iv) the trace SBS3 based on COSMIC attributes 27 , and (v) the trace SBS3 based on SigMA attributes 19 (Figure 1e). In all cases, the six genomic features were highly associated across major HRD annotations that were enriched for N[C>T]G and 2-4: Het: >40Mb in HRP samples, with other features enriched in HRD samples (Figure 1e).
[0115] Example 4. Training model for detecting HRD from WGS and WES breast cancers To determine whether defined genomic features can accurately predict HRD status at WGS resolution, a machine learning model called HRProfiler was trained based on a linear kernel support vector machine (SVM) using 311 samples, including 121 HRD and 190 HRP cancers, from the Sanger-WGS-Breast dataset (Figure 2a). For training purposes, patients were classified as HRD based on genomic changes in BRCA1 and BRCA2 or at least 42 HRD scores. 10-fold cross-validation was performed to determine the feature weights for the training model (Figure 2b). Features with positive weights (LOH: 1-40Mb, DEL.5.MH, 3-9: HET: 10-40Mb, and N[C>G]T) were enriched in HRD samples, while features with negative weights (N[C>T]G and 2-4: Het: >40Mb) were enriched in HRP samples. The performance of the model was tested on a total of 371 samples composed of 311 training samples and 60 holdout HRP samples. To ensure the robustness of the model performance, the model was run across 100 random test datasets generated by 20% randomly sampled from the entire dataset. HRProfiler had an average AUC of 0.97 and an F1 score of 0.86 across 100 test datasets, providing performance comparable to other tools in the same dataset 17(Figure 2c). To determine the applicability of genomic features at exome resolution for HRD prediction, the breast-specific exome HRProfiler model was further trained by applying it to 671 TCGA-WES-Breast cancers composed of 157 HRD and 514 HRP tumors with SVM (Figure 2d). For training purposes, patients were classified as HRD based on genomic changes in BRCA1 and BRCA2 or at least 42 HRD scores. The importance of features based on 10-fold cross-validation of the HRProfiler model demonstrated the robustness of genomic features enriched consistently in HRD samples for LOH: 1-40Mb, DEL.5.MH, and 3-9:HET: 10-40Mb and N[C>G], and in HRP samples for N[C>T]G and 2-4:Het: over 40Mb (Figure 2e). To compare the performance of HRProfiler in predicting HRD status for breast samples profiled at both whole-genome resolution and whole-exome resolution, the HRD status was determined for 65 holdout TCGA breast samples profiled using both WGS and WES by applying each of the whole-genome-based and whole-exome-based HRProfiler models (Figure 2f). At both WGS resolution and WES resolution, HRProfiler predicted the HRD status for breast samples, thereby outperforming the performance of SigMA and HRDetect in highlighting the generalizability of six features in predicting the HRD status for both WGS samples and WES samples. To further validate the performance of HRProfiler in an external independent dataset, HRProfiler was used to predict the HRD probability for 109 exome MSK-IMPACT breast samples, reporting higher sensitivity, AUC, and F1 scores compared to SigMA (Figure 2g).
[0116] Example 5. HRD Detection from WGS and Downsampled WGS Breast Samples To evaluate the predictive ability of the HRProfiler model of WGS in an independent WGS breast dataset, the HRD status was determined using the HRProfiler model of WGS for 237 triple-negative breast cancers (TNBC) with known HRD and HRP annotations and known responses to prior platinum therapy. 23 Next, the performance of HRProfiler was compared with the performance of HRDetect, CHORD, and SigMA. Similar to the previous WGS dataset, HRProfiler demonstrated performance comparable to other tools at WGS resolution (Figure 3a). Similarly, from the perspective of clinical endpoints, all tools presented results showing comparable prognostic benefits based on the disease-free survival (IDFS) for HRD-classified patients with prior chemotherapy (p-value < 0.05, log-rank test, Figures 3b(i)–(iv)). To determine the predictive ability and applicability of the HRProfiler model of WGS with lower genomic resolution, the genomic characteristics of 237 triple-negative breast cancers were first downsampled to exome resolution and to the WES models applied to the previously pre-trained HRDetect WGS model, HRProfiler, and SigMA. CHORD was not used in this data as this tool only supports samples that have been whole-genome sequenced. 17 HRProfiler was able to well-separate HRD samples and HRP samples from the downsampled dataset (Figure 3c). Importantly, HRProfiler was the only tool able to achieve significant stratification based on IDFS across HRD samples and HRP samples (p-value: 0.009; log-rank test; Figures 3d(i)–(iii)).
[0117] Example 6. Training and validation of HRProfiler for predicting HRD status from ovarian samples To determine whether the defined genomic features can be generalized to other HRD-related cancers, 182 TCGA ovarian exome patients (TCGA-WES-Ovarian) consisting of 82 HRD patients and 100 HRP patients were used to train a tissue-specific model for ovarian cancer (Figure 4a). For training purposes, patients were classified as HRD based on genomic changes in BRCA1 and BRCA2 or at least 63 HRD scores. 10-fold cross-validation was performed to determine the feature weights for the training model (Figure 4b). Features with positive weights (LOH: 1-40 Mb, DEL.5.MH, 3-9:HET: 10-40 Mb, and N[C>G]T) were enriched in HRD samples, while features with negative weights (N[C>T]G and 2-4:Het: over 40 Mb) were enriched in HRP samples. The performance of the model was tested by generating 100 training and test datasets by random sampling based on an 80 / 20 split of the training dataset and the test dataset. HRProfiler had an average AUC of 0.93 and an F1 score of 0.78 across 100 test datasets (Figure 4c). To validate the performance of the ovarian-specific HRProfiler model in an independent external dataset, the model was applied to predict the HRD status for 50 MSK-IMPACT ovarian samples using known HRD annotations, and its performance was comparable to SigMA (Figure 4d). To evaluate whether HRProfiler can serve as a prognostic biomarker, it was determined whether there was a statistically significant difference in survival between HRD patients and HRP patients in the holdout test dataset. Progression-free interval (PFI) analysis revealed good survival rates for stratified HRD patients based on HRProfiler (q-value = 0.0156; Cox proportional hazard model ratio), but not based on SigMA (q-value = 1; Cox proportional hazard model ratio) and BRCA1 / 2 mutation status (q-value = 1; Cox proportional hazard model ratio) in holdout TCGA ovarian patients pre-treated with platinum therapy after adjustment for age, clinical stage, and HRD score (Figure 4e(i)-(iii)).
[0118] Example 7. Online Method: Dataset In this disclosure, publicly available datasets were used for feature engineering, model development, and validation at both whole-genome resolution and whole-exome sequencing resolution. For the analysis at whole-genome resolution, CaVEman variant calls and ASCAT allele-specific copy number calls were used for 371 samples from 560 breast datasets 21 [ftp: / / ftp.sanger.ac.uk / pub / cancer / Nik-ZainalEtAl-560BreastGenomes / ]. The additional WGS datasets used in this study included 237 triple-negative breast (TNBC) sample parts from the SCAN-B trial 23 . The CaVEman variant calls and ASCAT copy numbers for the 237 TNBC sample parts were downloaded from https: / / data.mendeley.com / datasets / 2mn4ctdpxp / . For the PCAWG dataset, consensus variant and copy number calls were downloaded from the ICGC data portal: https: / / dcc.icgc.org / releases / PCAWG
[0119] For the analysis at whole-exome resolution, the TCGA dataset was utilized. The catalog of somatic mutations was downloaded from the GDC, and allele-specific exome copy number calls were derived in-house. 109 breast samples and 50 ovarian samples of the MSK-IMPACT exome were downloaded from dbGaP and processed in-house using the EVC pipeline
[0120] Example 8. HRD Definition Considering the lack of clinical response to PARP inhibitors or platinum therapies available for most of the data, ground truth for HRD was derived. This was based on the presence of germline or somatic changes in BRCA1 and BRCA2, or HRD scores of at least 42 for breast patients and 63 for ovarian patients
[0121] Example 9. Feature engineering for predicting HRD To identify features significantly enriched in HRD and HRP samples, average mutation profiles were generated based on the ratios across 96 mutations, 83 indels, and 48 copy number contexts. To determine significant channels at each resolution, Fisher's exact test was performed to determine if there was any significant difference in the average ratio of the channels given across HRD and HRP samples. If the log2 fold change (FC) is greater than 0.75 for WGS samples and greater than 0.25 for WES samples, significant channels are identified in all contexts and their -log 10(Adjusted p-value) is greater than 3. The same workflow was adopted for both the whole-genome samples and the whole-exome samples, and only the channels significantly enriched across both were considered for the feature engineering process. At single-nucleotide resolution, the A[C>T]G, C[C>T]G, G[C>T]G, and T[C>T]G channels were consistently enriched across HRP samples in both the whole-genome dataset and the exome dataset, had overlapping / similar mutation contexts, and were thus combined into a single feature called N[C>T]G. Here, N represents any of the four nucleotide bases (A / C / T / G). Similarly, A[C>G]T, C[C>G]T, and G[C>G]T are all the significant channels enriched in HRD samples and were combined into a single feature N[C>G]T. Here, N represents all possible nucleotide bases. At indel resolution, 5:Del:M:1, 5:Del:M:2, 5:Del:M:3, 5:Del:M:4, and 5:Del:M:5 are significant channels, all of which represent varying lengths of microhomology sequences at relatively large deletion sites where the length of the deletion is at least 5 base pairs. These indel channels were combined into a single feature, namely DEL.5.MH. Here, DEL.5 represents the length of the deletion of at least 5 bp, and MH represents the microhomology sequence. At copy number resolution, multiple significant channels were identified for loss of heterozygosity (LOH) representing LOH segments of sizes 1 - 40 Mb. These were combined into a single feature LOH.1.40Mb. A similar approach was applied to aggregate significant copy number channels for diploid / genome-doubled copy number segments into a single feature 2 - 4:HET:>40Mb, considering segments with a total copy number state of 2 - 4 and sizes of at least 40 Mb. Finally, significant copy number channels for amplification events were combined into a single feature, namely 3 - 9:HET:10 - 40Mb. Here, 3 - 9 represents segments with a total copy number state of at least 3 and a segment size of 10 - 40 Mb.
[0122] Example 10. Model Development and Performance in WGS To train a model for predicting HRD at WGS resolution, samples from 560 breast datasets were used. Only 371 / 560 samples labeled as such, as evaluated in the HRDetect publication, were considered. Six features derived from the feature engineering step were extracted from the 371 samples and normalized using min-max normalization. The initial training was based on 311 breast samples consisting of 121 HRD samples and 190 HRP samples. Next, 10-fold cross-validation was performed to tune hyperparameters and obtain feature weights for the model. The performance of the model was tested on the entire 371 breast dataset, using an HRD probability threshold of 0.3 to classify samples as HRD. The final HRD model was trained on the entire 371 breast samples using a linear kernel support vector machine (SVM) with L1 regularization and tuned hyperparameters. To validate the model on an external dataset, the HRD probability was predicted for 237 triple-negative breast (TNBC) samples, and its performance was evaluated against the ground truth based on molecular changes in the HR pathway or at least 42 HRD scores. The performance of the model was evaluated using conventional machine learning metrics such as AUC, sensitivity, specificity, precision, balanced accuracy (BA), and F1. To compare the performance of the other tools with HRProfiler, the default settings for HRDetect, CHORD, and SigMA were used to determine the HRD probability for 237 TNBC samples.
[0123] Example 11. Model Development and Performance in WES Samples from the TCGA breast dataset were used to train a model for predicting HRD at WES resolution. Only 736 samples with HRD annotation were used for both training and testing. Six features derived from the feature engineering step were extracted as ratios, except for DEL.5.MH, which was extracted as an absolute number. Next, all features were scaled individually by min-max normalization. The initial training was based on 671 breast samples composed of 157 HRD samples and 514 HRP samples. Next, 10-fold cross-validation was performed to tune the hyperparameters and obtain the feature weights for the model. The performance of this model was tested on 65 breast samples sequenced at both whole-genome resolution and whole-exome resolution. Samples with an HRD probability of at least 0.1 were considered HRD. To validate the model in an external dataset, the HRD probability was predicted for 109 MSK-IMPACT breast exome samples, and the performance of the model was evaluated against the ground truth based on molecular changes in the HR pathway or at least 42 HRD scores. The performance of the model was evaluated using conventional machine learning metrics such as AUC, sensitivity, specificity, precision, balanced accuracy (BA), and F1. To compare the performance of the model with other tools and HRProfiler, the default settings for SigMA were used to determine the HRD probability for the same samples. The WES model was also applied to 237 downsampled TNBC samples, and its performance was compared with that of other tools, including HRDetect and SigMA, using models pre-trained with default WGS and WES, respectively. Exome features for the 237 TNBC samples were derived by downsampling the available SNP6 ASCAT copy number calls into segments across the exome region. Variant and indel calls were downsampled to exome resolution using SigProfilerMatrixGenerator.
[0124] Example 12. Survival analysis and statistical analysis Survival analysis was performed using the Kaplan-Meier (KM) and Cox Proportional-Hazards Model (COXPH) functions from the survminer and survival packages in R. Interval Disease Free Survival (IDFS) was used to evaluate the prognostic benefit in patients treated with chemotherapy from 237 TNBC datasets. The Progression Free Interval (PFI) endpoint was used to evaluate the survival trend for TCGA ovarian cancer patients treated with platinum therapy.
[0125] All statistical analyses were performed in Python using the scikit-learn package in Python. All p-values were corrected for multiple hypothesis testing using Benjamini-Hochberg when required.
[0126] In summary, the present technology provides a machine learning approach called HRProfiler that uses a set of at least six genomic features to predict homologous recombination deficiency across both whole-genome sequencing data and whole-exome sequencing data. HRProfiler has performance similar to current tools when applied to the whole genome and outperforms all existing approaches when applied to whole-exome sequencing. HRProfiler incorporates features enriched in both HRD and HRP samples, which generally focus on variants enriched exclusively in HRD samples and are not considered in current methods 17~19 HRProfiler avoids the need to extract structural variants and mutation footprints that may be unreliable when using poor datasets from whole-exome sequencing and targeted panel sequencing. 27 SBS3 26Using the footprint of a single mutation such as this is not reliable for accurate HRD prediction. SBS3 is a flat mutation footprint that is likely to be incorrectly assigned in cancer genomes enriched for other correlated flat mutation footprints such as SBS5 and SBS40. The use of N[C>T]G and N[C>G]T as HRP specificity features serves as a reliable alternative to SBS3 and overcomes the problems associated with using flat mutation footprints as biomarkers at exome resolution.
[0127] The application of HRProfiler across both breast and ovarian cancers outlines the generalizability of features across different cancer types. Overall, the machine learning approach disclosed herein fills the gap in using the molecular phenotype footprint of a failed DNA repair process as a clinical biomarker to reliably stratify patients sensitive to PARP inhibitors and / or platinum therapies.
[0128] It is understood that the various disclosed embodiments may be implemented individually or collectively using various components, electronic device hardware, and / or software modules and components that make up an apparatus. These apparatuses may, for example, include a processor, a memory unit, and interfaces communicatively connected to each other, and may range from desktop and / or laptop computers to mobile devices. The processor and / or controller can perform the various disclosed operations based on the execution of program code stored in a storage medium. The processor and / or controller can communicate directly or indirectly, via communication links with other entities, apparatuses, and networks, for example, with at least one memory and at least one communication unit that enables the exchange of data and information. The communication unit may provide wired and / or wireless communication functions according to one or more communication protocols, and thereby may include appropriate transmission / reception antennas, circuits, and ports, and encoding / decoding functions that may be required for appropriate transmission / reception of data and other information. FIG. 6 shows an example of such an apparatus that includes at least one processor and / or controller, at least one memory unit in communication with the processor, and at least one communication unit capable of directly or indirectly exchanging data and information via communication links with other entities, apparatuses, and networks.
[0129] In one embodiment, the various information and data processing operations described herein may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer in a networked environment. The computer-readable medium may include removable and non-removable storage devices including, but not limited to, read only memory (ROM), random access memory (RAM), compact disc (CD), digital versatile disc (DVD), etc. Thus, the computer-readable media described herein include non-transitory storage media. In general, program modules may include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The computer-executable instructions associated with the data structures and program modules represent example program code for executing the steps of the methods disclosed herein. A particular sequence of such executable instructions or associated data structures represents a corresponding example of operations for implementing the functions described in such steps or processes.
[0130] The foregoing description of embodiments of the technology is not intended to be exhaustive or to limit the technology to the exact form described above. Specific embodiments and examples of the technology are described above for illustrative purposes, but as would be recognized by one of ordinary skill in the art, various equivalent modifications are possible within the scope of the technology. For example, although the steps are presented in a given order, alternative embodiments may perform the steps in a different order. It is also possible to combine the various embodiments described herein to provide further embodiments.
[0131] From the foregoing, specific embodiments of the present technology have been described herein for illustrative purposes, but well-known components and functions have not been illustrated or described in detail to avoid unnecessarily obscuring the description of the present technology. Where the context permits, singular or plural terms may each include plural or singular terms. Also, although the advantages associated with some embodiments of the present technology are described in the context of those embodiments, other embodiments may also exhibit such advantages, and not all embodiments necessarily exhibit such advantages in order to fall within the scope of the present technology. Accordingly, the present disclosure and related technologies may encompass other embodiments not explicitly illustrated or described herein.
[0132] References 1. Ceccaldi, R., Rondinelli, B. & D’Andrea, A.D. in Trends in Cell Biology Vol. 26 52 - 64 (Elsevier Ltd, 2016) 2. Konstantinopoulos, P.A., Ceccaldi, R., Shapiro, G.I. & D’Andrea, A.D. Homologous Recombination Deficiency: Exploiting the Fundamental Vulnerability of Ovarian Cancer. Cancer Discov 5, 1137 - 1154 (2015) https: / / doi.org:10.1158 / 2159 - 8290.CD - 15 - 0714 3. Kasi, A., Al - Jumayli, M., Park, R., Baranda, J. & Sun, W. in Journal of Pancreatic Cancer Vol. 6 107 - 115 (2020) 4. Abida, W. et al. Rucaparib in Men With Metastatic Castration-Resistant Prostate Cancer Harboring a BRCA1 or BRCA2 Gene Alteration. J Clin Oncol 38, 3763-3772 (2020) https: / / doi.org:10.1200 / JCO.20.01035 5. de Bono, J. et al. Olaparib for Metastatic Castration-Resistant Prostate Cancer. N Engl J Med 382, 2091-2102 (2020) https: / / doi.org:10.1056 / NEJMoa1911440 6. Moore, K. et al. Maintenance Olaparib in Patients with Newly Diagnosed Advanced Ovarian Cancer. N Engl J Med 379, 2495-2505 (2018) https: / / doi.org:10.1056 / NEJMoa1810858 7. Tutt, A. et al. in Nature Medicine Vol.24 628-637 (Springer US, 2018) 8. Curtin, N.J. & Szabo, C. in Nature Reviews Drug Discovery Vol.19 711-736 (Springer US, 2020) 9. Wang, D. & Lippard, S.J. Cellular processing of platinum anticancer drugs. Nat Rev Drug Discov 4, 307-320 (2005) https: / / doi.org:10.1038 / nrd1691 10. Abkevich, V. et al. in British Journal of Cancer Vol.107 1776-1782 (2012) 11. Melinda, L.T. et al. in Clinical Cancer Research Vol. 22 3764 - 3773 (2016) 12. Birkbak, N.J. et al. in Cancer Discovery Vol. 2 366 - 375 (2012) 13. Miller, R.E. et al. ESMO recommendations on predictive biomarker testing for homologous recombination deficiency and PARP inhibitor benefit in ovarian cancer. Ann Oncol 31, 1606 - 1622 (2020) https: / / doi.org:10.1016 / j.annonc.2020.08.2102 14. Popova, T. et al. in Cancer Research Vol. 72 5454 - 5462 (2012) 15. How, J.A. et al. in Cancers Vol. 13 1 - 18 (2021) 16. Takaya, H., Nakai, H., Takamatsu, S., Mandai, M. & Matsumura, N. in Scientific Reports Vol. 10 1 - 8 (2020) 17. Davies, H. et al. in Nature Medicine Vol. 23 517 - 525 (Nature Publishing Group, 2017) 18. Nguyen, L., W.M. Martens, J., Van Hoeck, A. & Cuppen, E. in Nature Communications Vol. 11 1 - 12 (2020) 19. Gulhan, D. C., Lee, J. J. K., Melloni, G. E. M., Cortes-Ciriano, I. & Park, P. J. Detecting the mutational signature of homologous recombination deficiency in clinical samples. Nature Genetics 51, 912 - 919 (2019) https: / / doi.org:10.1038 / s41588-019-0390-2 20. Alexandrov, L. B. et al. Signatures of mutational processes in human cancer. Nature 500, 415 - 421 (2013) https: / / doi.org:10.1038 / nature12477 21. Nik-Zainal, S. et al. in Nature Vol. 534 47 - 54 (Nature Publishing Group, 2016) 22. Alexandrov, L. B. et al. The repertoire of mutational signatures in human cancer. Nature 578, 94 - 101 (2020) https: / / doi.org:10.1038 / s41586-020-1943-3 23. Staaf, J. et al. in Nature Medicine Vol. 25 (Springer US, 2019) 24. Zehir, A. et al. in Nature Medicine Vol. 23 703 - 713 (2017) 25. Van Allen, E. M. et al. in Nature Medicine Vol. 20 682 - 688 (Nature Publishing Group, 2014) 26. Alexandrov, L. B. et al. in Nature Vol. 500 415 - 421 (2013) 27. Alexandrov, L.B. et al. in Nature Vol. 578 94 - 101 (2020) 28. Steele, C.D. et al. in Nature Vol. 606 984 - 991 (Springer US, 2022) 29. Steele, C.D. et al. Signatures of copy number alterations in human cancer. Nature 606, 984 - 991 (2022) https: / / doi.org:10.1038 / s41586 - 022 - 04738 - 6 30. Gao, G.F. et al. Before and After: Comparison of Legacy and Harmonized TCGA Genomic Data Commons’ Data. Cell Syst 9, 24 - 34 e10 (2019) https: / / doi.org:10.1016 / j.cels.2019.06.006 31. Pettitt, S.J. et al. Clinical brca1 / 2 reversion analysis identifies hotspot mutations and predicted neoantigens associated with therapy resistance. Cancer Discovery 10, 1475 - 1488 (2020) https: / / doi.org:10.1158 / 2159 - 8290.CD - 19 - 1485 32. Marquard, A.M. et al. Pan - cancer analysis of genomic scar signatures associated with homologous recombination deficiency suggests novel indications for existing cancer drugs. Biomarker Research 3, 1 - 10 (2015) https: / / doi.org:10.1186 / s40364 - 015 - 0033 - 4
Claims
Claim 1 A method for generating a homologous recombination feature set, the method comprising: (a) receiving sequencing data of interest and a corresponding homologous recombination classification; (b) generating a homologous recombination feature set, wherein the homologous recombination feature set includes a plurality of genomic features of the sequencing data of interest and a corresponding homologous recombination classification. Claim 2 The method according to claim 1, wherein the sequencing data includes whole genome sequencing data, whole exome sequencing data, a part thereof, or any combination thereof. Claim 3 The method according to claim 1 or 2, wherein the homologous recombination feature set includes the total number and ratio of deletions in microhomology features of the sequencing data, the total number and ratio of genomic segments having a loss of heterozygosity feature of the sequencing data, the total number and ratio of heterozygous genomic segment features of the sequencing data, the total number and ratio of C:G>T:A single nucleotide substitutions in the 5'-NpCpG-3' context feature of the sequencing data, the total number and ratio of C:G>G:C single nucleotide substitutions in the 5'-NpCpT-3' context feature of the sequencing data, or any combination thereof. Claim 4 The method according to claim 3, wherein the total number and ratio of the genomic segments having a loss of heterozygosity include a size of about 1 to about 40 megabases having at least one copy of the genomic segments having a loss of heterozygosity. Claim 5 The method according to claim 3, wherein the total number and ratio of the heterozygous genomic segments include a size of about 3 to about 40 megabases having 3 to 9 copies of each heterozygous genomic segment of the heterozygous genomic segments. Claim 6 The method according to claim 3, wherein the total number and ratio of the heterozygous genomic segments include a size of at least 40 megabases having 2 to 4 copies of each heterozygous genomic segment of the heterozygous genomic segments. Claim 7 The method according to claim 3, wherein the total number and ratio of deletions in microhomology include a size of at least 5 base pairs. Claim 8 The method according to any one of claims 1 to 7, wherein the homologous recombination classification includes homologous recombination deficiency positive or homologous recombination deficiency negative. Claim 9 The method according to any one of claims 1 to 8, wherein the sequencing data of the subject includes retrospective clinical trial sequencing data of a patient who participated in a clinical trial, and the patient is the same as or different from the subject.
10. A method of training a prediction model configured to predict the presence of homologous recombination deficiency in a subject, the method comprising: (a) receiving the sequencing data of the subject and the corresponding homologous recombination classification; (b) generating a homologous recombination feature set, the homologous recombination feature set including a plurality of genomic features of the sequencing data of the subject and the corresponding homologous recombination classification; (c) training the prediction model using the homologous recombination feature set, thereby generating a trained prediction model configured to predict the presence of homologous recombination deficiency in the subject.
11. The method according to claim 10, wherein the training includes a linear kernel support vector machine (SVM) with L1 regularization.
12. The method according to claim 10 or 11, wherein the prediction model includes a random forest prediction model, a naive Bayes classifier prediction model, a support vector machine prediction model, a logistic regression prediction model, or any combination thereof.
13. The method according to any one of claims 10 to 12, wherein the prediction model is configured to predict the presence of homologous recombination deficiency in the subject with an accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
14. The method according to any one of claims 10 to 13, wherein the prediction model is configured to predict the presence of homologous recombination deficiency in the subject with a precision rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
15. The method according to any one of claims 10 to 14, wherein the prediction model is configured to predict the presence of homologous recombination deficiency in the subject with an F1 of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
16. The method according to any one of claims 10 to 15, wherein the prediction model is configured to predict the presence of homologous recombination deficiency in the patient with a sensitivity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
17. The method according to any one of claims 10 to 16, wherein the prediction model is configured to predict the presence of homologous recombination deficiency in the patient with a specificity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
18. The method according to any one of claims 10 to 17, wherein the prediction model is configured to predict the presence of homologous recombination deficiency in the patient with a balanced accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
19. The method according to any one of claims 10 to 18, wherein the sequencing data includes whole-genome sequencing data, whole-exome sequencing data, a part thereof, or any combination thereof.
20. The method according to any one of claims 10 to 19, wherein the identical recombination feature set includes the total number and ratio of deletions in the microhomology features of the sequencing data, the total number and ratio of genomic segments having the heterozygosity loss feature of the sequencing data, the total number and ratio of heterozygous genomic segment features of the sequencing data, the total number and ratio of C:G>T:A single nucleotide substitutions in the 5'-NpCpG-3' context feature of the sequencing data, the total number and ratio of C:G>G:C single nucleotide substitutions in the 5'-NpCpT-3' context feature of the sequencing data, or any combination thereof.
21. The method according to claim 20, wherein the total number and the ratio of the genomic segments having heterozygosity loss include a size of about 1 to about 40 megabases having at least one copy of the genomic segments having heterozygosity loss.
22. The method according to claim 20, wherein the total number and the ratio of the heterozygous genomic segments include a size of about 3 to about 40 megabases having 3 to 9 copies of each heterozygous genomic segment of the heterozygous genomic segments.
23. The method according to claim 20, wherein the total number and the ratio of the heterozygous genomic segments include a size of at least 40 megabases having 2 to 4 copies of each heterozygous genomic segment of the heterozygous genomic segments.
24. The method according to claim 20, wherein the total number and the ratio of deletions in microhomology include a size of at least 5 base pairs.
25. The method according to any one of claims 10 to 24, wherein the homologous recombination classification includes homologous recombination deficiency positive or homologous recombination deficiency negative.
26. The method according to any one of claims 10 to 25, wherein the sequencing data of the subject includes retrospective study sequencing data of a patient who participated in a clinical trial, and the patient is the same as or different from the subject.
27. A method of administering a cancer therapeutic agent to a subject, the method comprising: (a)receiving the sequencing data of the subject; Determining the homologous recombination classification of the subject as an output of the trained prediction model, wherein the trained prediction model is provided with the sequencing data of the subject as an input, and the trained prediction model is trained with a homologous recombination feature set; Administering the cancer therapeutic agent to the subject according to at least the homologous recombination classification of the subject. A method comprising: **Claim 28** The method according to claim 27, wherein the cancer therapeutic agent comprises at least one selected from the group consisting of platinum therapy and poly(ADP-ribose) polymerase (PARP) inhibitors. **Claim 29** The method according to claim 28, wherein the PARP inhibitor comprises talazoparib, olaparib, niraparib, rucaparib, veliparib or any combination thereof. **Claim 30** The method according to claim 28, wherein the platinum therapy comprises cisplatin, oxaliplatin, carboplatin or any combination thereof. **Claim 31** The method according to any one of claims 27 to 30, wherein the cancer therapeutic agent causes double-strand breaks in the genomic molecules of the cells of the subject and causes apoptosis induced by p53. **Claim 32** The method according to any one of claims 27 to 31, wherein the subject is suspected of having cancer. **Claim 33** The method according to claim 32, wherein the cancer is at least one selected from the group consisting of breast cancer, ovarian cancer, prostate cancer, pancreatic cancer and sarcoma. **Claim 34** The method according to any one of claims 27 to 33, wherein the trained prediction model comprises a prediction model trained by a linear kernel support vector machine (SVM) with L1 regularization. **Claim 35** The method according to any one of claims 27 to 34, wherein the trained prediction model is configured to determine the homologous recombination classification of the subject with an accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%. **Claim 36** The method according to any one of claims 27 to 35, wherein the trained prediction model is configured to determine the homologous recombination classification of the subject with a precision rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
37. The method according to any one of claims 27 to 36, wherein the trained prediction model is configured to determine the homologous recombination classification of the subject with an F1 of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
38. The method according to any one of claims 27 to 37, wherein the trained prediction model is configured to determine the homologous recombination classification of the subject with a sensitivity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
39. The method according to any one of claims 27 to 38, wherein the trained prediction model is configured to determine the homologous recombination classification of the subject with a specificity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
40. The method according to any one of claims 27 to 39, wherein the trained prediction model is configured to determine the homologous recombination classification of the subject with a balanced accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
41. The method according to any one of claims 27 to 40, wherein the sequencing data includes whole genome sequencing data, whole exome sequencing data, a part of any of these, or any combination of these.
42. The method according to any one of claims 27 to 41, wherein the homologous recombination feature set includes genomic features.
43. The method according to claim 42, wherein the genomic features include the total number and ratio of deletions in the microhomology features of the sequencing data, the total number and ratio of genomic segments having a loss of heterozygosity feature in the sequencing data, the total number and ratio of heterozygous genomic segment features in the sequencing data, the total number and ratio of C:G>T:A single nucleotide substitutions in the 5'-NpCpG-3' context feature of the sequencing data, the total number and ratio of C:G>G:C single nucleotide substitutions in the 5'-NpCpT-3' context feature of the sequencing data, or any combination thereof.
44. The method according to claim 43, wherein the total number and the ratio of the genomic segments having a loss of heterozygosity include a size of about 1 to about 40 megabases having at least one copy of the genomic segments having a loss of heterozygosity.
45. The method according to claim 43, wherein the total number and the ratio of the heterozygous genomic segments include a size of about 3 to about 40 megabases having 3 to 9 copies of each heterozygous genomic segment of the heterozygous genomic segments.
46. The method according to claim 43, wherein the total number and the ratio of the heterozygous genomic segments include a size of at least 40 megabases having 2 to 4 copies of each heterozygous genomic segment of the heterozygous genomic segments.
47. The method according to claim 43, wherein the total number and the ratio of the deletions in microhomology include a size of at least 5 base pairs.
48. The method according to any one of claims 27 to 47, wherein the homologous recombination classification includes homologous recombination deficiency positive or homologous recombination deficiency negative.
49. The method according to any one of claims 27 to 48, wherein the sequencing data of the subject includes retrospective clinical trial sequencing data of a patient who participated in a clinical trial, and the patient is the same as or different from the subject.
50. A computer system configured to output a homologous recombination classification of a subject, the computer system comprising: (a) one or more processors; and (b) a non-transitory computer-readable storage medium including software, wherein the software, as a result of execution, causes the one or more processors of the computer system to: (i) receiving the sequencing data of the subject; (ii) when the trained prediction model is provided with the sequencing data of the subject as an input, outputting the homologous recombination classification of the subject as an output of the trained prediction model, wherein the trained prediction model is trained with a homologous recombination feature set; and a non-transitory computer-readable storage medium including executable instructions for causing the above to be performed; a computer system including the same.
51. The computer system according to claim 50, wherein the software includes determining a cancer therapeutic agent at least according to the homologous recombination classification of the subject.
52. The computer system according to claim 51, wherein the cancer therapeutic agent includes at least one selected from the group consisting of platinum therapy and poly(ADP-ribose) polymerase (PARP) inhibitors.
53. The computer system according to claim 52, wherein the PARP inhibitor includes talazoparib, olaparib, niraparib, rucaparib, veliparib, or any combination thereof.
54. The computer system according to claim 52, wherein the platinum therapy includes cisplatin, oxaliplatin, carboplatin, or any combination thereof.
55. The computer system according to any one of claims 51 to 54, wherein the cancer therapeutic agent causes double-strand breaks in the genomic molecules of the cells of the subject and causes apoptosis induced by p53.
56. The computer system according to any one of claims 50 to 55, wherein there is a suspicion that the subject has cancer.
57. The computer system according to claim 56, wherein the cancer is at least one selected from the group consisting of breast cancer, ovarian cancer, prostate cancer, pancreatic cancer, and sarcoma.
58. The computer system according to any one of claims 50 to 57, wherein the trained prediction model includes a prediction model trained by a linear kernel support vector machine (SVM) having L1 regularization.
59. The computer system according to any one of claims 50 to 58, wherein the sequencing data includes whole-genome sequencing data, whole-exome sequencing data, a part thereof, or any combination thereof.
60. The computer system according to any one of claims 50 to 59, wherein the identical recombination feature set includes genomic features.
61. The computer system according to claim 60, wherein the genomic features include the total number and ratio of deletions in microhomology features of the sequencing data, the total number and ratio of genomic segments having a heterozygosity loss feature in the sequencing data, the total number and ratio of heterozygous genomic segment features in the sequencing data, the total number and ratio of C:G>T:A single nucleotide substitutions in the 5'-NpCpG-3' context feature of the sequencing data, the total number and ratio of C:G>G:C single nucleotide substitutions in the 5'-NpCpT-3' context feature of the sequencing data, or any combination thereof.
62. The computer system according to claim 61, wherein the total number and ratio of the genomic segments having a heterozygosity loss include a size of about 1 to about 40 megabases having at least one copy of the genomic segments having a heterozygosity loss.
63. The computer system according to claim 61, wherein the total number and ratio of the heterozygous genomic segments include a size of about 3 to about 40 megabases having 3 to 9 copies of each heterozygous genomic segment of the heterozygous genomic segments.
64. The computer system according to claim 61, wherein the total number and ratio of the heterozygous genomic segments include a size of at least 40 megabases having 2 to 4 copies of each heterozygous genomic segment of the heterozygous genomic segments.
65. The computer system according to claim 61, wherein the total number and ratio of the deletions in microhomology include a size of at least 5 base pairs.
66. The computer system according to any one of claims 50 to 65, wherein the identical recombination classification includes identical recombination deficiency positive or identical recombination deficiency negative.
67. The computer system according to any one of claims 50 to 66, wherein the trained prediction model is configured to output the homologous recombination classification of the subject with an accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
68. The computer system according to any one of claims 50 to 67, wherein the trained prediction model is configured to output the homologous recombination classification of the subject with a precision rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
69. The computer system according to any one of claims 50 to 68, wherein the trained prediction model is configured to output the homologous recombination classification of the subject with an F1 of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
70. The computer system according to any one of claims 50 to 69, wherein the trained prediction model is configured to output the homologous recombination classification of the subject with a sensitivity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
71. The computer system according to any one of claims 50 to 70, wherein the trained prediction model is configured to output the homologous recombination classification of the subject with a specificity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
72. The computer system according to any one of claims 50 to 71, wherein the trained prediction model is configured to output the homologous recombination classification of the subject with a balanced accuracy rate of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99%.
73. The computer system according to any one of claims 50 to 72, wherein the sequencing data of the subject includes retrospective trial sequencing data of a patient who participated in a clinical trial, and the patient is the same as or different from the subject.