Systems and methods for classifying homology-directed repair deficiencies

JP2024528489A5Pending Publication Date: 2025-06-19FOUNDATION MEDICINE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023579476
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-25
Filing Date
2022-06-24
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods for classifying tumors as homologous recombination deficient (HRD) positive or negative are inaccurate and inefficient, often leading to overfitting and requiring extensive data, which hinders appropriate treatment selection.

Method used

A method involving feature selection using importance metrics to identify a subset of features from a plurality of characteristics, training an HRD model, and classifying tumors as HRD positive or negative based on these features, including copy number and short variant characteristics.

Benefits of technology

This approach reduces model overfitting, requires less data and processing power, and improves classification accuracy, enabling more efficient and accurate identification of HRD status for tailored cancer treatments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Described herein are methods, devices, and systems for identifying a subset of a plurality of features using one or more feature importance metrics for training and using a homology repair deficiency (HRD) classification model.Furthermore, described herein are methods, devices, and systems for classifying a tumor of a cancer, such as pancreatic cancer, as likely to be HRD positive or likely to be HRD negative, and for considering the tumor as HRD positive or HRD negative.Described herein are methods for treating a tumor of a cancer, such as pancreatic cancer, based on the classification.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 215,281, entitled “SYSTEM AND METHOD OF CLASSIFYING HOMOLOGOUS REPAIR DEFICIENCY,” filed June 25, 2021, the contents of which are incorporated herein by reference for all purposes. [Technical field]

[0002] FIELD OF THEINVENTION Described herein are methods, devices, and systems for selecting features of a homologous repair deficient (HRD) model, assessing tumors using the HRD model, and treating the tumor based on the assessment. [Background technology]

[0003] 2. Background of the Invention Copy number abnormalities involve the deletion or amplification of large contiguous segments of the genome and are common mutations in cancer. Certain copy number abnormalities are associated with an inability to repair the genome by the homologous recombination repair mechanism, called homologous repair deficiency (HRD). To identify some tumors with HRD, it is possible to sequence mutations in genes involved in the homologous repair pathway. Alternatively, it is possible to detect genomic scarring, which is a physical consequence of HRD, regardless of its cause.

[0004] Tumor genomes that exhibit HRD are associated with sensitivity to certain drugs, such as platinum chemotherapy or poly(ADP) ribose polymerase (PARP) inhibitors. However, it remains difficult to classify certain tumors as HRD positive. Thus, there remains a need to classify tumors, particularly of cancers such as pancreatic, breast, or prostate cancer, as HRD positive or HRD negative, which is particularly important, so that appropriate treatments can be selected and administered to the subject. In the past, techniques for identifying HRD have suffered from inaccuracies and inefficiencies that prevent them from being used in practice. One reason for this is that feature selection techniques are currently insufficient to be able to accurately determine the HRD status of a sample, for example due to overfitting, in order to efficiently and accurately identify (e.g., classify) said tumors as HRD positive or HRD negative. Another reason for this is that it can also be difficult to determine which features to identify in order to accurately determine the status of HRD. Thus, there is a need for techniques and systems that accurately and efficiently select a subset of features from a plurality of features that can be used to train a model to perform said identification. Summary of the Invention

[0005] The method includes the steps of: providing a genome obtained from a tumor of a subject; optionally ligating one or more adaptors onto the genome; amplifying nucleic acid molecules from the genome; capturing nucleic acid molecules from the amplified genome, where the captured nucleic acid molecules are captured by hybridization to one or more bait molecules; deriving a set of input features from the captured nucleic acid molecules; inputting, by one or more processors, the set of input features into a trained homologous recombination deficient (HRD) model; and using the trained HRD model to identify the tumor as HRD positive or HRD negative, where the model detects a feature of each of the plurality of features. Described herein is a method comprising: inputting a set of input features into a trained homologous recombination deficient (HRD) model trained by determining one or more feature importance metrics associated with the plurality of features; identifying a subset of features among the plurality of features using the one or more feature importance metrics; and training, by one or more processors, an HRD model based on the identified subset of features; and using the trained HRD model to identify the tumor as HRD positive or HRD negative; and classifying, by the one or more processors, the tumor as HRD positive or HRD negative using the trained HRD model.

[0006] Further described herein is a method that includes receiving, by one or more processors, a plurality of features; identifying, by the one or more processors, a subset of features of the plurality of features using an importance metric of the one or more features; and training, by the one or more processors, a homologous recombination deficiency (HRD) model based on the identified subset of the plurality of features, wherein the HRD model is configured to receive sample data associated with a genome of a tumor of a subject, and to use the sample data to identify the tumor of the subject as HRD positive or HRD negative.

[0007] Further described herein is a method comprising: receiving, by one or more processors, sample data related to a genome of a tumor in a subject; inputting, by the one or more processors, the sample data into a trained homologous recombination deficient (HRD) model, where the HRD model is trained by: determining one or more feature importance metrics associated with each feature of a plurality of features; identifying a subset of features of the plurality of features using the one or more feature importance metrics; and training, by the one or more processors, the HRD model based on the identified subset of features; and classifying, by the one or more processors, the tumor as HRD positive or HRD negative using the trained HRD model.

[0008] In some embodiments of the described methods, the plurality of features comprises one or more copy number features, one or more short variant features, or a combination thereof. In some embodiments of the described methods, the importance metric of the one or more features comprises one or more of a chi-square test, analysis of variance (ANOVA), random forest, or gradient boosting.

[0009] In some embodiments of the described methods, identifying the subset of features of the plurality of features includes obtaining, by the one or more processors, one or more feature rankings according to an importance metric of the one or more features, and selecting, by the one or more processors, the subset of the plurality of features based on the one or more feature rankings.

[0010] In some embodiments of the described method, identifying the subset of the multiple features includes (a) obtaining, by one or more processors, a feature ranking of the multiple features according to a feature importance metric; (b) obtaining, by the one or more processors, a new feature set by adding, by the one or more processors, one or more features from the multiple features based on the feature ranking to the existing feature set; (c) training, by the one or more processors, a new HRD model using the new feature set; (d) evaluating, by the one or more processors, the trained new HRD model to obtain an evaluation result; (e) storing, by the one or more processors, the evaluation result associated with the new HRD model and the new feature set; (f) repeating, by the one or more processors, steps (b)-(e) to obtain a plurality of evaluation results until a condition is satisfied; and (g) selecting, by the one or more processors, a subset of the multiple features based on the plurality of evaluation results.

[0011] In some embodiments of the described methods, the trained HRD model is a classification model and the method further comprises receiving new sample data associated with the genome of a tumor in a new subject, the new sample data being associated with a subset of the plurality of features, feeding the new sample data to the trained HRD classification model to generate a classification result of HRD positive or HRD negative, and outputting the classification result. In some embodiments, the classification result comprises at least one of an HRD positive likelihood score and an HRD negative likelihood score. In some embodiments, the method comprises recording at least one of the HRD positive likelihood score and the HRD negative likelihood score in a digital electronic file associated with the new subject. In some embodiments, the method comprises recording a designation in a digital electronic file associated with the new subject that the tumor is HRD positive based on the HRD positive likelihood score or that the tumor is HRD negative based on the HRD negative likelihood score.

[0012] In some embodiments of the described methods, the HRD model is a classification model, a regression model, a neural network, or any combination thereof. In some embodiments, the method includes recording at least one of the HRD-positive likelihood score and the HRD-negative likelihood score in a digital electronic file associated with the new subject. In some embodiments, the method includes recording in a digital electronic file associated with the new subject a designation that the tumor is HRD-positive based on the HRD-positive likelihood score, or that the tumor is HRD-negative based on the HRD-negative likelihood score.

[0013] In some embodiments of the described method, the plurality of features includes at least one of the following features: segment minor allele frequency (segMAF), number of sequencing reads, segment size, number of breakpoints per x megabase, changepoint copy number, segment copy number, number of breakpoints per chromosome arm, and number of segments of oscillating copy number. In some embodiments of the described method, at least one of the plurality of features is evaluated over a centromeric portion of the genome. In some embodiments of the described method, at least one of the plurality of features is evaluated over a telomeric portion of the genome.

[0014] In some embodiments of the described methods, at least one of the plurality of features is assessed across both centromeric and telomeric portions of the genome.

[0015] In some embodiments of the described methods, the plurality of features comprises a breakpoints per x megabases feature, the breakpoints per x megabases feature being based on the number of breakpoints occurring in a window of length x megabases across the genome. In some embodiments, the breakpoints per x megabases feature is evaluated across (i) a telomeric portion of the genome, (ii) a centromeric portion of the genome, or (iii) both a telomeric portion and a centromeric portion of the genome. In some embodiments, x is between about 1 and about 100 megabases. In some embodiments, x is about 10 megabases, about 25 megabases, about 50 megabases, or about 100 megabases. In some embodiments, the breakpoints per x megabases feature is a binned feature.

[0016] In some embodiments of the described method, the plurality of features comprises a changepoint copy number feature, and the changepoint copy number is based on the absolute difference in copy number between adjacent genomic segments across the genome of the tumor of interest. In some embodiments, the changepoint copy number feature is derived from ploidy-normalized copy number data. In some aspects, the changepoint copy number feature is evaluated across (i) the telomeric portion of the genome, (ii) the centromeric portion of the genome, or (iii) both the telomeric portion and the centromeric portion of the genome. In some embodiments, the changepoint copy number feature is a binned feature.

[0017] In some embodiments of the described methods, the plurality of features comprises segment copy number features, and the segment copy number is based on the copy number of each genome segment. In some aspects, the segment copy number features are evaluated over (i) the telomeric portion of the genome, (ii) the centromeric portion of the genome, or (iii) both the telomeric portion and the centromeric portion of the genome. In some embodiments, the segment copy number features are derived from ploidy-normalized copy number data. In some embodiments, the segment copy number features are binned features.

[0018] In some embodiments of the described methods, the plurality of features comprises a feature of the number of breakpoints per chromosome arm of the genome of the tumor of the subject. In some embodiments, the feature of the number of breakpoints per chromosome arm is evaluated over (i) the telomeric portion of the genome, (ii) the centromeric portion of the genome, or (iii) both the telomeric and centromeric portions of the genome. In some embodiments, the feature of the number of breakpoints per chromosome arm is a binned feature.

[0019] In some embodiments of the described method, the plurality of features includes a feature of the number of segments of oscillating copy number. In some embodiments, the feature of the number of segments of oscillating copy number is based on the number of repeated alternating segments between two copy numbers across the genome of the tumor of the subject. In some embodiments, the feature of the number of segments of oscillating copy number is evaluated across (i) the telomeric portion of the genome, (ii) the centromeric portion of the genome, or (iii) both the telomeric portion and the centromeric portion of the genome. In some embodiments, the feature of the number of segments of oscillating copy number is a binned feature.

[0020] In some embodiments of the described methods, the one or more copy number features include a segment minor allele frequency (segMAF) feature, and the segMAF is based on the minor allele frequency at heterozygous single nucleotide polymorphisms. In some aspects, the segMAF is evaluated over (i) the telomeric portion of the genome, (ii) the centromeric portion of the genome, or (iii) both the telomeric and centromeric portions of the genome. In some embodiments, the segMAF feature is a binned feature.

[0021] In some embodiments of the described methods, the one or more copy number features include a number of sequencing reads feature. In some embodiments, the number of sequencing reads feature is a binned feature.

[0022] In some embodiments of the described methods, the plurality of features further comprises a measure of genome-wide loss of heterozygosity in the genome of the tumor of the subject.

[0023] In some embodiments of the described methods, the plurality of features comprises one or more short variant features. In some embodiments, the one or more short variant features comprise at least one of a deletion in a microhomology or repeat region feature and a mutation signature derived from two or more short variant features. In some embodiments, the deletion in the microhomology or repeat region feature is a deletion of at least 5 base pairs.

[0024] In some embodiments of the described method, training the HRD model includes receiving, by one or more processors, an HRD positive training dataset, the HRD positive training dataset including a plurality of features associated with HRD positive tumors and HRD positive indicators; receiving, by one or more processors, an HRD negative training dataset, the HRD negative training dataset including a plurality of features associated with HRD negative tumors and HRD negative indicators; and training, by one or more processors, the HRD model using the HRD positive training dataset and the HRD negative training dataset. In some embodiments, the training includes using the HRD positive training dataset and the HRD negative training dataset. In some embodiments, the method includes balancing, by one or more processors, the HRD positive training dataset and the HRD negative training dataset before training the HRD model.

[0025] In some embodiments of the described methods, the method further comprises testing, by one or more processors, the trained model using an HRD positive test dataset comprising an HRD positive control derived from a genomic sequence comprising a loss-of-function mutation in BRCA1, BRCA2, both BRCA1 and BRCA2, or a biallelic mutation in BRCA1 and BRCA2. In some embodiments, the training comprises using an HRD positive training dataset and an HRD negative training dataset. In some embodiments, the method comprises balancing, by one or more processors, the HRD positive training dataset and the HRD negative training dataset prior to training the HRD model.

[0026] In some embodiments of the described methods, the method further comprises testing the trained model using an HRD-positive test dataset comprising an HRD-positive control derived from a genomic sequence comprising a loss-of-function mutation in at least one of ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, or RAD45L, by one or more processors. In some embodiments, the training comprises using an HRD-positive training dataset and an HRD-negative training dataset. In some embodiments, the method comprises balancing the HRD-positive training dataset and the HRD-negative training dataset, by one or more processors, before training the HRD model.

[0027] In some embodiments of the described methods, the method further comprises testing, by one or more processors, the trained model using an HRD-negative test dataset comprising an HRD-negative training dataset comprising an HRD-negative control derived from a consensus human genome sequence. In some embodiments, the training comprises using an HRD-positive training dataset and an HRD-negative training dataset. In some embodiments, the method comprises balancing, by one or more processors, the HRD-positive training dataset and the HRD-negative training dataset prior to training the HRD model.

[0028] In some embodiments of the methods described, the tumor in the subject is prostate cancer, non-small cell lung cancer (NSCLC), colorectal cancer (CRC), ovarian cancer, breast cancer, or pancreatic cancer.

[0029] In some embodiments of the described methods, the step of training the HRD model includes fitting the HRD model to sample data related to ovarian cancer, non-small cell lung cancer (NSCLC), colorectal cancer (CRC), breast cancer, pancreatic cancer, or prostate cancer, where the sample data includes a subset of the plurality of features.

[0030] In some embodiments of the described method, the tumor is obtained from a sample that is a solid tissue biopsy sample. In some embodiments, the solid tissue biopsy sample is a formalin-fixed paraffin-embedded (FFPE) sample. In some embodiments of the described method, the tumor is obtained from a sample that is a liquid biopsy sample that contains circulating tumor DNA (ctDNA). In some embodiments of the described method, the tumor is obtained from a sample that is a liquid biopsy sample that contains cell-free DNA (cfDNA).

[0031] In some embodiments of the described methods, the method further comprises determining, identifying, or applying the output of the tumor as HRD positive or HRD negative as a diagnostic value associated with the patient. In some embodiments of the described methods, the method further comprises generating a genomic profile of the subject based on the output of the tumor as HRD positive or HRD negative. In some embodiments, the method further comprises administering an anti-cancer agent or applying an anti-cancer treatment to the subject based on the generated genomic profile. In some embodiments of the described methods, the output of the tumor as HRD positive or HRD negative is used to generate a genomic profile of the subject. In some embodiments of the described methods, the output of the tumor as HRD positive or HRD negative is used in making a proposed treatment decision for the subject. In some embodiments of the described methods, the output of the tumor as HRD positive or HRD negative is used to apply or administer a treatment to the subject.

[0032] In some embodiments of the described methods, the HRD model is a machine learning model.

[0033] In some embodiments of the methods described, the subject has, is at risk of having, or is suspected of having cancer.

[0034] Further described herein is a method of treating cancer in a subject, the method comprising: (a) identifying a tumor as HRD positive or HRD negative according to any of the methods described above; and (b) administering to the subject a therapeutically effective amount of a drug effective against HRD-positive tumors if the tumor of the cancer is assessed as HRD positive. In some embodiments, the drug effective against HRD-positive tumors is a platinum-based drug or a PARP inhibitor. In some aspects, the method comprises administering to the subject a therapeutically effective amount of a drug that is not a platinum-based drug or a PARP inhibitor if the tumor is assessed as HRD negative.

[0035] Further described herein is a method for selecting a therapy for a subject's cancer, the method comprising: (a) assessing the cancer tumor as HRD positive or HRD negative according to any of the methods described above; and (b) selecting a therapy effective for HRD-positive tumors if the cancer is assessed as HRD positive. In some aspects, the method comprises selecting a therapy that is not a platinum-based drug or a PARP inhibitor if the tumor is assessed as HRD negative. In some embodiments, the therapy effective for HRD-positive tumors is a platinum-based drug or a PARP inhibitor.

[0036] Further described herein is a computer system comprising one or more processors, a memory, and one or more programs, the one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for performing a method according to any one of claims 1 to 65.

[0037] Further described herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform any one of the methods described above. [Brief description of the drawings]

[0038] [Figure 1] 1 shows an exemplary process for classifying a cancer tumor in a subject as HRD positive (HRD(+)) or HRD negative (HRD(-)).

[0039] [Diagram 2] We show different types of features that can be evaluated using different feature importance metrics such as ANOVA, random forests, gradient boosting (e.g., XGB), and chi-squared.

[0040] [Figure 3A] 1 illustrates an exemplary feature overlap analysis.

[0041] [Figure 3B] 1 illustrates an exemplary feature overlap analysis.

[0042] [Figure 4] 1 illustrates an exemplary iterative feature selection process.

[0043] [Diagram 5] 1 shows an exemplary plot of the performance of models resulting from an exemplary iterative feature selection process.

[0044] [Figure 6A] 1 illustrates an exemplary cross-validation process that may be used to evaluate and tune model performance.

[0045] [Figure 6B] 3 illustrates an exemplary division of a number of data elements into equal sized subsets.

[0046] [Figure 7] An exemplary method for training and operating an HRD classification model configured to classify a subject's cancer tumor as HRD positive (HRD(+)) or HRD negative (HRD(-)) is shown.

[0047] [Figure 8] We show examples of HRD score distributions for different machine learning models using logistic regression, gradient boosting (e.g., XGB), and random forest.

[0048] [Figure 9]Figure 1 shows the performance of an exemplary model in samples stratified by HRD and / or BRCA1 / 2 mutation status. On the left side, we show a pool of sample tumors referred to as "HRD wildtype: true" (N=245,050; -1 on the right side of the figure), "HRD wildtype: false" (N=30,799; 0 on the right side of the figure) and true HRD positive samples (biallelic BRCA mutation; N=6,851; 1 on the right side of the figure).

[0049] [Figure 10] Figure 1 shows the performance of exemplary models from subsets of Figure 9 in different tumor types (breast, ovarian, pancreatic and prostate cancer). For each tumor type, the subsets correspond to subsets-1, 0, and 1 of Figure 9 (i.e., HRD wildtype: true, HRD wildtype: false, and biallelic BRCA mutation for each cancer).

[0050] [Figure 11] 1 illustrates an example of a computing device that can be used in certain methods described herein, according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0051] Detailed Description of the Invention Described herein is a computer-implemented method for identifying a subset of a plurality of features using one or more feature importance metrics for training a homologous recombination deficiency (HRD) model (e.g., a classification model). The model is configured to receive test sample data on a subset of a plurality of features related to the genome of a tumor in a subject, and identify (e.g., classify) the tumor as likely to be HRD positive or likely to be HRD negative. Further described herein is a method for identifying (e.g., classifying) a tumor, such as a tumor of prostate cancer, ovarian cancer, breast cancer, colorectal cancer, NSCLC, or pancreatic cancer, as likely to be HRD positive (HRD(+)) or likely to be HRD negative (HRD(-)). Further described herein is a method for treating cancer, such as, but not limited to, pancreatic cancer, prostate cancer, ovarian cancer, breast cancer, non-small cell lung cancer (NSCLC), or colorectal cancer (CRC), based on the identification of the tumor as HRD positive (or likely to be HRD positive) or HRD negative (or likely to be HRD negative).

[0052] Selecting a subset of features can reduce overfitting of the model. Overfitting is problematic because it reduces the scalability of the model and can result in inaccurate classification (e.g., inaccurate HRD status) as the model ignores scenarios that are outside the scope of the data used to train the model. Furthermore, by selecting a subset of features with higher feature importance, the classification model can be trained with less training data and requires less input data. This not only allows for a more efficient modeling process, but also allows for more accurate classification from a wider range of samples from the model. Furthermore, models with a reduced set of input features may require less processing power to train and to perform classification tasks. Thus, the feature selection process improves the functionality of the computer system by improving processing speed and allowing for efficient use of computer memory and processing power. Furthermore, by selecting from specific derived copy number features and / or short variant features, the trained model results in greater efficiency and accuracy (e.g., fewer false positives / false negatives) when identifying tumors as HRD positive or HRD negative compared to previous methods. Previous methods of assessing HRD, such as loss of heterozygosity, telomeric allelic imbalance, and large scale transitions, are subject to noise and error compared to the assessment of derived copy number signatures and / or short variant signatures described herein. Proper identification of tumors is essential to being able to appropriately select patient (subject) treatment.

[0053] Tumor formation is driven, in part, by the accumulation of somatic alterations in a cell's genome. Among these alterations are copy number alterations, which are common in many cancers. Loss of function, gain of function, or gene regulatory mutations in certain genes involved in homologous repair defective pathways can lead to the accumulation of these copy number alterations. However, other than mutations in certain key genes such as BRCA1 and BRCA2, the exact combination of mutations that result in an HRD-positive status is unknown. Some tumors become HRD-positive through non-genomic means, for example, through promoter methylation of HRD-associated genes such as BRCA1. Instead of sequencing HRD-associated genes, an alternative approach is to identify and evaluate the consequences of HRD, such as specific copy number alteration signatures, or loss of heterozygosity signatures. However, while both HRD-positive and HRD-negative genomes can show copy number alterations, the exact values ​​and combinations of features that indicate the presence of HRD are unknown.

[0054] Thus, in one aspect, the methods of the invention relate to selecting a subset of features (from a larger plurality of potential features) that can be used to train and operate an HRD classifier process. In another aspect, the methods of the invention generally relate to a means of identifying (e.g., classifying) tumors that are likely to be HRD positive (HRD(+)) or likely to be HRD negative (HRD(-)) based at least in part on an evaluation of features, such as features corresponding to copy number aberrations. This classification is generally based on an evaluation of the likelihood that the tumor is HRD positive or HRD negative. Based on this evaluation, the HRD classifier process can further consider the tumor as HRD positive or HRD negative. Such classification and / or consideration can be used as a diagnostic value for patients with the tumor.

[0055] Existing methods for classifying tumors as likely HRD positive or likely HRD negative are often unreliable or inaccurate, especially for HRD-positive tumors with wild-type BRCA1 and BRCA2 (sometimes described as tumors with a "BRCAness" profile, i.e., tumors that show similarity to BRCA1 / 2-mutated tumors without having an associated BRCA1 / 2 mutation). Alternatively, not all mutations, even pathogenic mutations such as BRCA1 / 2 alterations, result in HRD (e.g., some mutations may be monoallelic passengers). Cancer-associated homologous repair deficiencies compromise tumor cell genomes, resulting in detectable changes in copy number (i.e., copy number abnormalities) and / or indel patterns. The specific patterns, distributions and morphologies of these copy number abnormalities and / or indel patterns can be used to classify tumors into HRD phenotypic classes. The present application, in various embodiments, provides means for selecting features associated with these patterns (i.e., copy number features) and indel patterns (i.e., short variant features) from among other potential features (such as cardinal features as described elsewhere herein) that can be used to identify HRD-positive tumors.

[0056] The present application further provides a specifically configured model based on one or more data features (such as one or more copy number features and / or one or more short variant features) related to the genome of a cancerous tumor in a subject, which can more reliably identify (e.g., classify) a tumor as likely to be HRD positive or likely to be HRD negative, and optionally consider the tumor as HRD positive or HRD negative. Identifying (e.g., classifying) a cancerous tumor in a subject indicates how the tumor should be treated. For example, a trained HRD model using test data including at least one or more copy number features, including one or more of segment size features, sequencing read features, absolute copy number features, number of breakpoints per x megabase features, change point copy number features, segment copy number features, number of breakpoints per chromosome arm features, number of segments of oscillating copy number features, and segment minor allele frequency features, can be used to identify (e.g., classify) a test tumor as likely to be HRD positive or likely to be HRD negative, and also consider the tumor as HRD positive or HRD negative based on the likelihood score. These categories of copy number features have been identified as being useful for this discrimination. Certain categories of short variant features have also been identified as being useful for this discrimination, including, but not limited to, for example, deletions (e.g., at least 5 base pairs) in microhomology or repeat region features and / or mutational signatures incorporating two or more short variant features.

[0057] In combination with one or more of these copy number features and / or one or more of these short variant features, other features or measurements may be useful in the described methods, including, but not limited to, certain fundamental features such as the subject's age, cancer type, cancer stage, tumor purity, tumor genomic ploidy, and / or tumor genomic loss of heterozygosity.

[0058] Once a cancer tumor in a subject has been identified (e.g., classified) as likely to be HRD positive or likely to be HRD negative, or considered to be HRD positive or HRD negative, it can be treated with an appropriate therapy. For example, if a tumor is identified as likely to be HRD positive, it can be treated with a drug that is effective against HRD-positive cancers, such as a platinum-based drug or a PARP inhibitor.

[0059] definition As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.

[0060] Reference herein to "about" a value or parameter includes (and accounts for) the variation that is directed to the value or parameter itself. For example, a statement of "about X" includes the statement "X."

[0061] The terms "cancer" and "cancerous" refer to or describe the physiological condition in mammals that is typically characterized by unregulated cell growth. This definition includes benign and malignant cancers. "Early cancer" or "early stage tumor" refers to a cancer that is not invasive or metastatic or is classified as stage 0, 1, or 2 cancer.Examples of cancer include lung cancer (e.g., non-small cell lung cancer (NSCLC)), kidney cancer (e.g., renal urothelial carcinoma), bladder cancer (e.g., bladder urothelial (transitional cell) carcinoma), breast cancer, colorectal cancer (e.g., colon adenocarcinoma), ovarian cancer, pancreatic cancer, gastric cancer, esophageal cancer, mesothelioma, melanoma (e.g., cutaneous melanoma), head and neck cancer (e.g., head and neck squamous cell carcinoma (HNSCC)), thyroid cancer, sarcoma (e.g., soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteosarcoma, osteoporosis ... sarcoma), osteosarcoma, chondrosarcoma, angiosarcoma, endothelial sarcoma, lymphangiosarcoma, lymphangioendothelial sarcoma, leiomyosarcoma, or rhabdomyosarcoma), prostate cancer, glioblastoma, cervical cancer, thymic carcinoma, leukemia (e.g., acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic myelogenous leukemia (CML), chronic eosinophilic leukemia, or chronic lymphocytic leukemia (CLL)), lymphoma (e.g., Hodgkin's lymphoma or non-Hodgkin's lymphoma (NHL), myeloma (e.g., multiple myeloma (MM)), mycosis fungoides, Merkel cell carcinoma, hematological malignancies, cancer of blood tissue, B-cell cancer, bronchial cancer, gastric cancer, brain or central nervous system cancer, peripheral nervous system cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, testicular cancer, biliary tract cancer, small intestine or appendix cancer, salivary gland cancer, adrenal gland cancer, adenocarcinoma, inflammatory myofibroblastic fibrosis, Cellular tumors, gastrointestinal stromal tumor (GIST), colon cancer, myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), polycythemia vera, chordoma, synovium, Ewing's tumor, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharynx These include, but are not limited to, cranioma, ependymoma, meningioma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid carcinoma, small cell carcinoma, essential thrombocythemia, amelanotic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, well-known eosinophilia, neuroendocrine carcinoma, or carcinoid tumor.

[0062] As used herein, "tumor" refers to all neoplastic cell growth and proliferation, whether malignant or benign, and all pre-cancerous and cancerous cells and tissues. The terms "cancer," "cancerous," and "tumor" are not mutually exclusive when referred to herein.

[0063] The terms "individual," "patient," and "subject" are used interchangeably and refer to a mammal, including, but not limited to, a human, bovine, equine, feline, canine, rodent, or primate. In some embodiments, the subject is a human.

[0064] As used herein, the term "effective amount" or "therapeutically effective amount" refers to an amount of a compound, drug, or composition sufficient to treat a particular disorder, condition, or disease, e.g., improve, alleviate, relieve, and / or delay one or more of its symptoms. In the context of cancer, an effective amount includes an amount sufficient to reduce the number of cancer cells present in a subject, in number and / or size, and / or slow the rate of growth of cancer cells. In some embodiments, an effective amount is an amount sufficient to prevent or delay recurrence of the disease. In the case of cancer, an effective amount of a compound or composition can (i) reduce the number of cancer cells; (ii) inhibit, delay, partially delay, preferably stop the proliferation of cancer cells; (iii) prevent or delay the onset and / or recurrence of cancer; and / or (iv) relieve to some extent one or more symptoms associated with cancer.

[0065] As used herein, "treatment" or "treating" is an approach to obtain beneficial or desired results, including clinical results. For purposes of the present invention, beneficial or desired clinical results include, but are not limited to, one or more of the following: alleviating one or more symptoms caused by a disease, reducing the extent of the disease, stabilizing the disease (e.g., preventing or delaying the worsening of the disease), preventing or delaying the spread of the disease (e.g., metastasis), preventing or delaying the recurrence of the disease, delaying or delaying the progression of the disease, improving the disease state, providing remission (partial or total) of the disease, reducing the dose of one or more other drugs required to treat the disease, delaying the progression of the disease, improving the quality of life, and / or prolonging survival. With respect to cancer, the number of cancer cells present in a subject may be reduced in number and / or size, and / or the rate of growth of the cancer cells may be slowed. In some embodiments, the treatment may prevent or delay the recurrence of the disease. In the case of cancer, treatment may be to: (i) reduce the number of cancer cells; (ii) inhibit, slow, slow to some extent, and preferably stop the proliferation of cancer cells; (iii) prevent or delay the onset and / or recurrence of cancer; and / or (iv) relieve to some extent one or more symptoms associated with cancer. The methods of the invention contemplate any one or more of these aspects of treatment.

[0066] It is understood that the embodiments and variations of the invention described herein include "consisting of" and / or "consisting essentially of" embodiments and variations.

[0067] When a range of values ​​is provided, it is understood that each intervening value between the upper and lower limits of that range, and any other stated or intervening value within that stated range, is included within the scope of the disclosure. When a stated range includes an upper or lower limit, ranges excluding any of those included limits are also included in the disclosure.

[0068] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention and is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be readily apparent to those skilled in the art, and the generic principles of the present specification may be applied to other embodiments. Thus, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.

[0069] The figures illustrate processes according to various embodiments. In the exemplary processes, some blocks are arbitrarily combined, the order of some blocks is arbitrarily changed, and some blocks are arbitrarily omitted. In some instances, additional steps may be performed in combination with the exemplary processes. Thus, the operations illustrated (and described in more detail below) are exemplary in nature and therefore should not be considered limiting.

[0070] The disclosures of all publications, patents, and patent applications referenced herein are each incorporated herein by reference in their entirety. To the extent that a reference incorporated by reference conflicts with the present disclosure, the present disclosure shall control.

[0071] Feature Selection Starting with a plurality of features, including those described elsewhere herein, a subset of the plurality of features can be identified using one or more feature importance metrics. In general, a feature importance metric allows for the evaluation of individual features to determine which features may be most relevant to the evaluation of HRD. Exemplary feature importance metrics include, but are not limited to, gradient boosting (e.g., XGBoost, also known as XGB), analysis of variance (ANOVA), chi-square analysis, and random forest. Individual features can be assigned values ​​based on their feature importance metrics, and features are assigned increasing importance based on their increasing contribution to the performance of the HRD model (e.g., improving the model's performance in classifying tumors as HRD positive or HRD negative). More important features, such as features that exceed a threshold (e.g., features that exceed the median value of the plurality of features), can then be selected for use in training or running the HRD model. Once a subset of features is identified, the subset of features can be used to train an HRD model (e.g., a classification model). The HRD model can then be used to identify (e.g., classify) tumors of interest using test data obtained from tumors and that include at least some of the features identified during feature selection.

[0072] By selecting this subset of features with higher feature importance, the model can be trained with less training data and requires less input data, thus improving memory usage and management. Furthermore, models with a reduced set of input features require less processing power to train and to perform discrimination (e.g., classification) tasks. Thus, the feature selection process improves the functionality of computer systems by improving processing speed and enabling efficient use of computer memory and processing power.

[0073] FIG. 1 illustrates an exemplary process for classifying a cancer tumor in a subject as HRD positive or HRD negative, including a block for identifying a subset of a plurality of features, according to some embodiments. In some embodiments, the process 100 is performed using, for example, one or more electronic devices implementing a software program. In some examples, the process 100 is performed using a client-server system, and the blocks of the process 100 are divided in any manner between a server and a client device. In other examples, the process 100 is performed using only a client device, or only multiple client devices. In the process 100, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some examples, additional steps can be performed in combination with the process 100. Thus, the operations illustrated (and described in more detail below) are exemplary in nature, and thus should not be considered limiting.

[0074] In block 102 of FIG. 1, an exemplary system (e.g., one or more electronic devices) receives a plurality of features. In some embodiments, the system receives a dataset including a plurality of data elements. The data elements can include data regarding a plurality of features and associated classification indicators (e.g., HRD positive or HRD negative). For example, the data elements can include data regarding a plurality of features of a sample from a particular subject, and associated classification indicators indicating whether the sample is HRD positive and HRD negative. The features can include features classified as basic features, copy number features, and / or short variant features (e.g., features corresponding to base substitutions or indels (insertions or deletions)). The basic features can include, but are not limited to, features regarding the age of the patient from whom the data was obtained, the type of cancer, the stage of the cancer, the purity of the tumor, the genomic ploidy of the tumor, and the tumor genomic loss of heterozygosity (such as the percentage of the genome under heterozygosity loss). The copy number feature may include, but is not limited to, a segment size feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a change point copy number feature, a segment copy number feature, a number of breakpoints per chromosome arm feature, a number of segments of oscillating copy number, and a segment minor allele frequency feature. The short variant feature may include, but is not limited to, for example, a deletion (e.g., at least 5 base pairs) in a homopolymer or repeat region feature, and / or a mutation signature incorporating two or more short variant features. In some embodiments, one or more of the features are binned features, and the values ​​are sorted into bins, such as 2nd quantile, 3rd quantile, 4th quantile, 5th quantile, 6th quantile, 7th quantile, or any other suitable binning configuration.

[0075] In block 104 of FIG. 1, the system and method select a subset of features from the plurality of features (i.e., basic features, copy number features, and / or short variant features). The selected subset of features may have a relatively high predictive value for classifying the subject's cancer tumor as HRD positive or HRD negative. In some embodiments, features that have a relatively low predictive value and / or are redundant may be excluded from the subset of features in block 104. In some embodiments, the predictive value of the features may be quantified using a feature importance metric. In some embodiments, the feature importance metric may be applied to obtain a feature importance score for each feature of the plurality of features. The feature importance score of the feature is obtained from a statistical correlation of the feature with the classification label (e.g., HRD positive or HRD negative). The statistical correlation between the feature and the classification label may be interpreted based on how much predictive value the feature has for the classification task. In other words, for example, a higher feature importance score may be achieved by having a higher statistical correlation between the feature and the classification label, which may indicate that the feature plays a more important role in predicting the classification label. By using features with higher feature importance, classification models can be trained with less data, thus providing a great degree of effectiveness to the training process and less constraints on computer resources (e.g., memory usage, processing speed, etc.). For example, a model with a reduced set of input features may require fewer processing resources to train and to perform the classification task. Finally, a model with a reduced set of input features may exhibit less noise and avoid overtraining. Thus, the feature selection process improves the overall effectiveness of the training process, improves processing speed, and improves the performance of computer systems by enabling efficient use of computer memory and processing power.

[0076] In some embodiments, the system selects a subset of features from the plurality of features received in block 102 of FIG. 1 by performing a feature overlap analysis, as indicated by block 104a. In block 104a, the importance metric of each feature is used to calculate a feature importance score for the plurality of features received from block 102. For each feature importance metric, the system can rank the plurality of features according to their feature importance scores. Thus, the system can obtain a plurality of feature rankings corresponding to the plurality of feature importance characteristics. The system can then identify the subset of features based on the plurality of rankings. The process of ranking the features and identifying the subset of features is described in more detail below.

[0077] In some embodiments, different types of features can be evaluated using different feature importance metrics. FIG. 2 illustrates multiple feature importance metrics that may be used to rank the multiple features in block 104a, according to some embodiments. The illustrated example feature importance metrics include ANOVA, Random Forest, Gradient Boosting (e.g., XGB), and Chi-Square. Additionally, ANOVA can be used to evaluate the numerical features of the multiple features and obtain a ranking of the numerical features. Chi-Square can be used to evaluate the categorical features of the multiple features and obtain a ranking of the categorical features. Random Forest can be used to evaluate all of the multiple features and rank all of the features. Similarly, Gradient Boosting (e.g., XGB) can be used to evaluate all of the multiple features and rank all of the features.

[0078] In some embodiments, the feature importance metric includes an analysis of variance (ANOVA) model. ANOVA evaluates whether there is equal variance between groups (i.e., HRD positive or HRD negative) when a numerical input variable is compared to a variable to be classified. If there is equal variance between groups, the feature does not affect the response and may not be considered for training the model. Based on the value of the variance (f-value), the features can be ranked, and for example, features above the median can be selected as useful features for the model.

[0079] In some embodiments, the feature importance metric includes chi-square analysis. For feature selection, chi-square analysis examines how much the expected count (i.e., if the feature is independent of the output) and the observed count deviate from each other. A higher chi-square value for a feature indicates that it is more dependent on the response variable and therefore more important. Chi-square analysis can be used to rank features, and features above the median, for example, can be selected as useful features for the model.

[0080] In some embodiments, the feature importance metric includes a random forest analysis. During feature selection, for each tree, the prediction accuracy of the out-of-bag portion of the data is recorded. This process is repeated after permuting each predictor variable. The difference between the two accuracies is then averaged across all trees and normalized by the standard error.

[0081] In some embodiments, the feature importance metric includes gradient boosting analysis (e.g., extreme gradient boosting (XGB) analysis). Gradient boosting, such as XGB, examines the contribution of each feature's gain to the model. In a boosted tree model, each gain of each feature in each tree is considered, and then the average per feature contribution is evaluated. The highest percentage contributing feature can then be selected.

[0082] After the features are ranked according to the feature importance metric in block 104a of Figure 1, the system uses the rankings to select a subset of features. An exemplary process for selecting a subset of features is described in further detail in Figures 3A and 3B below.

[0083] FIG. 3A illustrates an exemplary feature overlap analysis according to some embodiments. As described above in FIG. 2, multiple feature importance metrics may be used to rank the multiple features. In the example of FIG. 3A, the exemplary process uses ANOVA, random forest, and gradient boosting analysis to rank the features. However, one skilled in the art will appreciate that other learning techniques known in the art may be used as well. However, for the illustrative purposes of FIG. 3A, the ANOVA feature ranking 302 includes features 1, 4, 5, and 8 as the highest ranked features. The random forest ranking 304 includes features 8, 2, 3, and 1 as the highest ranked features. The gradient boosting ranking 306 includes features 6, 1, 4, and 2 as the highest ranked features. In some embodiments, other feature importance metrics may be used to evaluate the features. In some embodiments, fewer or more than three metrics may be used to evaluate the features. In some embodiments, more than four features may be considered top features, for example, any of more than five, more than six, more than seven, more than eight, more than nine, more than ten, more than eleven, more than twelve, more than thirteen, more than fourteen, more than fifteen, more than sixteen, more than seventeen, more than eighteen, more than nineteen, more than twenty, more than twenty-one, more than twenty-two, more than twenty-three, more than twenty-four, or more than twenty-five features may be considered top features.

[0084] Once the features are ranked, the system may perform a feature overlap analysis to determine which features were identified as top features by one or more metrics. In the example of FIG. 3A, the feature overlap analysis 308 identifies feature 1 as the top feature identified in the ANOVA feature ranking 302, the random forest ranking 304, and the gradient boosting ranking 306. The feature overlap analysis 308 also identifies features 2, 4, and 8 as the top features identified by two metrics. In some embodiments, the feature overlap analysis 308 may output a subset of the features by outputting the features identified as top by all metrics. In some embodiments, the feature overlap analysis 308 may output a subset of the features by outputting the features identified as top by one or more metrics. In some embodiments, the feature overlap analysis 308 may be represented graphically. In some embodiments, the feature overlap analysis 308 may output a list including the subset of the features.

[0085] FIG. 3B illustrates an example output 310 of a feature selection process for features used to classify a subject's cancer tumor as HRD positive or HRD negative, according to some embodiments. Feature importance rankings 312 are illustrated in graphs, with each graph showing a ranking of features by a particular feature importance metric. In each graph (ANOVA, Random Forest, and Gradient Boosting), each dot represents a feature whose y-axis value corresponds to its feature importance as calculated by the feature importance metric. In the example of FIG. 3B, feature overlap analysis 314 can include top features according to each feature importance metric. As illustrated, feature overlap analysis can identify features that are highly ranked by all of the metrics and / or a portion of the metrics.

[0086] Returning to Figure 1, in some embodiments, the system and method may determine a subset of the multiple features using an iterative feature selection process 104b in addition to or instead of process 104a. At block 104b, the system evaluates the features using one or more feature importance metrics (e.g., gradient boosting), as described below in Figure 4, and then performs an iterative feature selection process to gradually expand the feature set.

[0087] Figure 4 illustrates an iterative feature selection process that may be used by block 104b of Figure 1, according to some embodiments. In block 402, the system receives a data set having a plurality of features (e.g., the plurality of features received in block 102 of Figure 1).

[0088] 4, the system evaluates the features received in block 402 using one or more feature importance metrics (e.g., gradient boosting). The system can then rank the features according to their corresponding feature importance metric scores.

[0089] In block 408 of FIG. 4, the system and method obtain a new feature set. In a first iteration, the system may obtain the new feature set by including the highest ranked feature determined by block 404 in the feature set. In a subsequent iteration, the system may extend the existing feature set by adding the next highest ranked feature determined by block 404 to obtain the new feature set. The system further obtains a training dataset based on the new feature set. The training dataset may include multiple data elements, each data element including data related to the new feature set and a corresponding classification label (e.g., HRD positive or HRD negative). For example, the data element may include data related to the features of the new feature set from the sample and the corresponding classification label (e.g., HRD positive or HRD negative) of the sample.

[0090] In block 410 of Figure 4, the system and method train and evaluate a new classification model using the training dataset from block 408. The system records the performance of the model in relation to the list of features used in training and evaluating the model. In some embodiments, the training and evaluation of the classification model can be performed using a cross-validation method, as further described below with reference to Figures 6A and 6B. In some embodiments, the training and evaluation of the classification model can use separate subsets of the dataset from block 408.

[0091] In some embodiments, blocks 408 and 410 of FIG. 4 are repeated until all features received in block 402 are included in the data. In each iteration, block 408 adds the next highest ranked feature to the dataset. For example, in the first iteration, block 408 outputs a feature set including the highest ranked feature and a corresponding training set. In the second iteration, block 408 outputs a feature set including the two highest ranked features and a corresponding training set. In the third iteration, block 408 outputs a feature set including the three highest ranked features and a corresponding training set, and so on. In each iteration, block 410 then uses the training dataset from block 406 to train and evaluate a new classification model. The system repeats blocks 408 and 410 until a condition is met. In some embodiments, the condition includes block 412, where the system determines that there are no more features to be added (e.g., all features received in block 402 are included in the dataset used to train and evaluate the classification model in block 410). In some embodiments, the condition includes a determination that the performance of the new classification model exceeds a threshold. This iterative process allows the system to record the performance of the classification model as it is trained and evaluated with the highest ranked feature, the top two highest ranked features, the top three highest ranked features, etc., until all features received in block 402 have been used to train the classification model and evaluate its performance. An example of recorded performance data is shown in FIG.

[0092] In block 414 of Figure 4, the system and method utilize the recorded model performance from block 410 to determine a minimum subset of features that optimizes the performance of the classification model. In some embodiments, the system may determine the minimum subset of features such that the addition of additional features does not substantially improve the performance of the model. In some embodiments, the system may determine the minimum subset of features such that the performance of the classification model exceeds a certain predetermined threshold. The subset of features is output in block 414.

[0093] FIG. 5 illustrates an exemplary plot of the performance of the model determined in block 410 of FIG. 4. In the example illustrated in FIG. 5, the horizontal axis indicates the number of top features included in the data used to train and evaluate the classification model. The vertical axis indicates the performance of the model. In some embodiments, the performance of the model can be evaluated using the area under the receiver operating characteristic (ROC) curve (AUC). In the example of FIG. 5, block 416 can determine that the 26 highest ranked features are output as a subset of features, but a smaller number of features can be selected based on the change in the relative increase in model performance with each added feature.

[0094] FIG. 6A illustrates an exemplary cross-validation process that may be used to evaluate the performance of a model, according to some embodiments. In some embodiments, the process 600 in block 410 of FIG. 4 may be used to evaluate the performance of a model. In block 602, the system may receive a plurality of data elements. Each of the plurality of data elements may include one or more features and a known classification label. In block 604, the system divides the plurality of data elements from block 602 into n equally sized subsets. In block 606, the system holds out one of the subsets from block 604 as a "hold-out" set. In block 608, the system trains a model on all data elements that are not held out (e.g., data elements from the n-1 subsets that are not in the "hold-out" set). In block 610, the system uses the features of the data elements from the "hold-out" set as input to the model from block 608. The model generates a plurality of predicted classification labels corresponding to the features of the data elements. The predicted classification labels are then compared to the known classification labels of the "hold-out" set to evaluate the performance of the model on the "hold-out" set. Blocks 606, 608, and 610 are repeated until all n subsets from block 604 have been used once as a "holdout" set. That is, blocks 606, 608, and 610 are repeated n times, with a different subset being used as the "holdout" set for each iteration. Finally, in step 612, the performances from all n iterations of block 610 are averaged and the average performance is output.

[0095] FIG. 6B illustrates an example division of the plurality of data elements into five equal-sized subsets according to some embodiments. FIG. 6B may be an example of FIG. 6A where n=5. The plurality of data elements 622 may be an example of the plurality of data elements from block 602 of FIG. 6A. In the example of FIG. 6B, the plurality of data elements 622 are divided into Set 1, Set 2, Set 3, Set 4, and Set 5. In iteration 1 623, in the plurality of data elements 622, Set 1 may be used as a "hold-out" data set as described by block 606. A model may be trained on Set 2, Set 3, Set 4, and Set 5 as described by block 608. The performance of the model may then be evaluated on the "hold-out" data set 1. This process is then repeated four more times. In iteration 2 624, Set 2 is the "hold-out" set, the model is trained on Set 1, Set 3, Set 4, and Set 5, and the performance of the model is evaluated on Set 2. In iteration 3 626, set 3 is the "hold out" set, the model is trained on sets 1, 2, 4, and 5, and the model's performance is evaluated on set 3. In iteration 4 628, set 4 is the "hold out" set, the model is trained on sets 1, 2, 3, and 5, and the model's performance is evaluated on set 4. In iteration 5 630, set 5 is the "hold out" set, the model is trained on sets 1, 2, 3, and 4, and the model's performance is evaluated on set 5. In the example of FIG. 6B, the average performance may be the average of the model's performance from iteration 1 622, iteration 2 624, iteration 3 626, iteration 4 628, and iteration 5 630.

[0096] Returning to FIG. 1 , in block 106, the system obtains a subset of selected features, as determined by feature selection in block 104. A classification model 108 is trained using information from the selected features 106 and the labeled training data 110. In some embodiments, the dataset used for feature selection 104 is the same dataset that is the labeled training data 110. In some embodiments, the dataset used for feature selection 104 is a different dataset than the labeled training data 110. The process of training a classification model is described below in the following sections and in FIG. 7. Once the classification model 108 is trained, features from an unknown tumor of the subject's cancer (e.g., data elements not included in the data received in block 102 and not associated with a known classification label) can be input into the model 108 to predict whether the tumor of the subject's cancer is likely to be HRD positive or HRD negative.

[0097] Data characteristics A test sample from a tumor that has been identified (e.g., classified) can be obtained from a subject. Features, such as basic features, copy number features, and / or short variant features, associated with the test sample include one or more features that can be used as inputs for an HRD classification model. The HRD classification model is trained based on HRD-positive data associated with HRD-positive samples (e.g., tumor samples) and corresponding features (e.g., basic features, copy number features, and / or short variant features) from HRD-negative data associated with HRD-negative samples (e.g., tumor samples). The features can be used as functional readouts of HRD that can help identify tumors with a "BRCAness" profile associated with HRD. Tumors with such an HRD-positive phenotype can be suitable candidates for certain drug therapies that are not (or are less likely to be) effective in the HRD-negative phenotype.

[0098] Copy number features may include, but are not limited to, segment size features, number of sequencing reads features, number of sequencing reads features, absolute copy number features, number of breakpoints per x megabase features, change point copy number features, segment copy number features, number of breakpoints per chromosome arm features, and number of segments of oscillatory copy number. See Macintyre et al.,Copy-number signatures and mutational processes in ovarian carcinoma, Nat. Genet. 2018 Sep;50(9):1262-1270. Mixture modeling can be applied to separate each feature distribution into a mixture of Gaussian distributions or a mixture of Poisson distributions to achieve floating or binary component features. Copy number features may also include segment minor allele frequency features based on the A and B allele frequencies of germline SNPs in the segment.

[0099] In some embodiments, the HRD model (e.g., the HRD classifier model) can be trained using more features than are used as input. For example, the HRD classification model can be trained based on HRD-positive data and HRD-negative data, each of which includes a number of features associated with HRD-positive and / or HRD-negative tumors. The data input to the HRD classification model can then include fewer features. The HRD classifier model, in one example, can adjust the weights of the data features omitted from the sample data input to the trained HRD classifier model. In addition, the HRD classifier model can be trained using additional data features (e.g., measures of genome-wide loss of heterozygosity and / or one or more short variant features, each as described herein, etc.), although in some embodiments, the data input can only include one or more copy number features associated with the genome of the tumor associated with the cancer of interest.

[0100] Sequencing data is collected by sequencing at least a portion of at least one genome of the tumor to obtain genomic data features including copy number features, basic features including gLOH and tumor genome ploidy measurements, and / or short variant features. Absolute or relative copy number and segmentation can then be derived from whole genome sequencing data, such as shallow hole genome sequencing (sWGS) data. Circular binary segmentation (CBS) can also be used to divide the genome into segments of constant total copy number based on DNA microarray data, from which copy number features can be derived. Alternatively, absolute copy number and segmentation can be derived from any technique known in the art, including but not limited to exome sequencing (ES) or SNP arrays. Distributions of copy number features can be calculated from absolute copy number data, such as WGS data. Mixture modeling can be applied to divide each feature distribution into a mixture of Gaussian distributions or a mixture of Poisson distributions to achieve floating or binary component features. Thus, a particular "copy number feature" used to train an HRD classification model or to be input to a trained HRD classification model is expressed as its component features. For example, in the case of a copy number feature of segment size, when divided into z components, there are then z possible features that can be used to train an HRD classification model or to run an HRD classification model. In other words, for a particular test sample, a "copy number feature" in the category of "segment size" (assuming that the segment size is divided into z components) has z possible inputs, whether to train or run an HRD classification model. If z is equal to 3, at least one of the three segment size features can be input to the HRD classification model, i.e., segsize1, segsize2, or segsize3. Optimal model performance may depend, in part, on the number of component features selected for each particular category of features.However, a particular category of features can be split into any suitable number of component features and does not necessarily correspond to a particular probability distribution, and thus the model can work well and be efficiently validated with a greater or lesser number of component features, even if performance is not optimal.

[0101] When deriving copy number signatures, absolute copy number data can first be normalized by matching with a normal data set to determine the baseline level for calling copy number variant events.Normal panels are typically derived from healthy tissue samples (can be derived from the same individual as the tumor is derived).Analysis of healthy tissue samples allows to set the baseline copy number for deriving copy number signatures described herein.

[0102] Some of the copy number features described can be evaluated across subregions of the genome. For example, a particular copy number feature can be evaluated across the centromere portion of the genome. In another example, a copy number feature can be evaluated across the telomere portion of the genome. In yet a further example, a copy number feature can be evaluated across both the telomere and centromere portions of the genome. In an exemplary method, to define the telomere and centromere portions of the genome, a human reference sequence genome such as hg19 can be used to define the beginning and end of each chromosome arm. The length of a particular arm is then divided by 2 to define a midpoint. For each region analyzed for copy number features, the segment on the centromere side of this midpoint is defined as the centromere segment. The segment on the telomere side of this midpoint is defined as the telomere segment. If a segment spans the midpoint (e.g., a segment that starts centromeric and ends telomeric of the midpoint), the segment may be referred to as both a centromere and a telomere and may be used to assess both telomere and centromere copy number features. Thus, any of the data features described herein may be assessed across the telomeric region of the genome, the centromeric region of the genome, or both the telomeric and centromeric regions of the genome, as appropriate.

[0103] Copy number modeling may be influenced by the estimated base ploidy of the genome being evaluated. If the base ploidy is estimated higher, the floating point copy number feature may be shifted to the right, resulting in skewed component scores and ultimately erroneous classification. Normalizing copy number data to base ploidy involves dividing the copy number data by the average ploidy of the genome being evaluated. Thus, any of the copy number features described may be derived from ploidy-normalized copy number data, where the absolute copy number is normalized to the average ploidy of the genome of the tumor. An exemplary method for calculating the average ploidy is to obtain a weighted average copy number for all segments of the sample. For an exemplary method of calculating the average ploidy, see Sun et al., A computational approach to distinguish somatic vs. germline origin of genomic alterations from deep sequencing of cancer specimens without a matched normal, PLoS Comput.Biol.2018 Feb 7;14(2):e1005965.

[0104] The features described herein may, in some embodiments, be binned features. Feature binning involves organizing certain values ​​into certain categorical bins. For example, for a feature with values ​​ranging from 0 to 10, quartile binning may organize each of these values ​​from 0 to 10 into one of four bins, with lower values ​​being organized into lower bins and higher values ​​being organized into higher bins. In some embodiments, the binning is unsupervised. In some embodiments, the binning is supervised. In equal width binning, the bins have approximately the same width range. For example, for a feature with values ​​from 1 to 8, equal width binning with four bins would organize values ​​of 1 and 2 into a first bin, values ​​of 3 and 4 into a second bin, and so on. In some embodiments, the binning is equal frequency binning. In equal frequency binning, the bins are organized such that each bin has approximately the same number of values, and the values ​​are approximately equally distributed among the bins. For example, for a feature with values ​​from 1 to 10, with lower values ​​being much more frequent, the binning could be organized with 1 in the first bin, 2 in the second bin, and 3 to 10 in the third bin. The binning could be 2nd quantiles, 3rd quantiles, 4th quantiles, 5th quantiles, 6th quantiles, 7th quantiles, or any other suitable binning organization.

[0105] In some embodiments of any of the described methods, the copy number feature includes a segment size feature. The segment size is derived from the length in genomic bases of each copy number segment across the genome. For example, if a segment has a copy number of x and the next segment has a copy number of y, the length of the segment with copy number x and the length of the segment with copy number y are factors of the copy number category of segment size. In an exemplary embodiment, the segment size distribution is divided into 10 component features. Lower numbered segment size features represent smaller segment sizes (e.g., segsize1), while higher numbered segment size features represent larger segment sizes (e.g., segsize10). In some embodiments, the segment size distribution is divided into at least 5 component features, such as at least 6, at least 7, at least 8, at least 9, at least 10, or at least 11 component features. In some embodiments, the segment size distribution is divided into any of 5, 6, 7, 8, 9, 10, or 11 component features. In some embodiments, the segment size feature is evaluated across the telomeric portion of the genome. In some embodiments, the segment size features are evaluated across a centromeric portion of the genome. In some embodiments, the segment size features are evaluated across both telomeric and centromeric portions of the genome. In some embodiments, the segment size features are evaluated across the entire genome. In some embodiments, the segment size features are derived from ploidy-normalized copy number data. In some embodiments, the segment size features are binned features.

[0106] In some embodiments of any of the described methods, the copy number feature comprises a feature of x number of breakpoints per megabase. In some embodiments, x is between about 1 megabase (MB) and about 150 megabases. In some embodiments, x is any of about 10 MB, about 25 MB, about 50 MB, about 100 MB, and about 150 MB. The number of breakpoints per section represents the number of breakpoints per section across a genome or a portion of a genome. For example, for the number of breakpoints per 10 MB, a processing contiguous window (or alternatively, a sliding window) of 10 MB may be analyzed across the entire genome, and then the number of breakpoints for each frame of the sliding window may be evaluated. In this approach, a contiguous window was used, but it should be noted that a sliding window or any other technique suitable for evaluating the number of breakpoints may be used. Nevertheless, in some exemplary embodiments, the number of breakpoints per 1x megabase is divided into three component features. Lower numbered breakpoint count features represent fewer breakpoints (e.g., breakpoints per 10MB: bp10MB1 indicates fewer breakpoints per frame of a 10MB sliding window or per frame of a 10MB processing adjacent window), while higher numbered features represent more breakpoints per section (e.g., breakpoints per 10MB: bp10MB3 indicates more breakpoints per frame of a 10MB sliding window compared to a lower numbered feature such as bp10MB1). In some embodiments, the distribution of breakpoint counts is split into at least two component features, such as at least three or at least four component features. In some embodiments, the number of breakpoints per section is split into either two, three, four, or five component features. In some embodiments, the number of breakpoints per x megabase feature is evaluated across a telomeric portion of the genome.In some embodiments, the feature of the number of breakpoints per x megabases is evaluated over the centromere portion of the genome. In some embodiments, the feature of the number of breakpoints per x megabases is evaluated over the entire genome. In some embodiments, the feature of the number of breakpoints per x megabases is derived from ploidy-normalized copy number data. In some embodiments, the feature of the number of breakpoints per x megabases is a binned feature.

[0107] In some embodiments of any of the described methods, the copy number feature comprises a feature of the number of sequencing reads obtained from sequencing the genome segment. For a particular genome segment, this value refers to the average number of sequencing reads that align (i.e., "cover") to the sequenced segment. For genome segments with abnormally high copy numbers, the number of sequencing reads increases. In contrast, for genome segments that have lost copy numbers (such as homozygous deletions), there will be fewer sequencing reads. The sequencing read feature can be expressed as an actual number of reads (such as the average of the reads for each analyzed segment) or as a bin of sequencing reads. A lower numbered sequencing read feature represents a lower absolute sequencing read, while a higher numbered sequencing read feature represents a higher absolute sequencing read. In some embodiments, the sequencing read feature is evaluated over a telomeric portion of the genome. In some embodiments, the sequencing read feature is evaluated over a centromeric portion of the genome. In some embodiments, the feature of sequencing reads is evaluated across both telomeric and centromeric parts of the genome. In some embodiments, the feature of sequencing reads is derived from ploidy normalized data. In some embodiments, the feature of sequencing reads is a binned feature. In some embodiments, the feature of the number of sequencing reads is a measurement of the number of reads from next generation sequencing (NGS). In some embodiments, the feature of the number of sequencing reads is expressed as the ratio of the number of sequencing reads for a genome segment of a tumor sample compared to the number of sequencing reads for that genome segment in a control.

[0108] In some embodiments of any of the described methods, the copy number feature comprises an absolute copy number feature. An absolute copy number may be calculated for each genome segment and assigned a value. For example, the assigned value may include 0 (indicating homozygous deletion), 1 (may indicate heterozygous deletion), 2 (may be a normal count) or more (may indicate copy number amplification). The absolute copy number feature may represent an actual copy number count (such as an average of the copy numbers for each segment analyzed) or a bin of copy number values. For example, a copy number of at least 6 may be binned as representing a high copy number for the segment. A copy number of 3 to 5 may be binned as representing a moderately increased copy number. A copy number of 1 and 2 may be normal and a copy number of 0 may be binned as a homozygous deletion. A low numbered absolute copy number feature represents a low absolute copy number and a high numbered absolute copy number feature represents a high absolute copy number. In some embodiments, the absolute copy number is divided into any of 3, 4, 5, 6, 7, 8 or 9 component features. In some embodiments, absolute copy number features are assessed across a telomeric portion of the genome. In some embodiments, absolute copy number features are assessed across a centromeric portion of the genome. In some embodiments, absolute copy number features are assessed across both telomeric and centromeric portions of the genome. In some embodiments, absolute copy number features are derived from fold-normalized data. In some embodiments, absolute copy number features are binned features.

[0109] In some embodiments of any of the described methods, the copy number features include change point copy number features. Change point copy number refers to the absolute difference in copy number between genome segments across the genome. For example, adjacent segments modeled with copy number 7 and 2 have an absolute difference of 5. In exemplary embodiments, the change point copy number distribution is split into 7 component features. Lower numbered change point copy number features represent smaller absolute differences in copy number change (e.g., change point 1), while higher numbered features represent larger absolute differences in copy number change (e.g., change point 7). In some embodiments, the change point copy number distribution is split into at least 4 component features, such as at least 5, at least 6, at least 7, or at least 8 component features. In some embodiments, the change point copy number is split into any of 3, 4, 5, 6, 7, 8, or 9 component features. In some embodiments, the change point copy number features are evaluated across a telomeric portion of the genome. In some embodiments, the change point copy number features are evaluated across a centromeric portion of the genome. In some embodiments, the changepoint copy number features are evaluated across both telomeric and centromeric portions of the genome. In some embodiments, the changepoint copy number features are derived from ploidy-normalized copy number data. In some embodiments, the changepoint copy number features are binned features.

[0110] In some embodiments of any of the described methods, the copy number features include segment copy number features. The segment copy number is derived from the copy number of each segment across the genome or a portion of the genome. In an exemplary embodiment, the segment copy number distribution is divided into eight component features. Lower numbered segment copy number features represent lower copy numbers (e.g., copy number 1 can represent copy number levels of 0 or 1, or 0 to 1), while higher numbered copy number features represent higher copy numbers (e.g., copy number 8). In some embodiments, the segment copy number distribution is divided into at least four component features, such as at least five, at least six, at least seven, at least eight, or at least nine component features. In some embodiments, the segment copy number distribution is divided into any of four, five, six, seven, eight, nine, or ten component features. In some embodiments, the segment copy number features are evaluated across a telomeric portion of the genome. In some embodiments, the segment copy number features are evaluated across a centromeric portion of the genome. In some embodiments, the segment copy number features are evaluated across the entire genome. In some embodiments, the segment copy number features are derived from ploidy-normalized copy number data. In some embodiments, the segment copy number features are binned features.

[0111] In some embodiments of any of the described methods, the copy number feature comprises a breakpoint number per chromosome arm feature. In an exemplary embodiment, the distribution of breakpoint number per chromosome arm is divided into five component features. Lower numbered breakpoint number per chromosome arm features represent fewer breakpoints per arm (e.g., bpchrarm1), while higher numbered breakpoint number per chromosome arm features represent more breakpoints per chromosome arm (e.g., bpchrarm5). In some embodiments, the distribution of breakpoint number per chromosome arm is divided into at least three component features, such as at least four component features, at least five component features, at least six component features, or at least seven component features. In some embodiments, the distribution of breakpoint number per chromosome arm is divided into any of four, five, six, seven, or eight component features. In some embodiments, the number of breakpoints per chromosome arm is derived from ploidy-normalized copy number data. In some embodiments, the breakpoint number per chromosome arm feature is a binned feature.

[0112] In some embodiments, the copy number feature includes several segments with an oscillating copy number segment number (osCN) feature. The oscillating copy number segment number represents a cross section of a genome or a portion of a genome that counts the number of alternating segments repeated between two copy numbers. In an exemplary embodiment, the oscillating copy number segment number distribution is divided into three component features. The lower numbered oscillating copy number segment number features represent fewer repeated changes between the two copy numbers (e.g., osCN1), while the higher numbered oscillating copy number segment number features represent more repeated changes between the two copy numbers (e.g., osCN3). In some embodiments, the oscillating copy number segment number distribution is divided into at least two, e.g., at least three, or at least four component features. In some embodiments, the oscillating copy number segment number distribution is divided into either two, three, four, or five component features. In some embodiments, the oscillating copy number segment number features are evaluated across the telomeric portion of the genome. In some embodiments, the oscillating copy number segment number feature is evaluated across a centromeric portion of the genome. In some embodiments, the oscillating copy number segment number feature is evaluated across the entire genome. In some embodiments, the oscillating copy number segment number feature is derived from ploidy-normalized copy number data. In some embodiments, the oscillating copy number segment number feature is a binned feature.

[0113] In some embodiments, the copy number feature includes a segment minor allele frequency (segMAF) feature. The segMAF feature can be derived from either the average segMAF of the tumor genome, or the median segMAF. In a normal genome at a heterozygous allele site, the expected copy number of each allele is 1.0. HRD is associated with a complete loss of allele (loss of heterozygosity) or an increase in copy number of one allele relative to the other. Thus, segMAF is a cross-sectional view of the genome by segment that compares the ratio of minor alleles to major alleles. Specifically, each heterozygous SNP is analyzed for A allele and B allele frequency. The frequency of minor alleles is obtained as minor allele fraction. A balanced locus has a ratio of about 0.5:0.5, with a minor allele frequency of 0.5. A loss of heterozygosity event causes imbalance and skew of minor allele frequency to less than about 0.5 for the minor allele fraction. In some embodiments, the segMAF signature is evaluated across a telomeric portion of the genome. In some embodiments, the segMAF signature is evaluated across a centromeric portion of the genome. In some embodiments, the segMAF signature is evaluated across the entire genome. In some embodiments, the segment minor allele frequency signature is a binned signature.

[0114] The HRD classification model is trained by HRD-positive data, which includes one or more features and HRD-positive markers associated with the HRD-positive tumor for each of the plurality of HRD-positive tumors, and HRD-negative data, which includes one or more copy number features and HRD-negative markers associated with the HRD-negative tumor for each of the plurality of HRD-negative training tumors. The HRD classification model may also be trained based on other features or measures. Thus, test data including these other features or measures (including one or in combination with copy number features) can be input into the HRD classification model. For example, basic features including a measure of genomic loss of heterozygosity and / or one or more short variant features can be used in the HRD classification model (to train the HRD classification model or as test data input into the HRD classification model).

[0115] In some embodiments, the basic characteristics include the age of the subject from which the tumor was obtained. The patient may be any age, including at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, or at least 80 years old. The age characteristic may be an integer value of the subject. Alternatively, the age characteristic may be a qualitative characteristic, such as either an infant, a juvenile, a child, a young adult, or an elderly subject. In some embodiments, the age characteristic is a binned characteristic.

[0116] In some embodiments, the basic characteristic comprises a cancer type characteristic. The cancer type characteristic refers to the origin of the tumor. The cancer type may comprise, for example, one of adrenal, bile duct, bone / soft tissue, breast, colon / rectum, esophagus, eye, head and neck, kidney, liver, lung, lymphatic, medulloblastoma, mesothelioma, myeloid, nervous system, neuroendocrine, ovarian, pancreatic, prostate, skin, stomach, testicular, thymus, thyroid, urinary tract, uterus, or vulvar cancer. In some embodiments, the cancer type characteristic is a binned characteristic.

[0117] In some embodiments, the basic features include cancer stage features. Cancer stage is often based on cancer type (e.g., pancreatic cancer stage, prostate cancer stage, breast cancer stage, ovarian cancer stage, etc.), but universal staging systems are also known in the art. Any suitable cancer stage system may be used and may depend, for example, on tumor location, cell type, tumor size, tumor spread and distribution, tumor metastasis, and tumor grade. As a data feature, cancer stage is typically expressed as a range from less severe to more severe stages. For example, for a cancer stage feature that includes four component features, stage 1 may indicate early stage cancer, and stage 4 may indicate late stage cancer. In some embodiments, the cancer stage feature is a binned feature.

[0118] The HRD positive and HRD negative data are typically split into a training dataset, a validation dataset, and / or a test dataset. During training, the HRD classification model is provided with only the training set. Optionally, the training set may be balanced. Once trained, the model may be validated and adjusted by its performance on the validation set. If the model shows overfitting on the validation set, training may be adjusted and repeated. Once trained, and optionally after validation, the trained model may be evaluated using the test dataset.

[0119] A measure of genomic loss of heterozygosity (gLOH) (e.g., genome-wide loss of heterozygosity or exome-wide loss of heterozygosity) may be included as a basic feature in some embodiments. Whole exome sequencing or targeted sequencing across a sufficiently large portion of the genome may be interpreted as a proxy for genomic loss of heterozygosity, so it is not necessary to analyze the whole genome to determine genomic loss of heterozygosity. In some embodiments, gLOH is encoded as a continuous numeric feature. In some embodiments, gLOH is encoded as a categorical feature, for example, if gLOH is above or below a predetermined threshold. The predetermined threshold may be set, for example, at about 10% or more, about 12% or more, about 14% or more, or about 16% or more. The predetermined threshold may be set, for example, at about 16%. gLOH can be determined, for example, using the method described in Swisher et al., Rucaparib in relapsed, platinum-sensitive high-grade ovarian carcinoma (ARIEL2 Part 1): an international, multicenter, open-label, phase 2 trial, Lancet Oncology, vol. 18, no. 1, pp. 75-87 (2017).

[0120] One or more short variant features can be used in the HRD classification model (to train the HRD classification model and / or as test data input to the HRD classification model). These short variant features can include, but are not limited to, one or more deletions (e.g., deletions of at least 5 base pairs, etc.) in repeat or microhomology region features and / or mutation signatures incorporating two or more short variant features. These short variant features can be identified, in an exemplary method, by comparing sequencing data corresponding to tumor samples with a consensus human genome sequence (e.g., hg19). In some embodiments, the short variant features are binned features.

[0121] Multiple short variant features can be combined and expressed as a mutational signature score. For example, one or more short variant features can include a mutation profile, such as from the COSMIC cancer database. In one example, one or more short variant features include an indel-based signature, such as the COSMIC ID6 or COSMIC ID8 indel signature of the COSMIC cancer database. Sample profiles can be mapped to these COSMIC profiles, for example, using NNMF methodology. In another example, one or more short variant features include the COSMIC ID8 of the COSMIC cancer database. In yet another example, one or more short variant features include the SBS3 mutation signature of the COSMIC cancer database. For a summary of exemplary COSMIC ID signatures, see Alexandrov et al., The repertoire of mutational signatures in human cancer, Nature 2020;578(7793):94-101. See also Forbes et al., COSMIC: mining complete cancer genomes in the Catalogue of Somatic Mutations in Cancer, Nuc. Acids Res. 2011 Jan;39:D945-D950.

[0122] In some embodiments, the one or more short variant features include deletions of microhomology or repeat regions. In some embodiments, the deletion is at least 1 base pair. In some embodiments, the deletion is at least 5 base pairs. Deletions in microhomology regions are a characteristic result of microhomology-mediated end joining (MMEJ), which occurs in the absence of homologous recombination. In this process, short similar regions (microhomologies) are used to guide the repair of double-strand breaks in the genome. A distinguishing characteristic of these deletions is that the 3' end of the deleted sequence shares similarity with the situation upstream of the deletion. Thus, the feature of deletions in microhomology regions is a measure of the number of deletions that exhibit this behavior, and can also be based on the length of the microhomology (i.e., more deletion pairs with longer lengths, fewer deletions with shorter lengths).

[0123] In an exemplary embodiment, the test data includes a segment minor allele frequency feature and a segment size feature. In some embodiments, the segment minor allele frequency feature is a binned feature. In some embodiments, the segment size feature is a binned feature. The test data may further include at least one of a number of breakpoints per x megabase feature, a changepoint copy number feature, a number of sequencing reads feature, an absolute copy number feature, a segment copy number feature, a number of breakpoints per chromosome arm feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0124] In another exemplary embodiment, the test data includes a segment minor allele frequency feature and a breakpoint number per x megabase feature. In some embodiments, the segment minor allele frequency feature is a binned feature. In some embodiments, the breakpoint number per x megabase feature is a binned feature. The test data may further include at least one of a segment size feature, a number of sequencing reads feature, an absolute copy number feature, a changepoint copy number feature, a segment copy number feature, a number of breakpoints per chromosome arm feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0125] In another exemplary embodiment, the test data includes a segment minor allele frequency feature and a change point copy number feature. In some embodiments, the segment minor allele frequency feature is a binned feature. In some embodiments, the change point copy number feature is a binned feature. The test data may further include at least one of a segment size feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a segment copy number feature, a number of breakpoints per chromosome arm feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0126] In another exemplary embodiment, the test data includes a segment minor allele frequency feature and a segment copy number feature. In some embodiments, the segment minor allele frequency feature is a binned feature. In some embodiments, the segment copy number feature is a binned feature. The test data may further include at least one of a segment size feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a changepoint copy number feature, a number of breakpoints per chromosome arm feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0127] In another exemplary embodiment, the test data includes a segment minor allele frequency feature and a breakpoint number per chromosome arm feature. In some embodiments, the segment minor allele frequency feature is a binned feature. In some embodiments, the breakpoint number per chromosome arm feature is a binned feature. The test data may further include at least one of a segment size feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a changepoint copy number feature, a segment copy number feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0128] In another exemplary embodiment, the test data includes a segment minor allele frequency feature and a number of segments of oscillating copy number feature. In some embodiments, the segment minor allele frequency feature is a binned feature. In some embodiments, the number of segments of oscillating copy number feature is a binned feature. The test data may further include at least one of a segment size feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a changepoint copy number feature, a segment copy number feature, and a number of breakpoints per chromosome arm feature. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0129] In another exemplary embodiment, the test data includes a segment size feature and a number of breakpoints per x megabase feature. In some embodiments, the segment size feature is a binned feature. In some embodiments, the number of breakpoints per x megabase feature is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a changepoint copy number feature, a segment copy number feature, a number of breakpoints per chromosome arm feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0130] In another exemplary embodiment, the test data includes a segment size feature and a change point copy number feature. In some embodiments, the segment size feature is a binned feature. In some embodiments, the change point copy number feature is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a segment copy number feature, a number of breakpoints per chromosome arm feature, and a number of segments of oscillating copy number feature. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0131] In another exemplary embodiment, the test data includes a segment size feature and a segment copy number feature. In some embodiments, the segment size feature is a binned feature. In some embodiments, the segment copy number is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a changepoint copy number feature, a number of breakpoints per chromosome arm feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0132] In another exemplary embodiment, the test data includes a segment size feature and a breakpoint number per chromosome arm feature. In some embodiments, the segment size feature is a binned feature. In some embodiments, the breakpoint number per chromosome arm feature is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a changepoint copy number feature, a segment copy number feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0133] In another exemplary embodiment, the test data includes a segment size feature and a number of segments of oscillating copy number feature. In some embodiments, the segment size feature is a binned feature. In some embodiments, the number of segments of oscillating copy number feature is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a number of breakpoints per x megabase feature, a changepoint copy number feature, a segment copy number feature, and a number of breakpoints per chromosome arm feature. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0134] In another exemplary embodiment, the test data includes a feature of number of breakpoints per x megabases and a feature of copy number of changepoints. In some embodiments, the feature of number of breakpoints per x megabases is a binned feature. In some embodiments, the feature of copy number of changepoints is a binned feature. The test data may further include at least one of a feature of segment minor allele frequency (segMAF), a feature of number of sequencing reads, a feature of absolute copy number, a feature of segment size, a feature of segment copy number, a feature of number of breakpoints per chromosome arm, and a feature of number of segments of oscillating copy number. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a feature of cancer type, a feature of cancer stage, a feature of tumor purity, and a feature of tumor genome ploidy.

[0135] In another exemplary embodiment, the test data includes a feature of number of breakpoints per x megabases and a feature of segment copy number. In some embodiments, the feature of number of breakpoints per x megabases is a binned feature. In some embodiments, the feature of segment copy number is a binned feature. The test data may further include at least one of a feature of segment minor allele frequency (segMAF), a feature of number of sequencing reads, a feature of absolute copy number, a feature of segment size, a feature of changepoint copy number, a feature of number of breakpoints per chromosome arm, and a feature of number of segments of oscillating copy number. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a feature of cancer type, a feature of cancer stage, a feature of tumor purity, and a feature of tumor genome ploidy.

[0136] In another exemplary embodiment, the test data includes a feature of number of breakpoints per x megabases and a feature of number of breakpoints per chromosome arm. In some embodiments, the feature of number of breakpoints per x megabases is a binned feature. In some embodiments, the feature of number of breakpoints per chromosome arm is a binned feature. The test data may further include at least one of a feature of segment minor allele frequency (segMAF), a feature of number of sequencing reads, a feature of absolute copy number, a feature of segment size, a feature of change point copy number, a feature of segment copy number, and a feature of number of segments of oscillating copy number. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a feature of cancer type, a feature of cancer stage, a feature of tumor purity, and a feature of tumor genome ploidy.

[0137] In another exemplary embodiment, the test data includes a feature of number of breakpoints per x megabases and a feature of number of segments of oscillating copy number. In some embodiments, the feature of number of breakpoints per x megabases is a binned feature. In some embodiments, the feature of number of segments of oscillating copy number is a binned feature. The test data may further include at least one of a feature of segment minor allele frequency (segMAF), a feature of number of sequencing reads, a feature of absolute copy number, a feature of segment size, a feature of changepoint copy number, a feature of segment copy number, and a feature of number of breakpoints per chromosome arm. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a feature of cancer type, a feature of cancer stage, a feature of tumor purity, and a feature of tumor genome ploidy.

[0138] In another exemplary embodiment, the test data includes a changepoint copy number feature and a segment copy number feature. In some embodiments, the changepoint copy number feature is a binned feature. In some embodiments, the segment copy number feature is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a segment size feature, a number of breakpoints per x megabase feature, a number of breakpoints per chromosome arm feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0139] In another exemplary embodiment, the test data includes a changepoint copy number feature and a breakpoint number per chromosome arm feature. In some embodiments, the changepoint number feature is a binned feature. In some embodiments, the breakpoint number per chromosome arm feature is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a segment size feature, a number of breakpoints per x megabase feature, a segment copy number feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0140] In another exemplary embodiment, the test data includes a change point copy number feature and a number of segments of oscillating copy number features. In some embodiments, the change point copy number feature is a binned feature. In some embodiments, the number of segments of oscillating copy number features is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a segment size feature, a number of breakpoints per x megabase feature, a segment copy number feature, and a number of breakpoints per chromosome arm feature. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0141] In another exemplary embodiment, the test data includes a segment copy number feature and a breakpoint number per chromosome arm feature. In some embodiments, the segment copy number feature is a binned feature. In some embodiments, the breakpoint number per chromosome arm feature is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a segment size feature, a number of breakpoints per x megabase feature, a changepoint copy number feature, and a number of segments of oscillating copy number features. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0142] In another exemplary embodiment, the test data includes a segment copy number feature and an oscillatory copy number segment number feature. In some embodiments, the segment copy number feature is a binned feature. In some embodiments, the oscillatory copy number segment number feature is a binned feature. The test data may further include at least one of a segment minor allele frequency (segMAF) feature, a number of sequencing reads feature, an absolute copy number feature, a segment size feature, a number of breakpoints per x megabase feature, a changepoint copy number feature, and a number of breakpoints per chromosome arm feature. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following: the age of the subject from whom the test data was obtained, a cancer type feature, a cancer stage feature, a tumor purity feature, and a tumor genome ploidy feature.

[0143] In another exemplary embodiment, the test data includes a feature of the number of breakpoints per chromosome arm and a feature of the number of segments of oscillating copy number. In some embodiments, the feature of the number of breakpoints per chromosome arm is a binned feature. In some embodiments, the feature of the number of segments of oscillating copy number is a binned feature. The test data may further include at least one of a feature of segment minor allele frequency (segMAF), a feature of number of sequencing reads, a feature of absolute copy number, a feature of segment size, a feature of number of breakpoints per x megabase, a feature of changepoint copy number, and a feature of segment copy number. The test data may further include a measure of gLOH and / or one or more short variant features. The test data may further include one or more of the following features: the age of the subject from whom the test data was obtained, a feature of cancer type, a feature of cancer stage, a feature of tumor purity, and a feature of tumor genome ploidy.

[0144] HRD Model Tumors of the subject's cancer are classified using a trained HRD classification model configured to classify the tumor as HRD positive (or likely to be HRD positive) or HRD negative (or likely to be HRD negative). The HRD classification model is trained using HRD-positive data including, for each HRD-positive tumor of the plurality of HRD-positive tumors, one or more data features associated with the HRD-positive tumor (e.g., one or more copy number features and / or one or more short variant features, among other possible features), and an HRD-positive indicator. The HRD classification model is further trained using HRD-negative data including, for each HRD-negative tumor of the plurality of HRD-negative tumors, one or more data features associated with the HRD-negative tumor (e.g., one or more copy number features and / or one or more short variant features, among other possible features), and an HRD-negative indicator. Test data including one or more data features associated with the genome of a tumor of interest (e.g., one or more copy number features and / or one or more short variant features, among other possible features) are input into the trained HRD classification model, which then classifies the tumor as HRD positive (or likely to be HRD positive) or HRD negative (or likely to be HRD negative) based on the test data.

[0145] The models described herein may include one or more machine learning models, one or more non-machine learning models, or any combination thereof. The machine learning models described herein include any computer algorithm that improves automatically through experience and with the use of data. The machine learning models may include supervised models, unsupervised models, semi-supervised models, self-supervised models, and the like. Exemplary machine learning models include, but are not limited to, linear regression, logistic regression, decision trees, SVM, naive Bayes, neural networks, K-means, analysis of variance (ANOVA), chi-square analysis, random forests, dimensionality reduction algorithms, and gradient boosting algorithms (such as XGB). The non-machine learning models may include any computer algorithm that does not necessarily require training and retraining.

[0146] The HRD classifier may be a probabilistic classifier, such as a gradient boosting model. The probabilistic classifier may be configured to calculate the probability that the tumor is HRD positive or HRD negative, such as by outputting an HRD positive likelihood score or an HRD negative likelihood score. Based on the probability output from the HRD classification model, the tumor may be considered to be HRD positive or HRD negative. Optionally, the tumor may be considered ambiguous, for example, if neither the probability that the tumor is HRD positive nor the probability that the tumor is HRD negative exceeds a predetermined probability threshold. The HRD positive data and the HRD negative data may include copy number features and / or short variant features as described herein.

[0147] HRD negative data may include genomes with wild type alleles (i.e., alleles not associated with HRD) in certain HRD-associated genes. For example, in some embodiments, HRD negative data includes data associated with genomes with wild type alleles of one or more genes associated with HRD, including, but not limited to, BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and / or RAD45L. In some embodiments, HRD negative data includes promoter methylation data of one or more genes associated with HRD, including, but not limited to, BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and / or RAD45L. In some embodiments, the HRD-negative data comprises RNA expression data of one or more genes associated with HRD, including, but not limited to, BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and / or RAD45L. In some embodiments, the HRD-negative data comprises data associated with genomes associated with tumors found to be resistant to platinum-based drugs (e.g., chemotherapy) and / or PARP inhibitors. In some embodiments, the HRD-negative data comprises data associated with genomes associated with tumors previously classified as HRD-negative. In some embodiments, the HRD-negative data is derived, at least in part, from a consensus human genome sequence or a portion thereof.

[0148] HRD positive data may include data related to genomes with HRD-associated alleles in specific HRD-associated genes. For example, in some embodiments, HRD positive data includes data related to genomes with one or more mutations of genes associated with HRD, including but not limited to BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and / or RAD45L, particularly biallelic mutations thereof. In some embodiments, HRD positive data includes one or more promoter methylation data of genes associated with HRD, including but not limited to BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and / or RAD45L. In some embodiments, the HRD positive data includes RNA expression data of one or more genes associated with HRD, including, but not limited to, BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and / or RAD45L. In some embodiments, the HRD positive data includes data associated with genomes associated with tumors found to be sensitive to platinum-based drugs and / or PARP inhibitors. In some embodiments, the HRD positive data includes data associated with genomes associated with tumors previously classified as HRD positive. In some embodiments, the HRD positive data includes data associated with tumors harboring biallelic BRCA1 and BRCA2 mutations associated with HRD.

[0149] HRD positive data can be balanced with HRD negative data. For example, in an imbalanced training dataset, the number of HRD positive training tumors may exceed the number of HRD negative tumors (or vice versa). Balancing the data ensures that the model has a sufficient number of each label to avoid biasing the labels to one label. When balanced, the number of HRD positive tumors or the number of HRD negative tumors is adjusted so that the ratio between them is at a desired level (such as about 1:1 or any other desired ratio). The balanced dataset can be used to train an HRD classifier, which can then be tested against a test dataset containing HRD positive and HRD negative tumors.

[0150] Each tumor used to train the HRD classifier includes an HRD positive or HRD negative label. Any suitable methodology can be used to computationally label a tumor as HRD positive or HRD negative (e.g., by applying metadata tags to the tumor). An HRD positive label can be assigned by the presence of a change, particularly a biallelic change, in one of the HRD-related genes, such as, but not limited to, BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and / or RAD45L. Mutations in one or both of BRCA1 and BRCA2 are particularly indicative of HRD positivity, particularly biallelic BRCA1 / BRCA2 mutations. A tumor can also be labeled as HRD positive based on clinical history. For example, if the tumor was sensitive to a PARP inhibitor or platinum-based drug regimen, the tumor is more likely to be HRD positive. HRD-negative labeling can be assigned based on the absence of a change, particularly a biallelic change, in one of the HRD-associated genes, including but not limited to BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and / or RAD45L. Mutations in HRD-associated genes can be detected by comparing gene sequences with a reference genome, such as a consensus human genome sequence, such as hg19. Similarly, tumors can also be labeled as HRD-negative based on clinical history. For example, if a tumor is resistant to a PARP inhibitor or platinum-based drug regimen, the tumor is more likely to be HRD-negative. This is especially true if the tumor was treatment-naive before treatment with a PARP inhibitor or platinum-based drug regimen, since HRD-positive tumors can develop resistance to these drugs after several treatments. Although each tumor may contain HRD-positive or HRD-negative labeling, this labeling does not require absolute certainty that the tumor is HRD-positive or HRD-negative.Instead, a robust training dataset containing a large number of HRD-positive tumors and a large number of HRD-negative tumors is given, and the contributions of false positives and false negatives are averaged out in the model by avoiding overfitting these data as known in the art. Furthermore, by using a larger training dataset, especially a balanced training dataset, as well as a dataset with clearly defined positive and negative labels (e.g., by using a verified consensus genome for HRD-negative labeling, and by using a verified biallelic BRCA1 / 2 mutant or a verified well-characterized BRCAness sample for HRD-positive labeling), the model can properly evaluate the subtle differences between HRD-negative phenotypes and phenotypes showing HRD scars (i.e., HRD-positive phenotypes).

[0151] The classification method is a computer-implemented method. The classification can be performed on a specifically configured machine or system that includes program instructions for executing a trained HRD classifier model, which can be stored in a non-transitory computer-readable memory of the computer or system. The computer generally includes one or more processors that can access the memory. The one or more processors can receive data that may be stored in the memory (e.g., test data, such as one or more copy number features and / or one or more short variant features associated with the genome of the tumor in the subject, and in some embodiments, other features and measurements). The one or more processors can access the trained HRD classifier model and input the test data into the model. The one or more processors and the trained HRD classifier model can then classify the cancer as likely to be HRD positive or likely to be HRD negative.

[0152] The HRD classifier model can classify a cancer tumor as HRD positive or HRD negative. In some embodiments, the HRD classifier model can classify a tumor as likely to be HRD positive, likely to be HRD negative, or ambiguous. For example, if the HRD classifier model cannot classify a tumor as likely to be HRD positive or likely to be HRD negative with a sufficiently high confidence or probability, the tumor can be classified as ambiguous. The confidence or probability thresholds may be set by the user as needed, taking into account the tolerance for inaccurate classification. In one example, the user may set the threshold for the HRD positive likelihood score to 0.8 and the threshold for the HRD negative likelihood score to 0.2. If the HRD positive likelihood score is less than 0.8 and / or the HRD negative likelihood score is greater than 0.2, the HRD model cannot classify the tumor as HRD positive and classifies the tumor as HRD negative (depending on how low the HRD positive likelihood score is and how high the HRD negative likelihood score is) or ambiguous.

[0153] In some embodiments, the HRD classifier outputs a likelihood score that the tumor is HRD positive. In some embodiments, the HRD classifier outputs a likelihood score that the tumor is HRD negative. The HRD classifier may be configured to output either or both of an HRD positive likelihood score and an HRD negative likelihood score. The HRD classifier may also be configured to output a ratio of the HRD positive likelihood score to the HRD negative likelihood score and / or a ratio of the HRD negative likelihood score to the HRD positive likelihood score. The likelihood score may be expressed as a value between 0.0 (indicating certainty that the tumor is neither HRD positive nor HRD negative) and 1.0 (indicating certainty that the tumor is HRD positive or HRD negative). For example, a trained HRD classifier may receive test sample data including a plurality of data features associated with a tumor of a cancer of interest and output an HRD positive likelihood score of 0.8 and an HRD negative likelihood score of 0.15. The HRD classifier may be configured to consider a tumor as HRD positive or HRD negative based on one or more likelihood scores. In the above example, the HRD classifier may consider the tumor to be HRD positive based on an HRD positive likelihood score of 0.8 and an HRD negative likelihood score of 0.15. In some embodiments, the HRD classifier considers the tumor to be HRD positive if the HRD positive likelihood score is at least 0.4, e.g., at least 0.45, at least 0.5, at least 0.55, at least 0.6, at least 0.65, at least 0.70, at least 0.75, at least 0.80, at least 0.85, at least 0.90, at least 0.95, or at least 0.99. In some embodiments, the HRD classifier considers the tumor to be HRD positive if the HRD positive likelihood score is at least 0.7. In some embodiments, the HRD classifier considers the tumor to be HRD positive if the HRD positive likelihood score is at least 0.8. In some embodiments, the HRD classifier considers the tumor to be HRD positive if the HRD positive likelihood score is at least 0.9.In some embodiments, the HRD classifier considers a tumor to be HRD negative if the HRD-negative likelihood score is at least 0.4, e.g., at least 0.5, at least 0.6, at least 0.65, at least 0.70, at least 0.75, at least 0.80, at least 0.85, at least 0.90, at least 0.95, or at least 0.99. In some embodiments, the HRD classifier considers a tumor to be HRD negative if the HRD-negative likelihood score is at least 0.7. In some embodiments, the HRD classifier considers a tumor to be HRD negative if the HRD-negative likelihood score is at least 0.8. In some embodiments, the HRD classifier considers a tumor to be HRD negative if the HRD-negative likelihood score is at least 0.9. In some embodiments, the HRD classifier considers a tumor to be HRD positive if the HRD negative likelihood score is less than 0.5, e.g., less than 0.45, less than 0.40, less than 0.35, less than 0.30, less than 0.30, less than 0.25, less than 0.20, less than 0.15, less than 0.10, or less than 0.05. In some embodiments, the HRD classifier considers a tumor to be HRD negative if the HRD positive likelihood score is less than 0.5, e.g., less than 0.45, less than 0.40, less than 0.35, less than 0.30, less than 0.30, less than 0.25, less than 0.20, less than 0.15, less than 0.10, or less than 0.05. In some embodiments, the HRD classifier considers a tumor to be HRD positive if the HRD positive likelihood score is above a certain threshold (such as at least 0.80) and the HRD negative likelihood score is below a certain threshold (such as less than 0.25). In some embodiments, the HRD classifier considers a tumor to be HRD negative if the HRD negative likelihood score is above a certain threshold (such as at least 0.80) and the HRD positive likelihood score is below a certain threshold (such as less than 0.25). In some embodiments, the HRD classifier considers a tumor to be ambiguous if the HRD positive likelihood score is below a certain threshold and the HRD negative likelihood score is below a threshold, or if the absolute values ​​of the likelihood scores are within a threshold similarity percentage.

[0154] A report can be generated that identifies the cancer as likely HRD positive or likely HRD negative (or equivocal). The report can be, for example, an electronic medical record or a printed report that can be sent to the subject or a health care provider associated with the subject (e.g., doctor, nurse, clinic, etc.). The report can be used to make medical decisions, such as methods or drugs to treat the cancer tumor.

[0155] The report may be displayed on an electronic display or a customized interface. For example, in some embodiments, the computer-implemented method can automatically generate the report and automatically display the generated report on an electronic display or a customized interface.

[0156] FIG. 7 illustrates an exemplary method for training and operating an HRD classification model 702 configured to classify a subject's cancer tumor as HRD positive or HRD negative. The HRD classification model 702 is trained using a dataset including an HRD positive training dataset 704 and an HRD negative training dataset 706. The HRD positive training dataset 704 includes one or more HRD positive sample data elements (i.e., data of HRD positive sample 1 through HRD positive sample i). Each HRD positive sample data element is associated with a feature of the HRD positive tumor (e.g., copy number feature, basic feature, short variant feature, etc.). The HRD positive sample data elements may also include other data features, such as a measure of gLOH and / or a short variant feature (not shown). The feature is labeled as associated with an HRD positive label. Similarly, the HRD negative training dataset 706 includes one or more HRD negative training sample data elements (i.e., data of HRD(-) sample 1 through HRD(-) sample j). Each HRD negative sample data element is associated with an HRD negative tumor feature (e.g., a copy number feature, a basal feature, a short variant feature, etc.). The HRD negative sample data elements may also include other data features, such as a measure of gLOH and / or a short variant feature (not shown). The HRD negative samples are labeled as being associated with the HRD negative label.

[0157] In some embodiments, the HRD classification model 702 is a tree-based gradient boosting model (such as XGBoost). In this model, rather than training all models in isolation from each other (e.g., by random forests), the models are trained successively such that each new model fits the residuals from the previous model. Thus, the model achieves a strong classifier from many weaker classifiers connected in sequence. Iterative cross-validation can be used on the training data to estimate the performance of the HRD classification model.

[0158] After the classification model 702 is trained on the training dataset, the classification model 702 can be used to classify the subject's cancer tumor as HRD positive or HRD negative. To classify the subject's cancer tumor as HRD positive or HRD negative, the classification model 702 receives test data 708 including test feature data associated with the tumor to be classified. The test data 708 includes one or more copy number features and may include one or more base features, one or more short variant features, etc. The classification model 702 can determine a probability 710 that the tumor is HRD positive and / or a probability 712 that the tumor is HRD negative. The probabilities 710 and 712 are optionally input to an HRD calling module 714. The HRD calling module 714 can consider the cancer as HRD positive or HRD negative. For example, if the probability 710 that the tumor test sample is HRD positive is greater than the probability 712 that the tumor test sample is HRD negative, the tumor test sample can be considered HRD positive. If the probability 712 that the tumor test sample is HRD negative is greater than the probability 710 that the tumor test sample is HRD positive, the tumor test sample may be considered HRD negative. Optionally, if neither of probabilities 710 and 712 is above a predetermined threshold, the tumor test sample may be considered equivocal.

[0159] The methods described herein can be implemented using one or more computer systems. Such a computer system can include one or more programs configured for the computer system to execute on one or more processors to perform such methods. One or more steps of the computer-implemented methods may be performed automatically. The computer system can include one or more computing nodes. For example, the system can include two or more computing nodes (e.g., servers, computers, routers, or other types of electronic devices including network interfaces) that can be connected and configured to communicate and perform the methods over a network on one or more computing nodes of the network.

[0160] FIG. 8 illustrates an example of a computing device according to an embodiment. The device 1100 may be a host computer connected to a network. The device 1100 may be a client computer or a server. As illustrated in FIG. 8, the device 1100 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (a portable electronic device, e.g., a phone or tablet). The device may include, for example, one or more of a processor 1110, an input device 1120, an output device 1130, a storage 1140, and a communication device 1160. The input device 1120 and the output device 1130 may generally correspond to those described above and may be connectable or integrated with the computer.

[0161] The input device 1120 may be any suitable device that provides input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. The output device 1130 may be any suitable device that provides output, such as a display, a touch screen, a tactile device, or a speaker.

[0162] Storage 1140 can be any suitable device with storage, such as electrical, magnetic, or optical memory, including RAM, cache, a hard drive, or a removable storage disk. Communications device 1160 can include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of a computer can be connected in any suitable manner, such as by a physical bus or wirelessly.

[0163] The HRD classification module 1150 may be stored in the storage 1140 and executed by the processor 1110 and may include one or more program instructions for executing and implementing methods and processes related to the HRD model (e.g., as implemented in a device such as those described above).

[0164] The HRD module 1150 may also be stored in and / or transferred to any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (e.g., those described above) and may fetch instructions associated with the software from and execute the instructions. In the context of the present disclosure, a computer-readable storage medium may be any medium, such as storage 1140, that may contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.

[0165] The HRD module 1150 may also propagate in any transmission medium for use by or in connection with an instruction execution system, apparatus, or device (such as those mentioned above) and may fetch instructions associated with the software from and execute the instructions. In the context of this disclosure, a transmission medium may be any medium that may communicate, propagate, or transmit transmission programming for use by or in connection with an instruction execution system, apparatus, or device. Transmission-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.

[0166] The device 1100 can be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communication protocol and can be protected by any suitable security protocol. The network can include any suitable configuration of network links capable of implementing transmission and reception of network signals, such as wireless network connections (T1 or T3 lines), cable networks, DSL, or telephone lines.

[0167] The device 1100 may implement any operating system suitable for operating on a network. The software 350 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying the functionality of the present disclosure may be deployed in different configurations, such as in a client / server configuration or via a web browser as a web-based application or web service.

[0168] Treatment method Characterizing a tumor as HRD positive or HRD negative (or likely to be HRD positive or likely to be HRD negative) is particularly useful for selecting an effective treatment for a subject with a tumor. Tumors classified as HRD positive are often more sensitive to certain drugs and therapies to which HRD negative tumors may be resistant. Different drugs or therapies can be selected based on the classification of a tumor as HRD positive, likely to be HRD positive, HRD negative, or likely to be HRD negative. Thus, a method of treating a subject's cancer can include assessing a cancer tumor as likely to be HRD positive or HRD negative (or considering a cancer tumor as HRD positive or HRD negative) according to the methods described herein, and then administering a therapeutically effective amount of a drug to the subject based on classifying the tumor as likely to be HRD positive or likely to be HRD negative (or considering the tumor as HRD positive or HRD negative).

[0169] The method of treating cancer in a subject can include obtaining a classification of a tumor of the cancer of the subject as likely to be HRD positive or likely to be HRD negative. To obtain this classification, the HRD classification model described herein can be used. One or more copy number features associated with the genome of the cancer tumor can be input into an HRD classification model configured to classify the tumor as likely to be HRD positive or likely to be HRD negative based on one or more copy number features associated with the genome of the tumor of the subject. The HRD classification model is trained using HRD positive data from a plurality of HRD positive tumors and HRD negative data from a plurality of HRD negative tumors. The classification can be obtained, for example, by operating the HRD classification model or by receiving results from another that has operated the HRD classification model.

[0170] The one or more basic features and / or the one or more short variant features can be input into an HRD classification model configured to classify a tumor as likely to be HRD positive or likely to be HRD negative based on the one or more basic features and / or the one or more short variant features. The one or more short variant features and the one or more basic features can be in addition to or instead of the one or more copy number features.

[0171] In some embodiments, the treatment method may include obtaining test sample data including one or more copy number signatures. In some embodiments, the treatment method may include obtaining one or more basic signatures. In some aspects, the treatment method may include obtaining a measure of genome-wide loss of heterozygosity. In some embodiments, the treatment method may include obtaining one or more short variant signatures. A test sample may be obtained from a subject, and the nucleic acid molecule may be derived from the test sample. The test sample may be, for example, a solid tissue biopsy of a cancer, and nucleic acid may be isolated from the solid tissue sample. Optionally, the test sample may be preserved, for example, by freezing the test sample or fixing the sample (e.g., by forming a formalin-fixed paraffin-embedded (FFPE) sample) before isolating the nucleic acid molecule. Alternatively, the test sample may be a liquid biopsy sample (e.g., blood, plasma, or other liquid sample from a subject), and nucleic acid including circulating tumor DNA (ctDNA) may be obtained from the liquid sample. The nucleic acid from the sample may be assayed and then analyzed to generate either one or more copy number signatures, one or more basic signatures, or one or more short variant signatures.

[0172] Obtaining a classification of the tumor as likely to be HRD positive or likely to be HRD negative can include inputting the described features and / or measures into an HRD classification model and using the features and / or measures to classify the cancer as likely to be HRD positive or likely to be HRD negative based on the data inputted into the HRD classification model. Alternatively, obtaining a classification of the tumor as likely to be HRD positive or likely to be HRD negative can include receiving a report from another entity. The report may be generated by the other entity, and the report can include a classification of the tumor as likely to be HRD positive or likely to be HRD negative, the classification being generated using the HRD classification model described herein. In some embodiments, the report includes a likelihood score that the tumor is HRD positive and / or a likelihood score that the tumor is HRD negative, and a final classification can be made based on the likelihood score.

[0173] Once the tumor is classified as likely to be HRD positive or likely to be HRD negative, a treatment can be selected based on the classification. If the tumor is classified as likely to be HRD positive, a treatment that is effective for HRD positive tumors can be selected. The selected treatment can then be administered to the subject to treat the tumor classified as likely to be HRD positive. If the tumor is classified as likely to be HRD negative, a treatment that is not a platinum-based drug or a PARP inhibitor can be selected. The selected treatment can then be administered to the subject to treat the tumor classified as likely to be HRD negative.

[0174] Effective treatments for HRD-positive tumors can include one or more PARP inhibitors and / or one or more platinum-based drugs. PARP inhibitors can include, but are not limited to, veliparib, olaparib, talazoparib, iniparib, rucaparib, and niraparib. PARP inhibitors are described in Murphy and Muggia, PARP inhibitors: clinical development, emerging differences, and the current therapeutic issues, Cancer Drug Resist 2019;2:665-79. Platinum-based drugs can include, but are not limited to, cisplatin, oxaliplatin, and carboplatin. Platinum-based drugs are described in Rottenberg et al., The rediscovery of platinum-based cancer therapy, Nat.Rev.Cancer 2021 Jan;21(1):37-50.

[0175] The tumor to be treated is a tumor of the subject. In one embodiment, the tumor is pancreatic cancer. In another embodiment, the tumor is prostate cancer. In some embodiments, the tumor is ovarian cancer, breast cancer or prostate cancer. In some embodiments, the tumor is a tumor associated with HRD, and may include, but is not limited to, one of adrenal, bile duct, bone / soft tissue, breast, colon / rectum, esophagus, eye, head and neck, kidney, liver, lung, lymphatic system, medulloblastoma, mesothelioma, myeloid system, nervous system, neuroendocrine, ovarian, pancreatic, prostate, skin, stomach, testis, thymus, thyroid, urinary tract, uterus, or vulvar cancer. See Nguyen et al., Pan-cancer landscape of homologous recombination deficiency, Nat.Commun.2020 Nov 4;11(1):5584.

[0176] Although the present disclosure has been fully described with reference to the accompanying drawings, it should be noted that various modifications and changes will become apparent to those skilled in the art. Such modifications and changes should be understood to be included within the scope of the present disclosure as defined by the claims.

[0177] The above description has been described with reference to specific embodiments for illustrative purposes. However, the above illustrative description is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described to best explain the principles of the technology and their practical applications. This will enable others skilled in the art to best utilize the technology and various embodiments with various modifications suited to the particular use contemplated.

Claims

1. A step of preparing a genome obtained from a target tumor, Optionally, a step of ligating one or more adapters onto the genome, A step of amplifying nucleic acid molecules from the genome, A step of capturing nucleic acid molecules from the amplified genome, wherein the captured nucleic acid molecules are captured by hybridization to one or more bait molecules, A step of deriving a set of input features from the captured nucleic acid molecules, A step of inputting the set of input features into a trained homologous recombination deficiency (HRD) model by one or more processors and using the trained HRD model to identify the tumor as HRD positive or HRD negative, wherein the model Determining a metric of the importance of one or more features associated with each feature of a plurality of features, Identifying a subset of features among the plurality of features using the metric of the importance of the one or more features, and Training the HRD model by the one or more processors based on the identified subset of features Inputting the set of input features trained into a trained homologous recombination deficiency (HRD) model trained thereby, using the trained HRD model to identify the tumor as HRD positive or HRD negative, and Using the trained HRD model to classify the tumor as HRD positive or HRD negative by the one or more processors A method comprising.

2. A step of receiving a plurality of features by one or more processors, A step of identifying a subset of features among the plurality of features using a metric of the importance of one or more features by the one or more processors, and A step of training a homologous recombination deficiency (HRD) model based on the identified subset of the plurality of features by the one or more processors, wherein the HRD model is configured to receive sample data associated with the genome of a tumor of interest and use the sample data to identify the tumor of interest as HRD positive or HRD negative, the step of training the HRD model A method comprising. **Claim 3** A step of receiving, by one or more processors, sample data related to the genome of a tumor in a subject, A step of inputting, by the one or more processors, the sample data into a trained homologous recombination deficiency (HRD) model, wherein the HRD model Determining a metric of the importance of one or more features associated with each feature of the plurality of features, Identifying a subset of features of the plurality of features using the metric of the importance of the one or more features, and Training the HRD model by the one or more processors based on the identified subset of features The step of inputting the sample data into the trained homologous recombination deficiency (HRD) model trained thereby, and A step of classifying the tumor as HRD positive or HRD negative by the one or more processors using the trained HRD model A method comprising. **Claim 4** The method according to any one of claims 1 to 3, wherein the plurality of features includes one or more copy number features, one or more short variant features, or a combination thereof. **Claim 5** The metric of the importance of the one or more features includes one or more of chi-squared test, analysis of variance (ANOVA), random forest, or gradient boosting, the method according to any one of claims 1 to 3.

6. The step of identifying the subset of the features among the plurality of features comprises obtaining, by the one or more processors, one or more feature rankings according to the metric of the importance of the one or more features, and selecting, by the one or more processors, the subset of the plurality of features based on the one or more feature rankings The method according to any one of claims 1 to 3, comprising.

7. The step of identifying the subset of the plurality of features comprises (a) obtaining, by one or more processors, a feature ranking of the plurality of features according to the metric of the importance of the features, (b) obtaining, by the one or more processors, a new feature set by adding one or more features from the plurality of features to an existing feature set based on the feature ranking, (c) training, by the one or more processors, a new HRD model using the new feature set, (d) evaluating, by the one or more processors, the trained new HRD model to obtain an evaluation result, (e) storing, by the one or more processors, the evaluation result related to the new HRD model and the new feature set, (f) repeating steps (b) to (e) by the one or more processors to obtain a plurality of evaluation results until a condition is met, and (g) selecting, by the one or more processors, the subset of the plurality of features based on the plurality of evaluation results The method according to any one of claims 1 to 3, comprising:

8. The trained HRD model is a classification model, and the method comprises: receiving new sample data associated with the genome of a tumor in a new subject, the new sample data being associated with the subset of the plurality of features, providing the new sample data to the trained HRD classification model to generate a classification result of HRD positive or HRD negative, and outputting the classification result The method according to any one of claims 1 to 3, further comprising:

9. The method according to claim 8, wherein the classification result includes at least one of an HRD positive likelihood score and an HRD negative likelihood score.

10. The method according to any one of claims 1 to 3, wherein the HRD model is a classification model, a regression model, a neural network, or any combination thereof.

11. The method according to claim 9, comprising recording at least one of the HRD positive likelihood score and the HRD negative likelihood score in a digital electronic file associated with the new subject.

12. The method according to claim 9, comprising recording in a digital electronic file associated with the new subject a designation that the tumor is HRD positive based on the HRD positive likelihood score or that the tumor is HRD negative based on the HRD negative likelihood score.

13. The method according to any one of claims 1 to 3, wherein the plurality of features includes at least one of a feature of segment minor allele frequency (segMAF), a feature of the number of sequencing reads, a feature of segment size, a feature of the number of breakpoints per x megabase, a feature of the copy number of change points, a feature of the segment copy number, a feature of the number of breakpoints per chromosome arm, and a feature of the number of segments with oscillating copy number.

14. The method according to any one of claims 1 to 3, wherein at least one of the plurality of features is evaluated across the centromeric portion of the genome.

15. The method according to any one of claims 1 to 3, wherein at least one of the plurality of features is evaluated across the telomeric portion of the genome.

16. The method according to any one of claims 1 to 3, wherein at least one of the plurality of features is evaluated across both the centromeric portion and the telomeric portion of the genome.

17. The method according to any one of claims 1 to 3, wherein the plurality of features includes a feature of the number of breakpoints per x megabase, and the feature of the number of breakpoints per x megabase is based on the number of breakpoints that appear in a window of length x megabase across the entire genome.

18. The method according to claim 17, wherein the feature of the number of breakpoints per x megabase is evaluated across (i) the telomeric portion of the genome, (ii) the centromeric portion of the genome, or (iii) both the telomeric portion and the centromeric portion of the genome.

19. The method according to claim 17, wherein x is from about 1 to about 100 megabases.

20. The method according to claim 17, wherein x is about 10 megabases, about 25 megabases, about 50 megabases, or about 100 megabases.

21. The method according to claim 17, wherein the characteristic of the number of breakpoints per x megabase is a binned characteristic. **Claim 22** The method according to any one of claims 1 to 3, wherein the plurality of characteristics includes a characteristic of the number of change point copies, and the number of change point copies is based on an absolute difference in copy number between adjacent genomic segments across the genome of the tumor of interest. **Claim 23** The method according to claim 22, wherein the characteristic of the number of change point copies is derived from copy number data normalized by a multiple. **Claim 24** The method according to claim 22, wherein the characteristic of the number of change point copies is evaluated across (i) the telomeric portion of the genome, (ii) the centromeric portion of the genome, or (iii) both the telomeric portion and the centromeric portion of the genome. **Claim 25** The method according to claim 22, wherein the characteristic of the number of change point copies is a binned characteristic. **Claim 26** The method according to any one of claims 1 to 3, wherein the plurality of characteristics includes a characteristic of segment copy number, and the segment copy number is based on the copy number of each genomic segment. **Claim 27** The method according to claim 26, wherein the characteristic of the segment copy number is evaluated across (i) the telomeric portion of the genome, (ii) the centromeric portion of the genome, or (iii) both the telomeric portion and the centromeric portion of the genome. **Claim 28** The method according to claim 26, wherein the characteristic of the segment copy number is derived from copy number data normalized by a multiple. **Claim 29** The method according to claim 26, wherein the characteristic of the segment copy number is a binned characteristic. **Claim 30** The method according to any one of claims 1 to 3, wherein the plurality of characteristics includes a characteristic of the number of breakpoints per chromosomal arm of the genome of the tumor of interest.

31. The method according to claim 30, wherein the characteristic of the number of breakpoints per chromosomal arm is evaluated across (i) the telomeric part of the genome, (ii) the centromeric part of the genome, or (iii) both the telomeric part and the centromeric part of the genome.

32. The method according to claim 30, wherein the characteristic of the number of breakpoints per chromosomal arm is a binned characteristic.

33. The method according to any one of claims 1 to 3, wherein the plurality of characteristics includes a characteristic of the number of segments of the oscillating copy number.

34. The method according to claim 33, wherein the characteristic of the number of segments of the oscillating copy number is based on the number of repeated alternating segments between two copy numbers across the genome of the tumor of the subject.

35. The method according to claim 33, wherein the characteristic of the number of segments of the oscillating copy number is evaluated across (i) the telomeric part of the genome, (ii) the centromeric part of the genome, or (iii) both the telomeric part and the centromeric part of the genome.

36. The method according to claim 33, wherein the characteristic of the number of segments of the oscillating copy number is a binned characteristic.

37. The method according to any one of claims 1 to 3, wherein the one or more copy number characteristics include a segment minor allele frequency (segMAF) characteristic, and segMAF is based on the minor allele frequency in a heterozygous single nucleotide polymorphism.

38. The method according to claim 37, wherein segMAF is evaluated across (i) the telomeric part of the genome, (ii) the centromeric part of the genome, or (iii) both the telomeric part and the centromeric part of the genome.

39. The method according to claim 37, wherein the characteristics of the segment minor allele frequency are binned characteristics.

40. The method according to any one of claims 1 to 3, wherein the characteristics of the one or more copy numbers include the characteristics of the number of sequencing reads.

41. The method according to claim 40, wherein the characteristics of the number of sequencing reads are binned characteristics.

42. The method according to any one of claims 1 to 3, wherein the plurality of characteristics further includes a measure of the loss of overall genomic heterozygosity of the genome of the tumor of the subject.

43. The method according to any one of claims 1 to 3, wherein the plurality of characteristics includes the characteristics of one or more short variants.

44. The method according to claim 43, wherein the characteristics of the one or more short variants include at least one of a deletion of microhomology or repeat region characteristics and a mutation signature derived from the characteristics of two or more short variants.

45. The method according to claim 44, wherein the deletion of microhomology or repeat region characteristics is a deletion of at least 5 base pairs.

46. The step of training the HRD model is receiving, by the one or more processors, an HRD positive training dataset, the HRD positive training dataset including a plurality of characteristics related to an HRD positive tumor and an HRD positive label, the step of receiving the HRD positive training dataset, receiving, by the one or more processors, an HRD negative training dataset, the HRD negative training dataset including a plurality of characteristics related to an HRD negative tumor and an HRD negative label, the step of receiving the HRD negative training dataset, The step of training the HRD model by using the HRD positive training data set and the HRD negative training data set by the one or more processors The method according to any one of claims 1 to 3, comprising:

47. The method according to any one of claims 1 to 3, further comprising the step of testing the trained model by using an HRD positive test data set including an HRD positive control derived from a genomic sequence including loss-of-function mutations in BRCA1, BRCA2, both BRCA1 and BRCA2, or biallelic mutations of BRCA1 and BRCA2, by the one or more processors.

48. The method according to any one of claims 1 to 3, further comprising the step of testing the trained model by using an HRD positive test data set including an HRD positive control derived from a genomic sequence including a loss-of-function mutation in at least one of ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, or RAD45L, by the one or more processors.

49. The method according to any one of claims 1 to 3, further comprising the step of testing the trained model by using an HRD negative test data set including an HRD negative training data set including an HRD negative control derived from a consensus human genomic sequence, by the one or more processors.

50. The method according to claim 46, wherein the training step includes using an HRD positive training data set and an HRD negative training data set.

51. The method according to claim 50, further comprising the step of balancing the HRD positive training data set and the HRD negative training data set by the one or more processors before training the HRD model.

52. The method according to any one of claims 1 to 3, wherein the tumor in the subject is prostate cancer, ovarian cancer, breast cancer, non-small cell lung cancer (NSCLC), colorectal cancer (CRC), or pancreatic cancer.

53. The method according to any one of claims 1 to 3, wherein the step of training the HRD model comprises fitting the HRD model to sample data related to ovarian cancer, non-small cell lung cancer (NSCLC), colorectal cancer (CRC), breast cancer, pancreatic cancer, or prostate cancer, and the sample data comprises the subset of the plurality of features.

54. The method according to any one of claims 1 to 3, wherein the tumor is obtained from a sample that is a solid tissue biopsy sample.

55. The method according to claim 54, wherein the solid tissue biopsy sample is a formalin-fixed paraffin-embedded (FFPE) sample.

56. The method according to any one of claims 1 to 3, wherein the tumor is obtained from a sample that is a liquid biopsy sample containing circulating tumor DNA (ctDNA).

57. The method according to any one of claims 1 to 3, wherein the tumor is obtained from a sample that is a liquid biopsy sample containing cell-free DNA (cfDNA).

58. The method according to any one of claims 1 to 3, further comprising the step of determining, identifying, or applying the output of the tumor as HRD positive or HRD negative as a diagnostic value related to the patient.

59. The method according to any one of claims 1 to 3, further comprising the step of generating a genomic profile of the subject based on the output of the tumor as HRD positive or HRD negative.

60. The method according to claim 59, further comprising the step of administering an anti-cancer agent or applying an anti-cancer treatment to the subject based on the generated genomic profile.

61. The method according to any one of claims 1 to 3, wherein the output as HRD positive or HRD negative of the tumor is used to generate the genomic profile of the subject.

62. The method according to any one of claims 1 to 3, wherein the output as HRD positive or HRD negative of the tumor is used when making a decision on a proposed treatment for the subject.

63. The method according to any one of claims 1 to 3, wherein the output as HRD positive or HRD negative of the tumor is used to apply or administer treatment to the subject.

64. The method according to any one of claims 1 to 3, wherein the HRD model is a machine learning model.

65. The method according to any one of claims 1 to 3, wherein the subject has cancer, is at risk of having cancer, or is suspected of having cancer.

66. A method for treating a subject's cancer, comprising: (a) identifying the tumor as HRD positive or HRD negative according to the method according to any one of claims 1 to 3; (b) when the tumor of the cancer is evaluated as HRD positive, administering to the subject a therapeutically effective amount of a drug effective for HRD positive tumors.

67. The method according to claim 66, wherein the drug effective for HRD positive tumors is a platinum-based drug or a PARP inhibitor.

68. The method according to claim 66, comprising, when the tumor is evaluated as HRD negative, administering to the subject a therapeutically effective amount of a drug that is neither a platinum-based drug nor a PARP inhibitor.

69. A method for selecting a treatment method for a subject's cancer, comprising: (a) evaluating the tumor of the cancer as HRD-positive or HRD-negative according to the method according to any one of claims 1 to 3; (b) when the cancer is evaluated as HRD-positive, selecting a treatment effective in HRD-positive tumors A method comprising:

70. The method according to claim 69, comprising selecting a treatment method that is neither a platinum-based drug nor a PARP inhibitor when the tumor is evaluated as HRD-negative.

71. The method according to claim 70, wherein the treatment method effective for HRD-positive tumors is a platinum-based drug or a PARP inhibitor.

72. A computer system, comprising: one or more processors; a memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 1 to 3; one or more programs A computer system comprising:

73. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 3. Non-transitory computer-readable storage medium.