Method for constructing a genomic scar model
By constructing a gene biomarker model and using machine learning to train weights, the limitations of existing BRCA gene detection technologies have been overcome, enabling accurate prediction of BRCAness events and patient identification, simplifying the interpretation of test results, and improving treatment sensitivity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- AMOY DIAGNOSTICS CO LTD
- Filing Date
- 2020-12-29
- Publication Date
- 2026-05-07
AI Technical Summary
Current technologies for detecting BRCA gene mutations cannot fully cover all gene abnormalities that cause homologous recombination repair defects, resulting in some patients not being included in effective treatments. Furthermore, the interpretation of test results is complex and fails to meet clinical needs.
A gene biomarker model is constructed, and weights are trained using machine learning methods. Based on CNV type and number, a gene biomarker score (GSS) model is built to predict BRCAness events, avoiding redundant calculations and high-cost whole-genome sequencing.
Accurate prediction of BRCAness status improves the identification rate of patients with HRR-related gene mutations, enhances the identification of patients sensitive to platinum-based drugs and PARPi, and simplifies the interpretation of test results.
Smart Images

Figure 0007854805000002 
Figure 0007854805000003 
Figure 0007854805000004
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of gene diagnosis, and specifically relates to a method for constructing a genomic scar model.
Background Art
[0002] Homologous Recombination Repair (HRR) is an important repair method for double-stranded DNA damage and is frequently seen in cells to accurately repair harmful breaks in double-stranded DNA. HRR is a complex signaling pathway involving multiple steps, among which breast cancer susceptibility genes (BRCA1 / 2) are important genes related to homologous recombination function. When mutations occur in the BRCA gene and the BRCA1 and BRCA2 proteins lose their functions, it leads to HRR dysfunction. Generally, it is called Homologous Recombination Deficiency (HRD), and HRD widely exists in breast cancer, ovarian cancer, prostate cancer and pancreatic cancer as a tumor driver event. Usually, tumors with mutations or abnormal expression of BRCA1 / 2 are sensitive to platinum preparations and poly[ADP-ribose] polymerase inhibitors (PARPi). Therefore, the detection of mutations in the BRCA1 / 2 gene plays a prominent role in the clinical classification and dosing guidance of this type of disease.
[0003] However, as research progresses, the detection of BRCA gene mutations is increasingly unable to meet existing clinical needs. The concentration of patients with BRCA gene mutations in the treatment response group is decreasing, and some patients may be missed from treatment. For example, in triple-negative breast cancer, 20% of patients have BRCA gene mutations, but the overall response rate to platinum-based chemotherapy in this group is approximately 30%. Furthermore, in high-grade serous ovarian cancer, 30% of patients have BRCA gene mutations, but the overall response rate to PARPi in this group is approximately 50%. This indicates that some BRCA-negative patients respond to platinum-based chemotherapy or PARPi, and therefore, BRCA detection may miss some patients from treatment. The main reason for this omission is, firstly, that the detection of BRCA gene mutations is relatively limited. There are many HRR function-related genes, and BRCA1 / 2 are two of these with relatively high mutation frequencies. From an analysis based on the principle of drug action, any HRR gene that can form a synthetic lethal effect with platinum-based drugs and PARPi is noteworthy (collectively referred to as BRCAness events). For example, the results of the PROfound trial showed that deletion of the HRR-related gene ATM is effective against prostate cancer with PARPi olaparib, and therefore, many clinical trials are gradually shifting their focus from BRCA to other HRR-related genes. Next, BRCA gene mutation detection cannot cover all types of genomic abnormalities that cause HRR dysfunction. In addition to gene mutations, methylation of the BRCA1 promoter region and loss of heterozygosity (LoH) within the BRCA gene region are also major causes of HRR dysfunction. Finally, the interpretation of results from BRCA gene mutation detection is complex and prone to omissions, making clinical application relatively difficult. Currently, many authoritative organizations, such as the American Society of Clinical Genetics and Genomics (ACMG) and the European Molecular Genetics Quality Network (EMQN), provide optimal guidelines for molecular genetic analysis of hereditary breast / ovarian cancer.The classification of BRCA mutations—pathogenic, possibly pathogenic, significance unknown, possibly benign, and benign—varies somewhat in terms of the level of evidence used in different guidelines, which presents a significant obstacle to clinical application.
[0004] For the reasons described above, there is a very urgent need for novel clinical molecular markers that can easily quantitatively evaluate homologous recombination repair deficiencies in cells. Searching for molecular markers characteristic of downstream genome mutations (including mutations, copy number variations (CNVs), and gene expression abnormalities) that induce HRR deficiency is a major research goal today. In 2009, Olafur et al. discovered that the characteristics of CNV mutations are closely related to BRCAness, and in 2012, Abkevich et al. discovered that the number of LoHs in the whole genome is significantly related to BRCAness events. In the same year, Popova et al. discovered that large-scale state transitions (LSTs) in genes are related to BRCA1 / 2 gene inactivation, and Birkbak et al. discovered that telomeric allelic imbalance (TAI) is related to BRCAness events in triple-negative breast cancer and is significantly enriched in the platinum-sensitive group. In 2016, Myriad, Inc. in the United States quantitatively calculated an HRD score (HRD assessment) by aggregating the number of LoH, LST, and TAI occurrences across the entire genome. This aggregated index can accurately predict BRCAness events and simultaneously effectively enrich the platinum- and PARPi-sensitive patient population. Compared to detecting the BRCA gene alone, the HRD assessment can screen 40% more potential patients. In addition, in 2017, Davies et al. discovered that whole-genome single-nucleotide mutation models, genome rearrangement models, and indel models are all closely related to HRR deficiencies, and that combining several of these models with the HRD assessment allows for accurate prediction of BRCAness events using logistic regression. However, each of these methods has its limitations. For example, the HRD assessment simply adds LoH, LST, and TAI without much thought, but in reality, TAI and LoH overlap in some situations, resulting in duplicate values. Furthermore, some CNV types are not considered within the HRD assessment.While Davies' model takes a comprehensive approach, it requires whole-genome sequencing to aggregate models for each type of mutation, which makes the testing extremely costly. [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] The objective of this invention is to overcome the shortcomings of existing technologies and to provide a method for constructing a genomic scar model.
[0006] Another objective of the present invention is to provide applications for the genomic scar model constructed by the above construction method. [Means for solving the problem]
[0007] The technical proposal of this invention is as follows: The method for constructing a genome scar model includes the following steps. (1) Collect known positive and negative samples to form a training set; (2) Analyze the CNV status of the above training set and determine the type of CNV and the corresponding quantity; (3) Confirm BRCAness positive events and BRCAness negative events; (4) Using machine learning methods, weights for different types of CNVs determined in step (2) are trained based on BRCAness-positive and BRCAness-negative events in the training set, and then the weighted CNVs of different types are accumulated to obtain a genomic scar model used to calculate the GSS (Genomic Scar Score); (5) Collect other known positive and negative samples to construct a test set, and obtain the CNV type and corresponding quantity of the test set based on step (2); (6) The results obtained in step (5) are substituted into the genome scar model obtained in step (4), the GSS of the test set is calculated, and the genome scar model is further validated based on the GSS score.
[0008] In one preferred embodiment of the present invention, step (2) is performed to sequence the training set obtained in step (1) and calculate the status of CNVs in the sequence analysis results. Furthermore, adjacent regions with the same copy number variation are connected to form fragments to prevent calculation duplication and to determine the type and corresponding quantity of CNVs. More preferably, the sequencing analysis described above is based on whole genome, whole exome, targeted capture sequencing, or copy number variant chips.
[0009] In one preferred embodiment of the present invention, the BRCAness positive event includes: a pathogenic or possibly pathogenic mutation occurring in one allele of any one of BRCA1 / 2 and a loss of heterozygosity occurring in another allele; two pathogenic or possibly pathogenic mutations occurring in any one of BRCA1 / 2; or a loss of heterozygosity occurring in one allele of BCRA1 and methylation occurring in the promoter region of another allele.
[0010] The aforementioned BRCAness-positive events further include related events in which other homologous recombination repair-related genes other than the BRCA1 / 2 genes cause genomic instability through mutation, loss of homozygosity, or silencing of the corresponding genes.
[0011] In one preferred embodiment of the present invention, the BRCAness-negative event is characterized by the HRR-related gene being wild-type, and furthermore, the loss of heterozygosity in the corresponding allele or the absence of methylation in its promoter region.
[0012] In one preferred embodiment of the present invention, the type of CNV in step (2) is determined based on the length of the mutant fragment, the type of the mutant fragment, and the genomic location of the mutant fragment.
[0013] More preferably, the length of the mutant fragments can be divided into short fragments of 5 to 10 M, medium fragments longer than 10 M and 15 M or less, and long fragments longer than 15 M.
[0014] More preferably, the length of the mutant fragment can be divided into continuous variables such as CNV fragments of length 5, 6, 7, 8, 9…30M.
[0015] More preferably, the type of mutant fragment includes loss of heterozygosity, unbalanced amplification of the mutant fragment, and balanced amplification of the mutant fragment.
[0016] More preferably, the genomic location of the mutant fragment includes the mutant fragment being located on the telomere side, the mutant fragment being located within the centromere region, and the mutant fragment being located outside the telomere side and within the centromere region.
[0017] Another technical option 1 of the present invention: Application of the genomic scar model constructed using the above method to enrich the HRR mutation-related group.
[0018] Another technical option 2 of the present invention: Application of the genomic scar model constructed using the above method to enrich the platinum-sensitive group.
[0019] Another technical proposal 3 of the present invention: Application of the genomic scar model constructed using the above method to enrich the PARPi-sensitive group. [Effects of the Invention]
[0020] The beneficial effects of this invention are as follows: 1. Compared to detecting mutations in the BRCA gene, the present invention does not require the detection of methylation of the BRCA1 promoter region and loss of heterozygosity in the BRCA gene, and can accurately predict the BRCAness status of the sample being measured. In addition, for the complex decoding of BRCA mutations, the present invention can directly present the interpretation results based on the genome scar score.
[0021] 2. Compared to detecting mutations in the BRCA gene, the present invention can concentrate patients with HRR-related gene mutations.
[0022] 3. Compared with the detection of BRCA gene mutations, the present invention can enrich more platinum agent-sensitive patients.
[0023] 4. Compared with the detection of BRCA gene mutations, the present invention can enrich more PARPi-sensitive patients.
Brief Description of Drawings
[0024] [Figure 1] Figure 1 is a diagram of experimental results in Example 3 of the present invention. [Figure 2] Figure 2 is Diagram 1 of experimental results in Example 4 of the present invention. [Figure 3] Figure 3 is Diagram 2 of experimental results in Example 4 of the present invention.
Modes for Carrying Out the Invention
[0025] Hereinafter, specific embodiments will be combined with the attached drawings to further explain and describe the technical solution of the present invention.
Examples
[0026] The types of 110 and 18 test samples collected respectively are FFPE samples of ovarian cancer patients and control blood samples, which are used as training sets and test sets for constructing a genomic scar model. Subsequently, database construction and capture were performed using the homologous recombination repair deficiency (HRD) test kit of Xiamen艾德生物医药科技股份有限公司. This kit includes 35 HRR-related genes and 70,000 snp sites as capture regions. The captured and concentrated DNA was finally sequenced using an Illumina Novaseq sequencer. <00001Raw data is matched against a human reference genome sequence (hg19) using BWA (Li H. and Durbin R. 2009). A BAM file is generated from the matched data and used as the input file for mutations and copy number variations. For mutation detection, Varscan (Koboldt, D. 2012) is used, and for strand-specific copy number variations, sequenceza (Favero F. 2015) is used.
[0028] BRCAness samples are identified. The most common BRCAness events are selected and labeled as positive samples. Specifically, these include: a. a pathogenic or probable pathogenic mutation in one allele of any one BRCA1 / 2 gene, resulting in loss of heterozygosity in another allele; b. two pathogenic or probable pathogenic mutations, i.e., loss of functionality, in any one BRCA1 / 2 gene; and c. loss of heterozygosity in one allele of BCRA1, resulting in methylation of the promoter region of another allele, of which methylation of the BCRA1 promoter region is obtained using pyrophosphate sequencing technology.
[0029] Confirm BRCAness-negative samples. The HRR-related mutant gene is wild-type, and furthermore, no LoH occurs in the corresponding gene. The Xiamen Aide Bio HRD test kit includes HRR-related genes such as ATM, FAM175A, FANCI, NBN, RAD51C, ATR, FANCA, FANCL, PALB2, RAD51D, ATRX, FANCC, FANCM, RAD50, RAD52, BAP1, FANCD2, KMT2D, RAD51, RAD54L, BARD1, FANCE, MDC1, RAD51B, SLX4, BLM, FANCF, MRE11A, WRN, XRCC2, BRCA1, FANCG, BRCA2, BRIP1, and EMSY.
[0030] The copy number variant fragments in this embodiment are classified according to the length of the mutation, including short fragments (5-10M), medium fragments (greater than 10 and less than or equal to 15M), and long fragments (>15M). The copy number variant fragments can also be classified according to the type of mutation, including loss of heterozygosity (LOH), unbalanced amplification of the mutant fragment (Allele-specific CNV, ASCNV), and balanced amplification of the mutant fragment (Balance CNV, BCNV). The copy number variant fragments can also be classified according to their location on the genome, including the location of the mutant fragment on the telomere side, and the location of the mutant fragment within the centromere region and the remaining other regions. Ultimately, the copy number variant fragments can be divided into 27 types (i.e., classification by length × classification by mutation type × classification by genomic location = 27).
[0031] Based on the above processing, the training set consists of 68 BCRAness samples and 42 negative samples, while the test set consists of 10 BRCAness samples and 8 negative samples. To prevent overfitting during training, only copy number variant fragments of types that occur more frequently than the number of samples in the training set are retained in the training samples. Subsequently, logistic regression is used to train the weights of the copy number variant fragment types based on the BRCAness type of the sample, thereby constructing a genomic scar model. The measured samples are calculated using the genomic scar model, and samples with a GSS score less than 0.5 are determined to be BRCAness-negative samples, while samples with a GSS score greater than 0.5 are determined to be BRCAness-positive samples.
[0032] In the test set, the GSS of the sample being measured is calculated using a genomic scar model to determine the BRCAness status of the sample. This is then compared to the pre-labeled BRCAness status of the samples in the test set. Of these, 10 BRCAness-positive samples could be accurately identified as positive by the genomic scar model, and 8 BRCAness-negative samples could be accurately identified as negative by the genomic scar model. In other words, the GSS of the genomic scar model can accurately predict the BRCAness status of a sample, with a sensitivity of 100%, specificity of 100%, and accuracy of 100%. [Examples]
[0033] The types of 191 test samples collected were FFPE samples from ovarian cancer patients and control blood samples. Library construction and capture were then performed using the homologous recombination repair deficiency (HRD) test kit from Xiamen Aide Biomedical Technology Co., Ltd., and sequencing was performed using an Illumina Novaseq sequencer. Subsequently, the GSS of the measured samples was calculated using the genomic scar model already trained in Example 1.
[0034] The relationship between the high GSS group and the HRR mutation group is summarized in the table below.
[0035] [Table 1]
[0036] Hypergeometric distribution testing showed that the high GSS group was significantly enriched with patients with HRR-related gene mutations, with a p-value of 0.0003. In addition, compared to HRR-related genes, the GSS model of the genome scar obtained in Example 1 can enrich a larger number of patients with genomic instability. [Examples]
[0037] The 44 test samples collected were FFPE samples from ovarian cancer patients and control blood samples, the initial postoperative treatment for these patients was platinum-based chemotherapy. Subsequently, library construction and capture were performed using the homologous recombination repair deficiency (HRD) testing kit from Xiamen Aide Biomedical Technology Co., Ltd., and sequencing was performed using an Illumina Novaseq sequencer. Finally, the GSS of the measured samples was calculated using the genomic scar model already trained in Example 1.
[0038] The effect of GSS in enriching the platinum-sensitive group was evaluated by comparing progression-free survival (PFS) in patients with high GSS and low GSS status. The results are shown in Figure 1. As shown in the figure, GSS status can be divided into high GSS (GSS+) and low GSS (GSS-). There was a significant difference in progression-free survival among patients, with platinum-based therapy significantly extending progression-free survival in the high GSS group (P=0.05). Among these, the median PFS was 11 months for patients in the high GSS group and 8.5 months for patients in the low GSS group. [Examples]
[0039] The collected 14 and 20 test samples were FFPE samples from ovarian cancer patients and control blood samples, both receiving primary maintenance therapy and subsequent therapy with PARPi, respectively. Library construction and capture were then performed using the homologous recombination repair deficiency (HRD) testing kit from Xiamen Aide Biomedical Technology Co., Ltd., followed by sequencing using an Illumina Novaseq sequencer. The GSS of the measured samples was then calculated using the genomic scar model already trained in Example 1.
[0040] As shown in Figure 2, among ovarian cancer patients who received primary maintenance therapy with PARPi, there was a significant difference in progression-free survival between the high GSS group (GSS+) and the low GSS group (GSS-). PARPi treatment significantly extended progression-free survival in the high GSS group (P=0.03), with a median PFS of 10.5 months for the high GSS group and 7 months for the low GSS group.
[0041] In ovarian cancer patients receiving PARPi as a second-line or later treatment, the objective response rate (ORR) in the high GSS group was 38.5% (5 / 13), the highest compared to the high HRD group (35.7%) and the BRCA mutation group (33.3%). The ORR in the low GSS group was 14.3% (1 / 7), the lowest compared to the low HRD group (16.7%) and the BRCA wild-type group (21.4%). The results are shown in Figure 3, indicating that the GSS in the genomic scar model obtained in Example 1 can enrich the PARPi-sensitive group.
[0042] The above description is merely a preferred embodiment of the present invention and therefore does not limit the scope of implementation of the present invention. In other words, equivalent modifications and alterations made based on the patent scope and specification of the present invention should all fall within the scope encompassed by the present invention. [Industrial applicability]
[0043] This invention discloses a method for constructing a genome scar model and its applications, comprising: (1) collecting known positive and negative samples to constitute a training set; (2) analyzing the CNV status of the training set to determine the type and corresponding quantity of CNVs; (3) determining BRCAness positive and negative events; and (4) using a machine learning method to train and obtain weights for different types of CNVs determined in step (2) based on the BRCAness positive and negative events of the training set, and then accumulating the weighted CNVs of different types to obtain a genome scar model used for calculating the GSS. This invention can accurately predict the BRCAness status of a sample being measured, and can directly interpret results based on the genome scar score, thus having industrial applicability.
Claims
1. A method for constructing a genomic scar model, The following steps (1) Collect BRCAness-positive and BRCAness-negative samples to form a training set; (2) Analyze the CNV status of the above training set and determine the type of CNV and the corresponding quantity; (3) Determine the BRCAness positive and BRCAness negative events of the training set; (4) Machine learning utilizes logistic regression to train the weights of CNV types based on the BRCAness type of the sample to construct a genomic scar model; based on the BRCAness-positive and BRCAness-negative events of the training set, the weights of the different types of CNV determined in step (2) are trained and obtained; and then the weighted CNVs of the different types are accumulated to obtain a genomic scar model used for calculating the GSS; (5) Collect other known BRCAness-positive and BRCAness-negative samples to construct a test set, and obtain the type and corresponding quantity of CNV of the test set based on step (2); (6) Substitute the results obtained in step (5) into the genome scar model obtained in step (4) to calculate the GSS of the test set, and further validate the genome scar model based on the GSS score; A method for constructing a genomic scar model, characterized by including, The aforementioned genome scar model includes a training set and a test set. The BRCAness positivity includes: in any one of BRCA1 / 2, a pathogenic or probably pathogenic mutation occurs in one allele and loss of heterozygosity occurs in another allele; in any one of BRCA1 / 2, two pathogenic or probably pathogenic mutations occur; in one allele of BRCA1, loss of heterozygosity occurs and methylation occurs in the promoter region of another allele; The aforementioned BRCAness positivity further includes related events in which other homologous recombination repair-related genes other than the BRCA1 / 2 genes cause genomic instability through mutation, loss of homozygosity, or silencing of the corresponding genes. The aforementioned BRCAness negativity means that the homologous recombination repair-related gene is wild-type, and furthermore, there is no loss of heterozygosity in the corresponding allele, or no methylation in its promoter region. The type of CNV in step (2) above is determined based on the length of the mutant fragment, the type of the mutant fragment, and the position of the mutant fragment on the genome. Here, the lengths of the mutant fragments are classified into short fragments of 5 to 10 M, medium fragments longer than 10 M and less than or equal to 15 M, and long fragments longer than 15 M. The types of mutant fragments are classified into categories including loss of heterozygosity, unbalanced amplification of mutant fragments, and balanced amplification of mutant fragments. A method for constructing a genomic scar model, wherein the genomic location of the mutated fragment is classified into types including the mutated fragment being located on the telomere side, the mutated fragment being located within the centromere region, and the mutated fragment being located outside the telomere side and within the centromere region.
2. The construction method according to claim 1, characterized in that a measurement sample is calculated using a genome scar model, samples with a GSS score less than 0.5 are determined to be BRCAness-negative samples, and samples with a GSS score greater than 0.5 are determined to be BRCAness-positive samples.
3. Use of a genomic scar model constructed by the construction method described in claim 1 in enriching the HRR mutation-associated patient population.
4. Use of a genomic scar model constructed by the construction method described in claim 1 in enriching the platinum-sensitive patient population.
5. Use of a genomic scar model constructed by the construction method described in claim 1 in enriching the PARPi-sensitive patient population.
Citation Information
Patent Citations
Methods and materials for evaluating loss of heterozygosity
JP2015506678A
Methods and Materials for Assessing Homologous Recombination Defects
JP2016523511A
Methods and Materials for Assessing Homologous Recombination Defects
JP2017533693A
An integrated machine-learning framework to estimate homologous recombination deficiency
WO2020168008A1