A method for detecting tumor single sample homologous recombination deficiency based on high-throughput sequencing
Patent Information
- Application Number
- CN202410985853.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2044-07-22
AI Technical Summary
但用WGS和WES评估HRD状态,各自存在缺陷,WGS评估HRD准确,但是数据冗余,成本过高,WES成本较低,但WES并不均匀覆盖整个基因组,因此在判断HRD上存在偏差,且目前判断HRD状态需要肿瘤单样本和对照样本配对计算,大大增加了分析成本
[0044] This invention provides a method for detecting homologous recombination defects in single tumor samples based on high-throughput sequencing. The invention first constructs a necessary database using high-throughput sequencing data from a normal population, then combines this database with SNP and copy number variation information from single tumor samples to simulate somatic SNP and copy number variation results. The simulated somatic SNP and copy number variation results are then analyzed together to calculate the HRD (Heart Rate Determination) of a single tumor sample. This method can calculate HRD values even when control samples are unavailable, and it is simple, easy to implement, and low-cost, providing a certain evidentiary basis for determining whether a patient is suitable for medications, including PARP inhibitors and platinum-based drugs.
Smart Images

Figure CN118942534B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of bioinformatics and high-throughput sequencing technology, specifically relating to a method for detecting homologous recombination defects in single tumor samples based on high-throughput sequencing. Background Technology
[0002] The human cell genome is formed by the double helix of DNA. During cell division, the double-stranded DNA unwinds, and the two strands serve as template strands for replication. Therefore, the accuracy of DNA replication is crucial for survival and reproduction. Accurate replication is not only the foundation of stable heredity in a species but also a guarantee for maintaining genome integrity and a necessary condition for cells to maintain basic growth and perform normal functions. However, DNA replication is not 100% accurate; the mismatch probability is approximately 10%. -10 Furthermore, environmental factors such as radiation, ultraviolet light irradiation, and chemical mutagens can also cause errors or damage to DNA sequences. When single-strand mismatches or damage occur in DNA, there are three main repair mechanisms: base cleavage repair, nucleotide cleavage repair, and mismatch repair. These primarily use the parent strand as a template to repair the damaged area. When double-strand mismatches or damage occur in DNA, there are two repair mechanisms: homologous recombination repair and non-homologous end joining repair. Non-homologous end joining repair directly links the two broken ends. During the repair process, some bases may be randomly introduced or removed, making it a high-efficiency but low-accuracy repair mechanism. Homologous recombination repair uses the homologous sequence of the undamaged sister chromatid as a template, making it a low-efficiency but high-accuracy repair mechanism.
[0003] When homologous recombination repair malfunctions, resulting in homologous recombination-deficient (HRD), double-strand damage is primarily repaired via non-homologous recombination. This leads to errors in DNA, such as insertions, deletions, and copy number abnormalities. Over time, these errors accumulate, causing apoptosis or transformation into cancer cells. HRD occurs in various tumors, commonly including ovarian cancer, bladder cancer, breast cancer, endometrial cancer, prostate cancer, and pancreatic cancer. Known causes of HRD include mutations in the BRCA1 / 2 genes, BRCA1 promoter methylation, mutations in other genes involved in the homologous recombination repair pathway, and downregulation of BRCA gene expression due to other factors.
[0004] Since HRD is caused by multiple factors, simply detecting mutations in genes along the homologous repair and recombination pathway is insufficient to fully assess the HRD status. Studies have found that when cells develop HRD, they over-rely on non-homologous recombination for repair, introducing insertions, deletions, and copy number abnormalities that result in specific traces on the genome, known as genomic scarring. Genomic scarring manifests in three forms: loss of heterozygosity, large fragment migration, and telomere allele imbalance. A combination of these three indicators can determine the genomic HRD status. Tumor patients with high HRD are highly sensitive to platinum-based drugs that cause DNA breaks, as well as PARP inhibitors, especially ovarian cancer patients. HRD has become one of the important biomarkers associated with PARP inhibitor therapy in ovarian cancer.
[0005] Early assessments of HRD status used SNP microarray technology, but this was expensive. With the development of next-generation sequencing (NGS), the costs of whole-genome sequencing (WGS) and whole-exome sequencing (WES) have gradually decreased, and HRD assessment has increasingly shifted to NGS platforms. However, both WGS and WES have their limitations. WGS is accurate but suffers from data redundancy and high costs, while WES is cheaper but does not uniformly cover the entire genome, leading to biases in HRD assessment. Furthermore, current HRD assessment requires paired calculations of single tumor samples and control samples, significantly increasing analysis costs. In addition, factors such as difficulty in obtaining adjacent normal control samples for certain cancer types, improper preservation of control samples, or the inability to obtain control samples due to patient-related reasons prevent the determination of patients' HRD status, impacting subsequent medication guidance. Summary of the Invention
[0006] The purpose of this invention is to provide a method for detecting homologous recombination defects in a single tumor sample based on high-throughput sequencing, wherein the detection method can realize the detection of homologous recombination defects based on a single tumor sample.
[0007] This invention provides a method for detecting homologous recombination defects in single tumor samples based on high-throughput sequencing, comprising the following steps:
[0008] Provide BAM files for the normal population;
[0009] SNP detection and analysis are performed based on the BAM file of the normal population to obtain the SNP information of each sample in the normal population. The SNP information of each sample in the normal population is then merged and normalized to obtain the SNP database of the normal population.
[0010] The BAM files of the normal population were subjected to copy number variation detection and analysis, and the obtained copy number variation data of the normal population was normalized to obtain a copy number variation database of the normal population.
[0011] High-throughput sequencing and quality control were performed on single tumor samples. The obtained single tumor sample data was compared and filtered with the human reference genome to obtain the single tumor sample BAM file.
[0012] The tumor single-sample BAM files were subjected to SNP detection analysis and copy number variation detection analysis to obtain tumor single-sample SNP information and tumor single-sample copy number variation information, respectively.
[0013] The SNP database of the normal population is supplemented into the tumor single-sample SNP information to obtain simulated somatic SNP data;
[0014] The copy number variation database of the normal population is supplemented into the single-sample copy number variation information of the tumor to obtain simulated somatic copy number variation data;
[0015] By jointly analyzing the simulated somatic SNP data and the simulated somatic copy number variation data, the results of homologous recombination defect detection were obtained.
[0016] Preferably, the method for obtaining the SNP database of the normal population includes the following steps:
[0017] The reads with an alignment quality greater than 1 in the BAM file of the normal population are counted, and the number of bases that are consistent with and inconsistent with the human reference genome at the alignment coverage sites of the reads are counted to form a data file;
[0018] The data file was analyzed to obtain SNP information for each sample in the normal population. The SNP information includes the chromosome where the SNP is located, the position of the SNP on the chromosome, the bases (ref and alt) that are consistent with the human reference genome, and the number of sequencing reads supporting the ref and alt bases.
[0019] The SNP information of each sample in the normal population is merged, the SNP sites of different normal population samples are normalized, the frequency of the normalized SNP sites is calculated, and the SNP information of the normal population is obtained.
[0020] The depth information of each site of each sample in the normal population BAM file is calculated, and the site depth information of each sample in the normal population is merged and averaged to obtain the depth information of the normal population.
[0021] The SNP information and depth information of the normal population are integrated to obtain the SNP database of the normal population.
[0022] Preferably, the step of integrating the SNP information and depth information of the normal population includes: if the depth information loci in the depth information of the normal population also exist in the SNP information of the normal population, then supplement the depth information loci with the corresponding SNP information and the normalized ref and alt read support counts and mutation frequencies of the SNP information; for depth information loci not detected in the depth information of the normal population, then the normalized depth is assigned to the ref read support count, and the alt read support count and mutation frequency are 0.
[0023] Preferably, the method for obtaining the copy number variation database of the normal population includes any one of the following three methods;
[0024] Method 1: 1) Apply a 120bp window to the capture region in the normal population BAM file; 2) Calculate the homogenization depth of the normal population samples within the 120bp window to obtain the normal population copy number variation database.
[0025] Method 2: 1) Slide a 120bp window over the capture region in the normal population BAM file; 2) Calculate the depth information and depth information log2 value within the 120bp sliding window capture region in step 1) to form the depth information file of the normal population; 3) Integrate the depth information files of the normal population to obtain the normal population copy number variation database.
[0026] Method 3: 1) Apply a 120bp window to the capture region in the normal population BAM file; 2) For the human reference genome, log2 is equal to 0 and depth is equal to 1 within the 120bp window of the capture region to construct the normal population copy number variation database.
[0027] Preferably, the step of obtaining the simulated somatic SNP variation data includes: based on the position coordinates of the SNP in the tumor single-sample SNP information, finding the SNP mutation with the same position coordinates in the SNP database of the normal population, and supplementing it after the corresponding SNP in the tumor single-sample SNP information to obtain the simulated somatic SNP variation data.
[0028] Preferably, the step of obtaining the simulated somatic copy number variation data includes any one of the following two items:
[0029] The first step: Based on the window coordinates in the tumor single-sample copy number variation information, find the depth information with the same coordinates in the normal population copy number variation database, and append it to the end of the corresponding window in the tumor single-sample copy number variation information; and calculate the log2 value of the normalized depth of the tumor single sample and the normal population within the same sliding window and place it after the corresponding window to obtain the simulated somatic copy number variation data; the normal population copy number variation database is obtained according to method one in the above technical solution.
[0030] The second step: Based on the copy number variation database of the normal population, the BAM file of the tumor single sample is detected to detect copy number variation, and simulated somatic copy number variation data is obtained. The copy number variation database of the normal population is obtained according to method two or method three in the above technical solution.
[0031] Preferably, during the high-throughput sequencing process, a DNA library is constructed using a genome-wide equidistant SNP site capture method, with an interval of 50Kb between SNP sites and 52,000 SNP sites.
[0032] Preferably, the quality control parameters include the following:
[0033] 1) Sequencing sequences in which more than 40% of the bases with a mass less than Q15 have been removed;
[0034] 2) Sequencing sequences with more than 5 N bases removed;
[0035] 3) Use a 4bp sliding window to slide from the end to the front of the sequencing sequence. If the base quality of all bases in the window is lower than Q20, then cut the sequence; otherwise, stop sliding.
[0036] 4) Remove sequencing sequences that are less than 15 bp in length after pruning low-quality bases in step 3).
[0037] Preferably, the human reference genome includes the hg19 reference genome.
[0038] Preferably, obtaining the homologous recombination defect includes: analyzing the simulated somatic SNP data and simulated somatic copy number variation data to obtain the number of loss of heterozygosity, large fragment migration, and telomere allele imbalance in a single tumor sample, and calculating the homologous recombination defect according to a predetermined formula (1):
[0039] Homologous recombination defect = number of heterozygous deletions + number of large fragment migrations + number of telomere allele imbalances, formula (1).
[0040] The present invention also provides a computer-readable storage medium storing a computer program or instructions, wherein the computer program or instructions, when executed by a processor, implement the steps of the detection method described above.
[0041] The present invention also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of the detection method described above.
[0042] The present invention also provides the application of the detection method described in the above technical solution, the computer-readable storage medium described in the above technical solution, or the computer program product described in the above technical solution in the preparation of products that guide the use of tumor drugs.
[0043] Beneficial effects:
[0044] This invention provides a method for detecting homologous recombination defects in single tumor samples based on high-throughput sequencing. The invention first constructs a necessary database using high-throughput sequencing data from a normal population, then combines this database with SNP and copy number variation information from single tumor samples to simulate somatic SNP and copy number variation results. The simulated somatic SNP and copy number variation results are then analyzed together to calculate the HRD (Heart Rate Determination) of a single tumor sample. This method can calculate HRD values even when control samples are unavailable, and it is simple, easy to implement, and low-cost, providing a certain evidentiary basis for determining whether a patient is suitable for medications, including PARP inhibitors and platinum-based drugs. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the embodiments will be briefly described below.
[0046] Figure 1 This is a flowchart of the single-sample HRD calculation in Example 1;
[0047] Figure 2 This is a graph showing the correlation results of single and double sample HRD values in the ICGC database in Example 2.
[0048] Figure 3 This is a graph showing the correlation results of HRD values for single and dual samples of real ovarian cancer in Example 3. Detailed Implementation
[0049] This invention provides a method for detecting homologous recombination defects in single tumor samples based on high-throughput sequencing, comprising the following steps:
[0050] Provide BAM files for the normal population;
[0051] SNP detection and analysis are performed based on the BAM file of the normal population to obtain the SNP information of each sample in the normal population. The SNP information of each sample in the normal population is then merged and normalized to obtain the SNP database of the normal population.
[0052] The BAM files of the normal population were subjected to copy number variation detection and analysis, and the obtained copy number variation data of the normal population was normalized to obtain a copy number variation database of the normal population.
[0053] High-throughput sequencing and quality control were performed on single tumor samples. The obtained single tumor sample data was compared and filtered with the human reference genome to obtain the single tumor sample BAM file.
[0054] The tumor single-sample BAM files were subjected to SNP detection analysis and copy number variation detection analysis to obtain tumor single-sample SNP information and tumor single-sample copy number variation information, respectively.
[0055] The SNP database of the normal population is supplemented into the tumor single-sample SNP information to obtain simulated somatic SNP data;
[0056] The copy number variation database of the normal population is supplemented into the single-sample copy number variation information of the tumor to obtain simulated somatic copy number variation data;
[0057] By jointly analyzing the simulated somatic SNP data and the simulated somatic copy number variation data, the results of homologous recombination defect detection were obtained.
[0058] This invention preferably performs high-throughput sequencing and quality control on normal population samples, comparing and filtering the obtained normal population sample data with a human reference genome to obtain a normal population BAM file. In the high-throughput sequencing process described in this invention, a genome-wide equidistant SNP site capture method is preferably used to construct the DNA library, with a preferred spacing of 50 kb between SNP sites and a preferred number of SNP sites (52,000). The genome-wide equidistant SNP site capture method is preferably performed using a genome-wide SNP backbone panel, HiSNPUltraPanel, which has been disclosed in Chinese patent CN202210278698.8. The high-throughput sequencing described in this invention is preferably next-generation sequencing; the sequencing depth is preferably 100× or higher. The quality control parameters of this invention preferably include the following: 1) removing sequencing sequences with more than 40% of bases having a base quality less than Q15; 2) removing sequencing sequences with more than 5 N bases; 3) using a 4bp sliding window from the end to the beginning of the sequencing sequence, removing sequences when the base quality within the window is all below Q20, otherwise stopping the sliding; 4) removing sequencing sequences less than 15bp in length after trimming low-quality bases. The quality control parameters of this invention can ensure high data quality and reliability. The specific process of quality control in this invention is not particularly limited; conventional operating procedures can be followed according to the above parameters. The human reference genome of this invention is preferably the hg19 reference genome. This invention preferably uses filtering to mark duplicates and correct base quality on the aligned data file. This invention does not particularly limit the software used for quality control, alignment, and filtering; conventional quality control, alignment, and filtering software in the art can be used.
[0059] The normal population samples described in this invention preferably include non-tumor samples, and more preferably include, but are not limited to, normal tissue and / or blood samples from tumor patients and tissue and / or blood samples from normal individuals.
[0060] After obtaining the normal population BAM file, the present invention performs SNP detection and analysis based on the normal population BAM file, statistically analyzes the SNP information of each sample in the normal population, and merges and normalizes them to obtain the SNP database of the normal population.
[0061] The preferred method for constructing the SNP database of the normal population according to the present invention includes: counting reads with an alignment quality greater than 1 in the normal population BAM file, and counting the number of bases consistent with and inconsistent with the human reference genome at the alignment coverage sites of the reads, forming a data file; analyzing the data file to obtain the SNP information of each sample in the normal population, wherein the SNP information includes the chromosome where the SNP is located, the position of the SNP on the chromosome, the bases consistent with the human reference genome (ref bases) and the mutant bases (alt bases) of the SNP, and the number of sequencing reads supporting the ref and alt bases; merging the SNP information of each sample in the normal population, normalizing the SNP sites of different normal population samples, calculating the frequency of the normalized SNP sites, and obtaining the SNP information of the normal population; calculating the depth information of each site of each sample in the normal population BAM file, merging the site depth information of each sample in the normal population, taking the average value, and obtaining the depth information of the normal population; integrating the SNP information and the depth information of the normal population to obtain the SNP database of the normal population.
[0062] This invention preferably uses the mpileup module of the samtools software to count reads with an alignment quality greater than 1 in the BAM files of the normal population, and counts the number of bases that are consistent with and inconsistent with the human reference genome at the alignment coverage sites of these reads, forming a data file. This invention does not have specific limitations on the specific process of this statistical analysis; the conventional usage of the mpileup module of the samtools software is sufficient.
[0063] After obtaining the data file, the present invention preferably uses the pileup2snp module of VarScan software to analyze the base ratio under each comparison site in the data file to obtain the SNP information of each sample in the normal population. The SNP information includes the chromosome where the SNP is located, the position of the SNP on the chromosome, the bases (ref and alt) that are consistent with the human reference genome, and the number of sequencing reads supported by the ref and alt bases.
[0064] After obtaining the SNP information of each sample in the normal population, the present invention preferably merges the SNP information of each sample in the normal population, normalizes the SNP sites of different normal population samples, calculates the frequency of the normalized SNP sites, and obtains the SNP information of the normal population. The normalization method of the present invention preferably includes: when the same SNP is detected in multiple samples, summing the support numbers of the ref and alt reads of the same SNP site detected in different samples in the normal population, and dividing by the total number of normal population samples for normalization; simultaneously, when an SNP is detected only in one sample, directly dividing the ref and alt reads of the unique SNP site detected in the normal population by the total number of normal population samples for normalization.
[0065] Furthermore, this invention preferably uses the depth module of the samtools software to calculate the depth information of each site of each sample in the normal population BAM file, and then merges the site depth information of each sample in the normal population, taking the average depth to obtain the depth information of the normal population. This invention does not have any particular limitations on the process of obtaining the depth information of the normal population using the depth module of the samtools software; conventional methods in the art can be used.
[0066] After obtaining the depth information of the normal population, the present invention preferably integrates the SNP information and the depth information of the normal population to obtain the SNP database of the normal population. The integration step of the present invention preferably includes: if the depth information loci in the depth information of the normal population also exist in the SNP information of the normal population, then supplement the depth information loci with corresponding SNP information, as well as the normalized ref and alt read support counts and mutation frequencies of the SNP information; for depth information loci not detected in the depth information of the normal population, then the normalized depth is assigned to the ref read support count, and the alt read support count and mutation frequency are 0.
[0067] After obtaining the normal population BAM file, this invention performs copy number variation detection and analysis on the normal population BAM file, and normalizes the obtained normal population copy number variation data to obtain a normal population copy number variation database. The construction steps of the normal population copy number variation database of this invention preferably include any one of the following three methods;
[0068] Method 1: 1) Apply a 120bp window to the capture region in the normal population BAM file; 2) Calculate the homogenization depth of the normal population samples within the 120bp window to obtain the normal population copy number variation database.
[0069] Method 2: 1) Slide a 120bp window over the capture region in the normal population BAM file; 2) Calculate the depth information and depth information log2 value within the 120bp sliding window capture region in step 1) to form the depth information file of the normal population; 3) Integrate the depth information files of the normal population to obtain the normal population copy number variation database.
[0070] Method 3: 1) Slide window the captured region with a 120bp window; 2) The human reference genome has log2 equal to 0 and depth equal to 1 within the 120bp sliding window of the captured region, thus constructing a normal population copy number variation database.
[0071] In the method one of the present invention, the calculation step of the homogenization depth in step 2) of the present invention preferably includes: statistically analyzing the depth information of the normal population in the 120bp sliding window capture area in step 1), and dividing it by the average depth, that is, the average depth information of the normal population, and the resulting homogenization depth of the normal population in the 120bp sliding window capture area is the normal copy number variation database.
[0072] In the second method described in this invention, the present invention preferably uses the coverage module of the cnvkit software to calculate the depth information and the depth information log2 value within the 120bp sliding window capture area in step 2). The present invention preferably uses the reference module of the cnvkit software to perform the integration described in step 3).
[0073] Of the three methods described in this invention, this invention preferably uses the reference module of the cnvkit software to perform the calculation described in step 2).
[0074] While constructing SNP and copy number variation databases for the normal population, this invention performs high-throughput sequencing and quality control on single tumor samples. The obtained single tumor sample data is then compared and filtered with the human reference genome to obtain a single tumor sample BAM file. The high-throughput sequencing, quality control, comparison, and filtering process for single tumor samples described in this invention is preferably the same as the corresponding process for normal population samples described in the above technical solution, and will not be repeated here; the high-throughput sequencing depth is preferably 500× or higher.
[0075] After obtaining the tumor single-sample BAM file, this invention performs SNP detection analysis and copy number variation detection analysis on the tumor single-sample BAM file to obtain tumor single-sample SNP information and tumor single-sample copy number variation information, respectively. This invention preferably uses the pileup2snp module of VarScan software for the SNP detection analysis. The copy number variation analysis of this invention preferably includes: using the depth function of samtools software to calculate the depth of each site in the capture region of the tumor single sample, and calculating the overall average depth in the capture region; performing a sliding window operation on the capture region with a 120bp window, and calculating the average depth within the 120bp sliding window; and normalizing the average depth within the 120bp sliding window by dividing it by the overall average depth in the capture region to obtain the tumor single-sample copy number variation information.
[0076] After obtaining the tumor single-sample SNP information, tumor single-sample copy number variation information, normal population SNP database, and normal population copy number variation database, this invention supplements the normal population SNP database into the tumor single-sample SNP information to obtain simulated somatic SNP data; simultaneously, this invention supplements the normal population copy number variation database into the tumor single-sample copy number variation information to obtain simulated somatic copy number variation data.
[0077] The method for obtaining simulated somatic SNP data according to the present invention preferably includes: based on the position coordinates of the SNP in the tumor single-sample SNP information, finding the SNP mutation with the same position coordinates in the SNP database of the normal population, and supplementing it after the corresponding SNP in the tumor single-sample SNP information to obtain the simulated somatic SNP variation data.
[0078] The preferred steps for obtaining the simulated somatic copy number variation data of the present invention include either of the following two items: First, based on the coordinates of the window in the tumor single-sample copy number variation information, find the depth information of the same coordinates in the normal population copy number variation database, and append it to the end of the corresponding window in the tumor single-sample copy number variation information; and calculate the log2 value of the normalized depth of the tumor single sample and the normal population within the same sliding window and place it after the corresponding window to obtain the simulated somatic copy number variation data; the normal population copy number variation database is obtained according to method one of the above technical solutions.
[0079] The second step: Based on the copy number variation database of the normal population, the BAM file of the tumor single sample is detected to detect copy number variation, and simulated somatic copy number variation data is obtained. The copy number variation database of the normal population is obtained according to method two or method three in the above technical solution.
[0080] In the second aspect of this invention, the present invention preferably uses the batch module of the cnvkit software to detect copy number mutations.
[0081] After obtaining the simulated somatic copy number variation data and simulated somatic SNP data, this invention uses the simulated somatic SNP data and simulated somatic copy number variation data for joint analysis to obtain the homologous recombination defect detection results. Specifically, the preferred method for obtaining the homologous recombination defect of this invention includes: analyzing the simulated somatic SNP data and simulated somatic copy number variation data to obtain the number of loss of heterozygosity, large fragment migration, and telomere allele imbalance in a single tumor sample, and calculating the homologous recombination defect according to a predetermined formula (1):
[0082] Homologous recombination defect = number of heterozygous deletions + number of large fragment migrations + number of telomere allele imbalances, formula (1).
[0083] The present invention also provides a computer-readable storage medium storing a computer program or instructions, wherein the computer program or instructions, when executed by a processor, implement the steps of the detection method described above.
[0084] The present invention also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of the detection method described above.
[0085] The detection method described in this invention can also calculate HRD values for single tumor data where control samples cannot be obtained. The method is simple, easy to implement, and low in cost, providing a certain evidence basis for determining whether a patient is suitable for drugs including PARP inhibitors and platinum-based drugs.
[0086] Based on the above advantages, this invention also provides the application of the detection method described in the above technical solutions, the computer-readable storage medium described in the above technical solutions, or the computer program product described in the above technical solutions in the preparation of products that guide the use of oncology drugs. The oncology drugs described in this invention preferably include, but are not limited to, platinum-based drugs and / or PARP inhibitors.
[0087] To further illustrate the present invention, the technical solutions provided by the present invention will be described in detail below with reference to the accompanying drawings and embodiments, but these should not be construed as limiting the scope of protection of the present invention.
[0088] Example 1
[0089] A method for calculating homologous recombination defect (HRD) in a single human tumor sample based on next-generation sequencing, the steps of which are as follows, and the flowchart is shown below. Figure 1 As shown
[0090] 1. DNA was extracted from tumor single samples and normal population samples, and hybridization capture library construction was performed using the whole genome SNP backbone panel HiSNP Ultra Panel v1.0 (the panel is disclosed in Chinese patent CN202210278698.8);
[0091] 2. Perform next-generation NGS sequencing on the hybridization capture libraries;
[0092] 3. Perform quality control on the second-generation NGS sequencing data. Remove sequencing sequences with more than 40% of bases having a quality of less than Q15, and remove sequencing sequences with more than 5 N bases. Use a 4bp sliding window to slide from the end to the beginning of the sequencing sequence. If the average quality of bases in the window is lower than Q20, remove it; otherwise, stop sliding. Remove sequencing sequences with a length of less than 15bp after trimming low-quality bases to ensure high data quality and reliability.
[0093] 4. Align the high-quality data of the tumor single sample and normal population sample after quality control to the human reference genome hg19, mark the repetitive sequences, and perform base quality recorrection to obtain the BAM file for subsequent analysis;
[0094] 5. Construction of a normal population mutant SNP database:
[0095] (1) The BAM files of normal population samples were analyzed using the mpileup module of samtools (version: 1.9). The reads that were aligned to the genome and had an alignment quality MQ greater than 1 were counted. The number of identical and inconsistent bases in these sequences at the genome alignment coverage sites was counted.
[0096] (2) Using the pileup2snp module of VarScan (version: 2.4.3) software, analyze the base ratio of each alignment site in the file obtained in (1), and then obtain the SNP information of each normal population sample: the chromosome where the SNP is located, the position of the SNP on the chromosome, the ref base and the mutated alt base of the SNP that are consistent with the genome, the number of sequencing reads supported by the ref and alt bases, and the mutation frequency calculated based on the ref and alt reads;
[0097] (3) Merge the SNPs detected in all normal samples, sum the support number of the ref and alt reads of the same SNP site detected in different samples, and divide by the total number of samples for normalization. Divide the ref and alt reads of the unique SNP site detected in the sample directly by the total number of samples for normalization, thereby obtaining the SNP information of the normal population.
[0098] (4) Using the depth module of samtools software, calculate the depth information of each site within the capture panel range for the BAM file of the normal population sample, then merge the site depth information of all normal samples and divide by the number of samples to obtain the depth information of the normal population.
[0099] (5) Integrate the SNP information and depth information of the normal population. If the depth information loci are also present in the SNP information, supplement the depth information loci with the corresponding SNP information and the normalized ref and alt read support counts. For depth information loci not detected in the SNP information, the normalized depth is assigned to the ref read support count, and the alt read support count and mutation frequency are 0, thus obtaining the SNP information database of the normal population.
[0100] 6. There are three methods for constructing a database of copy number variations in normal populations:
[0101] Method 1:
[0102] (1) First, use a 120bp window to slide the capture area of the capture panel;
[0103] (2) Calculate the average depth of the normal population obtained in step 5;
[0104] (3) Calculate the mean depth of the normal group in the 120bp sliding window capture area in step 5, and divide it by the average depth to obtain the uniform depth of the normal group in the 120bp sliding window capture area.
[0105] Method 2:
[0106] (1) First, use a 120bp window to slide the capture area of the capture panel;
[0107] (2) For the BAM files of normal population samples, the depth and log2 value of the depth information in the 120bp capture area were calculated using the coverage module of the cnvkit (version: 0.9.9) software.
[0108] (3) The reference module of the cnvkit software integrates the depth information files of all normal samples to generate a normal population copy number variation database.
[0109] Method 3:
[0110] (1) First, use a 120bp window to slide the capture area of the capture panel;
[0111] (2) The reference module of the cnvkit software calculates the normal population copy number variation database within the inner 120bp sliding window of the reference genome.
[0112] 7. Simulated somatic SNP results from a single tumor sample
[0113] (1) The mpileup module of samtools software was used to analyze the BAM files of tumor single samples, and the base changes of sequences with an alignment quality greater than 1 in the BAM files were statistically analyzed at the covered sites on the genome.
[0114] (2) Using the pileup2snp module of VarScan software, analyze the file obtained in (1) to obtain the SNP information of a single tumor sample;
[0115] (3) The SNP sites in the tumor single sample are supplemented with SNP database data of the normal population to obtain simulated somatic SNP files;
[0116] Specifically, the supplement refers to finding mutations with the same location coordinates in the SNP database of the normal population based on the location coordinates of the SNP in the tumor single sample, and adding them after the SNP information of the tumor single sample. This results in a single SNP line containing both the ref and alt reads information of the tumor single sample and the ref and alt reads information of the normal population.
[0117] 8. Results of simulated somatic copy number variation in a single tumor sample
[0118] Method 1:
[0119] (1) Use the depth function of samtools software to calculate the depth of each site in the capture region of a single tumor sample;
[0120] (2) Calculate the overall average depth of a single tumor sample within the capture area;
[0121] (3) Use a 120bp window to slide the capture area of the capture panel, calculate the average depth within the 120bp sliding window, and divide it by the overall average depth of the tumor single sample for homogenization.
[0122] (4) Integrate the normal population copy number variation database obtained by method one in step 6 with the copy number variation information of tumor single samples, and calculate the log2 value of the homogenization depth of tumor single samples and normal population within the same sliding window to obtain simulated somatic copy number variation information.
[0123] Specifically, the integration refers to: finding depth information at the same location coordinates in the normal population copy number variation database based on the coordinates of the tumor single sample window, and adding it after the tumor single sample window; at the same time, taking log2 of the ratio of the normalized depth of the tumor sample and the normal population, and also placing it after this window information.
[0124] Method 2: Using the batch module of the cnvkit software, based on the normal population copy number variation database obtained in Method 2 in step 6, copy number variations are detected in the BAM file of a single tumor sample to obtain simulated somatic copy number variation information.
[0125] Method 3: Using the batch module of the cnvkit software, based on the normal population copy number variation database obtained in Method 3 in step 6, copy number variations are detected in the BAM file of a single tumor sample to obtain simulated somatic copy number variation information.
[0126] 9. Calculate HRD: Based on the somatic mutation SNPs and somatic copy number variation information simulated in a single tumor sample, use sequenza (version: 2.1.2) and ScarHRD (version: 0) for analysis to obtain the number of loss of heterozygosity, large fragment migration, and telomere allele imbalance in a single tumor sample. The sum of the three is the homologous recombination defect HRD value.
[0127] Example 2
[0128] This embodiment uses WGS BAM data of 70 paired ovarian cancer tumors and controls downloaded from the International Cancer Genome Consortium (ICGC) database. Data within the capture region of HiSNP Ultra Panel v1.0 is extracted from the WGS BAM file to simulate the capture hybridization data of Hi SNP Ultra Panel v1.0.
[0129] Paired HRD analysis: For BAM data captured by simulated HiSNPs in paired tumor and control samples, the mpileup module of the samtools software was used to statistically analyze the base changes at genomic alignment positions for reads with alignment quality greater than 1. Then, the somatic module of VarScan software was used to jointly analyze the tumor and control samtools mpileup outputs to obtain somatic mutation SNP results. The copynumber module of VarScan software was used to obtain copy number variation results based on the tumor and control samtools mpileup outputs. Sequenza and scarHRD were used to analyze somatic mutation SNPs and copy number variations to obtain the paired calculated tumor sample purity and HRD value.
[0130] The HRD of a single sample was analyzed using the steps in Example 1: A SNP database and a copy number variation database for the normal population were constructed from the control samples of 70 paired samples. Then, the purity and HRD values of the tumor samples from the 70 paired samples were obtained using the SNP database from the normal population and the copy number variation databases constructed by three methods. The three methods for constructing the copy number variation databases are:
[0131] Method 1: The normal population copy number variation database is generated by homogenizing the normal population within the 120bp sliding window capture region;
[0132] Method 2: The cnvkit software calculates and integrates 120bp sliding window capture region depth information files for all normal samples to generate a normal population copy number variation database;
[0133] Method 3: The cnvkit software was used to calculate the genomic copy number variation database within a 120bp sliding window of the hg19 reference genome.
[0134] The paired and single-sample purity correlations of HiSNP simulation data from 70 ICGC samples are shown in [reference needed]. Figure 2 In (a), M1 represents method one, M2 represents method two, and M3 represents method three. Figure 2 As shown in (a), the correlation R between the purity calculated from a single sample and the purity calculated from paired samples is high under the three different methods for constructing copy number variation databases. 2 All values are greater than 0.85, indicating that the purity calculated from a single sample can be used for data filtering and screening.
[0135] Studies have shown that samples with low tumor purity contain too many normal cells, resulting in a significant deviation between the calculated HRD value and the actual tumor HRD status. Therefore, it is generally believed that the calculated HRD value is more reliable when the tumor purity is greater than 0.3. After filtering out samples with an HRD less than or equal to 0.3 based on the purity calculated from a single sample, the consistency between the HRD values calculated from single samples and paired samples is shown in [the figure]. Figure 2 In (b), under three different methods of constructing copy number variation databases, the correlation R... 2 The values are mostly around 0.93. This example shows that the single-sample HRD value calculated by this method is highly consistent with the HRD value calculated by paired samples.
[0136] Example 3
[0137] The HRD of real tumor samples was detected and calculated using the method disclosed in Example 1. The specific steps are as follows:
[0138] DNA was extracted from 81 FFPE samples of real ovarian cancer tumors and 84 FFPE samples adjacent to ovarian cancer tumors. Hybridization capture and sequencing were performed using HiSNP UltraPanel v1.0, and the sequencing data were compared with the genome after quality control. 84 normal control samples were used to construct a population SNP database and a copy number variation database. The three detection methods for the copy number variation database are as follows:
[0139] Method 1: The normal population copy number variation database is generated by homogenizing the normal population within the 120bp sliding window capture region;
[0140] Method 2: The cnvkit software calculates and integrates 120bp sliding window capture region depth information files for all normal samples to generate a normal population copy number variation database;
[0141] Method 3: The cnvkit software was used to calculate the genomic copy number variation database within a 120bp sliding window of the hg19 reference genome.
[0142] Paired HRD analysis: For 81 real ovarian cancer tumor samples and corresponding control samples, the mpileup module of Samtools software was used to analyze the base changes at genomic alignment positions of reads with an alignment quality greater than 1. Then, the somatic mutation SNP results were obtained by combining the Samtools mpileup outputs from tumor and control samples using the somatic module of VarScan software. Copy number variation results were obtained by the copy number module of VarScan software based on the Samtools mpileup outputs from tumor and control samples. Sequenza and ScarHRD were used to analyze somatic mutation SNPs and copy number variations to obtain the paired calculated tumor sample purity and HRD values.
[0143] Single-sample HRD analysis: The single-sample HRD values of 81 real ovarian cancer tumor samples were calculated using SNP and copy number variation databases constructed from SNP and copy number variation databases of 84 real control samples and simulated HiSNP data of 70 ICGC normal samples in Example 1.
[0144] A database of SNPs and copy number variations was constructed from 84 real ovarian cancer control samples. Single-sample HRDs were calculated for 81 real ovarian cancer tumor samples. After removing samples with a purity of less than 0.3 calculated from single samples, the correlation between the calculated HRD values and those from paired samples is shown in [the table below]. Figure 3 In (a), M1 represents method one, M2 represents method two, and M3 represents method three. Figure 3 As shown in (a), under the three different methods of constructing the copy number variation database, the correlation R between the HRD value calculated for a single sample and the HRD value of the paired sample is [missing information]. 2 All values are greater than 0.9, showing a high degree of consistency.
[0145] When the control sample used as the population is replaced with simulated data from ICGC, the correlation between the HRD calculated for single samples and the HRD calculated for paired samples is shown in [reference needed]. Figure 3 (b), where M1 represents Method 1, M2 represents Method 2, and M3 represents Method 3. The figure shows the correlation R between the HRD value calculated for a single sample and the HRD value of the paired sample under three different copy number variation databases and different control populations. 2 The result remained at 0.89. This example demonstrates that the single-sample HRD calculated by this tumor single-sample calculation method is highly correlated and consistent with the two-sample HRD. Even when using reference SNP and copy number variation databases constructed from different normal populations, the results remain consistent, indicating that single-sample HRD calculation can be used to replace paired-sample HRD.
[0146] Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, and not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.
Claims
1. A method for detecting homologous recombination defects in single tumor samples based on high-throughput sequencing, characterized in that, Includes the following steps: Provide BAM files for the normal population; SNP detection and analysis are performed based on the BAM file of the normal population to obtain the SNP information of each sample in the normal population. The SNP information of each sample in the normal population is then merged and normalized to obtain the SNP database of the normal population. The BAM files of the normal population were subjected to copy number variation detection and analysis, and the obtained copy number variation data of the normal population was normalized to obtain a copy number variation database of the normal population. High-throughput sequencing and quality control were performed on single tumor samples. The obtained single tumor sample data was compared and filtered with the human reference genome to obtain the single tumor sample BAM file. The tumor single-sample BAM files were subjected to SNP detection analysis and copy number variation detection analysis to obtain tumor single-sample SNP information and tumor single-sample copy number variation information, respectively. The SNP database of the normal population is supplemented into the tumor single-sample SNP information to obtain simulated somatic SNP data; The copy number variation database of the normal population is supplemented into the single-sample copy number variation information of the tumor to obtain simulated somatic copy number variation data; By combining the simulated somatic SNP data and the simulated somatic copy number variation data, the results of homologous recombination defect detection were obtained.
2. The detection method according to claim 1, characterized in that, The method for obtaining the SNP database of the normal population includes the following steps: The reads with an alignment quality greater than 1 in the BAM file of the normal population are counted, and the number of bases that are consistent with and inconsistent with the human reference genome at the alignment coverage sites of the reads are counted to form a data file; The data file was analyzed to obtain SNP information for each sample in the normal population. The SNP information includes the chromosome where the SNP is located, the position of the SNP on the chromosome, the bases (ref and alt) that are consistent with the human reference genome, and the number of sequencing reads supporting the ref and alt bases. The SNP information of each sample in the normal population is merged, the SNP sites of different normal population samples are normalized, the frequency of the normalized SNP sites is calculated, and the SNP information of the normal population is obtained. The depth information of each site of each sample in the normal population BAM file is calculated, and the site depth information of each sample in the normal population is merged and averaged to obtain the depth information of the normal population. The SNP information and depth information of the normal population are integrated to obtain the SNP database of the normal population.
3. The detection method according to claim 2, characterized in that, The step of integrating the SNP information and the depth information of the normal population includes: if the depth information loci in the depth information of the normal population also exist in the SNP information of the normal population, then supplement the depth information loci with the corresponding SNP information and the number of reads and mutation frequency of the normalized ref and alt of the SNP information. For depth information sites not detected in the depth information of the normal population, the normalized depth is assigned to the number of refreads, and the number of reads and mutation frequency of alt are 0.
4. The detection method according to claim 1, characterized in that, The method for obtaining the copy number variation database of the normal population includes any one of the following three methods; Method 1: 1) Use a 120bp window to slide across the capture region in the normal population BAM file; 2) Calculate the homogenization depth of the normal population sample within a 120bp sliding window to obtain the normal population copy number variation database; Method 2: 1) Slide a 120bp window over the capture area in the normal population BAM file; 2) Calculate the depth information and depth information log2 value within the 120bp sliding window capture area in step 1) to form the depth information file of the normal population. 3) Integrate the deep information files of the normal population to obtain a normal population copy number variation database; Method 3: 1) Use a 120bp window to slide across the capture area; 2) Within a 120bp capture region of the human reference genome, log2 is equal to 0 and depth is equal to 1, thus constructing a normal population copy number variation database.
5. The detection method according to any one of claims 1-3, characterized in that, The steps for obtaining the simulated somatic SNP data include: based on the position coordinates of the SNP in the tumor single-sample SNP information, finding the SNP mutation with the same position coordinates in the SNP database of the normal population, and adding it to the end of the corresponding SNP in the tumor single-sample SNP information to obtain the simulated somatic SNP variation data.
6. The detection method according to claim 4, characterized in that, The steps for obtaining the simulated somatic copy number variation data include either of the following two: The first step: Based on the coordinates of the window in the tumor single-sample copy number variation information, find the depth information of the same position coordinates in the normal population copy number variation database, and supplement it after the corresponding window in the tumor single-sample copy number variation information; and calculate the log2 value of the normalized depth of the tumor single sample and the normal population within the same sliding window and place it after the corresponding window to obtain the simulated somatic copy number variation data; the normal population copy number variation database is obtained according to the method in claim 4. The second step: Based on the copy number variation database of the normal population, the BAM file of the tumor single sample is detected to detect copy number variation, thereby obtaining simulated somatic copy number variation data. The copy number variation database of the normal population is obtained according to method two or method three in claim 4.
7. The detection method according to claim 1, characterized in that, During the high-throughput sequencing process, a DNA library was constructed using a genome-wide equidistant SNP site capture method, with an interval of 50Kb between SNP sites and a total of 52,000 SNP sites.
8. The detection method according to claim 1, characterized in that, The quality control parameters include the following: 1) Sequencing sequences in which more than 40% of the bases with a mass less than Q15 have been removed; 2) Sequencing sequences with more than 5 N bases removed; 3) Use a 4bp sliding window to slide from the end to the front of the sequencing sequence. If the base quality of all bases in the window is lower than Q20, then cut the sequence; otherwise, stop sliding. 4) Remove sequencing sequences that are less than 15 bp in length after pruning low-quality bases in step 3).
9. The detection method according to claim 1, characterized in that, The human reference genome includes the hg19 reference genome.
10. The detection method according to claim 1, characterized in that, The homologous recombination defect is obtained by analyzing the simulated somatic SNP data and simulated somatic copy number variation data to determine the number of loss of heterozygosity, large fragment migration, and telomere allele imbalance in a single tumor sample, and then calculating the homologous recombination defect according to a predetermined formula (1): Homologous recombination defect = number of heterozygous deletions + number of large fragment migrations + number of telomere allele imbalances, formula (1).
11. A computer-readable storage medium storing a computer program or instructions thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of any of the detection methods described in claims 1-10.
12. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the detection method according to any one of claims 1-10.
13. The use of the detection method according to any one of claims 1-10, the computer-readable storage medium according to claim 11, or the computer program product according to claim 12 in the preparation of products for guiding the use of oncology drugs.
Citation Information
Patent Citations
Method and device for constructing multi-population non-exon region SNP (Single Nucleotide Polymorphism) probe set
CN114678067A
Single-sample whole-genome allele specific copy number variation prediction method
CN112802548A
NGS platform-based homologous recombination defect detection method and quality control system
CN113658638A