Method and device for calculating whole genome copy number in tumor single sample
Patent Information
- Application Number
- CN202610867743.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-06-16
AI Technical Summary
[0009]本发明的主要目的在于提供一种肿瘤单样本中全基因组拷贝数的计算方法及装置,以解决现有技术中缺乏适用于全基因组、且在无对照样本条件下具有高准确性的等位基因拷贝数分析方法的问题
[0030]应用本发明的技术方案,利用上述肿瘤单样本中全基因组拷贝数的计算方法,能够对高通量测序数据(包括但不限于全基因组测序或外显子组测序)中肿瘤样本的拷贝数变异及等位基因不平衡进行分析。该计算方法在无需严格匹配正常对照样本的情况下,通过构建基线(参考)等位基因深度并结合等位基因频率信息,实现对肿瘤基因组结构变异的准确解析。
Smart Images

Figure CN122455093B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and more specifically, to a method and apparatus for calculating the whole genome copy number in a single tumor sample. Background Technology
[0002] In clinical practice and scientific research, obtaining matched normal samples (such as peripheral blood or adjacent normal tissue) from cancer patients presents numerous practical difficulties. On the one hand, in actual diagnosis and treatment, some patients only receive tumor tissue samples, especially paraffin-embedded (FFPE) samples or small-volume puncture samples, often making it impossible to obtain high-quality normal controls simultaneously. On the other hand, in retrospective studies or multicenter data integration analyses, historical samples generally lack information on matched normal samples. Furthermore, collecting additional normal samples not only increases the burden on patients but also involves ethical approval, sample management, and cost issues. Therefore, in many real-world application scenarios, only single-sample tumor data is available for analysis.
[0003] However, allele copy number analysis is crucial in tumor research and clinical decision-making. This analysis can identify not only loss of heterozygosity (LOH) and copy-neutral LOH (cnLOH), but also calculate homologous recombination deficiency (HRD) scores, further supporting the inference of tumor purity, ploidy, and clonal structure. Therefore, accurately recovering allele-level copy number information without control samples remains a key technical challenge in the field.
[0004] Currently, several mature algorithms have been developed in the field of allele copy number analysis. These methods are typically based on joint modeling of tumor samples and matched normal samples. Representative methods include ASCAT, FACETS, and CNVkit.
[0005] The core principles of the aforementioned methods typically include two key observables: LogR (log-ratio signal), which reflects the copy number change in tumor samples relative to normal samples; and BAF (allele frequency), which describes the proportional relationship between alleles. By combining LogR and BAF signals, these methods can achieve joint inference of total copy number and allele copy number, and further estimate tumor purity and ploidy. Among them, the ASCAT method, through its Joint Segmentation Functional Combination (ASPCF) and grid search strategy, is widely considered one of the standard methods in this field.
[0006] However, the methods described above all rely on matching normal samples to determine germline heterozygous sites, thus ensuring the interpretability of BAF signals. Without control samples, it is difficult to distinguish whether BAF signals originate from true allelic imbalance or from germline homozygosity, leading to significant uncertainty in the analytical results.
[0007] To address the challenges of single-sample analysis, some methods have attempted to approximate normal sample information by introducing reference data or constructing panels of normals (PoNs). Furthermore, while existing technologies have proposed single-sample analysis schemes based on target region sequencing data, these schemes are primarily applicable to specific panel regions and cannot be extended to the whole genome.
[0008] Therefore, current technology still lacks a method for allele copy number analysis that is applicable to the whole genome and has high accuracy in the absence of control samples. Summary of the Invention
[0009] The main objective of this invention is to provide a method and apparatus for calculating the whole genome copy number in a single tumor sample, in order to solve the problem that there is a lack of existing methods for allele copy number analysis that are applicable to the whole genome and have high accuracy in the absence of control samples.
[0010] To achieve the above objectives, according to a first aspect of the present invention, a method for calculating the whole genome copy number in a single tumor sample is provided. The method includes: S1) obtaining the copy number signal and allele frequency of a single tumor sample based on a population single nucleotide polymorphism (SNP) site dataset; S2) performing a first segmentation analysis using the copy number signal and the allele frequency to obtain a first segment set; S3) identifying low heterozygous density regions in the first segment set, and screening a corrected heterozygous site set from the population SNP site dataset based on the distribution of allele frequencies of the single tumor sample; S4) performing a second segmentation analysis using the corrected heterozygous site set and the copy number signal to obtain a second segment set; S5) calculating the purity and ploidy of the single tumor sample based on the second segment set, and calculating the whole genome copy number in the single tumor sample.
[0011] Further, S1) above includes: S1-1) obtaining single nucleotide polymorphism (SNP) sites from the population frequency database, extracting genotype information of each SNP site in the population, calculating the GC content of upstream and downstream windows of each SNP site, and obtaining a standard analysis site dataset; obtaining sequencing data of normal samples, calculating the sequencing depth of the SNP sites in the sequencing data, and recording it as the normal sample sequencing depth; the standard analysis site dataset and the normal sample sequencing depth constitute the population SNP dataset; S1-2) obtaining sequencing data of the tumor single sample; calculating the sequencing depth of each SNP site in the sequencing data of the tumor single sample, and recording it as the tumor single sample sequencing depth; calculating the number of supporting reads for each allele in the sequencing data of the tumor single sample; calculating the log ratio signal LogR, where LogR = log2. R R = sequencing depth of the tumor single sample ÷ sequencing depth of the normal sample; S1-3) Correct the above LogR, the correction includes: correcting the above LogR using the above GC content in the above standard analysis site dataset, and normalizing the corrected signal to obtain the corrected log ratio signal LogR. corr , which is the copy number signal mentioned above; S1-4) Use the model to judge the genotype status of each of the above single nucleotide polymorphism sites in the above standard analysis site dataset, obtain a set of credible heterozygous sites, and calculate the allele frequency of the above set of credible heterozygous sites.
[0012] Furthermore, the models in S1-4 above include Hidden Markov Models (HMMs).
[0013] Furthermore, S2 above includes: utilizing a piecewise algorithm, and using the aforementioned LogR corr Using the allele frequencies of the aforementioned set of credible heterozygous sites as the synchronous analysis object, the genome of the aforementioned single tumor sample was subjected to the aforementioned first segmentation process to obtain the aforementioned first segment set composed of different genomic fragments.
[0014] Further, S3 above includes: calculating the number N of heterozygous single nucleotide polymorphism sites in the genome segments in the first segment set; calculating the SNP density of each genome segment, wherein the SNP density = N ÷ length of the genome segment × 10 6 If the above SNP density is <5, calculate the mutation abundance (VAF) distribution of the above single nucleotide polymorphism sites, select the secondary significant peaks other than the main peaks at 0.5 or 1.0, and select the SNP sites located in the interval of the above secondary significant peaks as correction candidate sites and obtain the set of corrected heterozygous sites, and calculate the allele frequency of the above set of corrected heterozygous sites.
[0015] Furthermore, S4 above includes: utilizing a piecewise algorithm, and using the aforementioned LogR corr Using the allele frequencies of the aforementioned set of corrected heterozygous sites as synchronous analysis objects, the genome of the aforementioned single tumor sample was subjected to the aforementioned second segmentation process to obtain the aforementioned second segment set composed of different genomic fragments.
[0016] Further, S5 above includes: performing a grid search for purity and ploidy for each of the aforementioned genomic fragments in the aforementioned second segment set to obtain a candidate solution list; inputting the aforementioned candidate solution list and the corresponding genomic fragments in the aforementioned second segment set into ABSOLUTE software to determine the purity and ploidy of the aforementioned single tumor sample; and using the purity and ploidy of the aforementioned single tumor sample to calculate the whole genome copy number in the aforementioned single tumor sample.
[0017] Furthermore, the above calculation method also includes: iteratively optimizing the allele frequencies by repeating steps S1-4)-S4) in the above calculation method, thereby detecting copy number-neutral loss of heterozygosity or loss of heterozygosity in high tumor purity samples.
[0018] To achieve the above objectives, according to a second aspect of the present invention, a device for calculating the whole genome copy number in a single tumor sample is provided. The device includes: a copy number and allele frequency acquisition unit, a first segmentation analysis unit, a corrected heterozygous site screening unit, a second segmentation analysis unit, and an allele copy number calculation unit; wherein the copy number and allele frequency acquisition unit is used to acquire the copy number signal and allele frequency signal of a single tumor sample based on a population single nucleotide polymorphism (SNP) site dataset; and the first segmentation analysis unit is used to perform a first segmentation analysis using the copy number signal and the allele frequency signal. A first segment set is obtained; the corrected heterozygous site screening unit is used to identify low heterozygous density regions in the first segment set and, based on the distribution of allele frequencies in the tumor single sample, screen out a corrected heterozygous site set from the population single nucleotide polymorphism site dataset; the second segment analysis unit is used to perform a second segment analysis using the corrected heterozygous site set and the copy number signal to obtain a second segment set; the allele copy number calculation unit is used to calculate the purity and ploidy of the tumor single sample based on the second segment set, and calculate the whole genome allele copy number in the tumor single sample.
[0019] Furthermore, the copy number and allele frequency acquisition unit includes: a standard analysis site dataset construction module, a sequencing depth calculation module, a signal correction module, and an allele frequency calculation module; wherein, the standard analysis site dataset construction module is used to obtain single nucleotide polymorphism (SNP) sites in the population frequency database, extract the genotype information of each SNP site in the population, calculate the GC content of the upstream and downstream windows of each SNP site, and obtain a standard analysis site dataset; and obtain sequencing data of normal samples, calculate the sequencing depth of the SNP sites in the normal samples, and record it as the normal sample sequencing depth, forming the population SNP site dataset; the sequencing depth calculation module is used to obtain sequencing data of a single tumor sample to be tested, calculate the sequencing depth of each SNP site in the sequencing data of the single tumor sample, and record it as the tumor single sample sequencing depth; calculate the number of supporting reads of each allele in the sequencing data of the single tumor sample; and calculate the log ratio signal LogR, where LogR = log2. R R = sequencing depth of the tumor single sample ÷ sequencing depth of the normal sample; the signal correction module is used to correct the LogR using the GC content in the standard analysis site dataset, and outputs the corrected log-ratio signal LogR after normalization. corr As the copy number signal mentioned above; the allele frequency calculation module is used to infer the genotype status of each single nucleotide polymorphism site in the standard analysis site dataset using the model, obtain a set of credible heterozygous sites, and calculate the allele frequency of the set of credible heterozygous sites.
[0020] Furthermore, the model in the allele frequency calculation module mentioned above includes a hidden Markov model.
[0021] Furthermore, the aforementioned first segmentation analysis unit includes a first segmentation algorithm execution module, which is used to execute the segmentation algorithm to achieve the aforementioned LogR corr Using the allele frequencies of the aforementioned set of credible heterozygous sites as the synchronous analysis object, the genome of the aforementioned single tumor sample was subjected to the aforementioned first segmentation process to obtain the aforementioned first segment set composed of different genomic fragments.
[0022] Furthermore, the aforementioned corrected heterozygous site screening unit includes: a SNP density calculation module, a subsignificant peak identification module, and a corrected candidate site screening module; wherein, the aforementioned SNP density calculation module is used to calculate the number N of heterozygous single nucleotide polymorphism sites in each genomic fragment in the aforementioned first segment set, and calculate SNP density = N ÷ length of the aforementioned genomic fragment × 10 6The aforementioned secondary significant peak identification module is used to calculate the mutation abundance distribution of the aforementioned single nucleotide polymorphism sites when the aforementioned SNP density is less than 5, and to identify secondary significant peaks other than the main peaks at 0.5 or 1.0. The aforementioned correction candidate site screening module is used to select single nucleotide polymorphism sites located within the aforementioned secondary significant peak intervals as correction candidate sites, form a set of correction heterozygous sites, and calculate the allele frequencies of the aforementioned set of correction heterozygous sites.
[0023] Furthermore, the aforementioned second segmentation analysis unit includes a second segmentation algorithm execution module, which is used to execute the segmentation algorithm to achieve the aforementioned LogR corr Using the allele frequencies of the aforementioned set of corrected heterozygous sites as synchronous analysis objects, the genome of the aforementioned single tumor sample was subjected to the aforementioned second segmentation process to obtain the aforementioned second segment set composed of different genomic fragments.
[0024] Furthermore, the aforementioned allele copy number calculation unit includes: a grid search module, used to perform a grid search on each of the aforementioned genomic fragments in the aforementioned second segment set for purity and ploidy, to obtain a candidate solution list; an ABSOLUTE fitting module, used to input the aforementioned candidate solution list and the corresponding genomic fragments in the aforementioned second segment set into ABSOLUTE software to determine the purity and ploidy of the aforementioned single tumor sample; and a copy number inference module, used to calculate the whole genome allele copy number of the aforementioned single tumor sample based on the aforementioned purity and ploidy.
[0025] Furthermore, the aforementioned device is configured to cyclically execute the aforementioned allele frequency calculation module, the aforementioned first segmentation analysis unit, the aforementioned corrected heterozygous site screening unit, and the aforementioned second segmentation analysis unit to iteratively optimize the aforementioned allele frequencies, thereby detecting copy number-neutral loss of heterozygosity or loss of heterozygosity in high tumor purity samples.
[0026] To achieve the above objectives, according to a third aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method for calculating the whole genome copy number in a single tumor sample.
[0027] To achieve the above objectives, according to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, wherein when the computer program / instructions are executed by a processor, the steps of the method for calculating the whole genome copy number in a single tumor sample are implemented.
[0028] To achieve the above objectives, according to a fifth aspect of the present invention, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of the method for calculating the whole genome copy number in a single tumor sample.
[0029] To achieve the above objectives, according to a sixth aspect of the present invention, a method for calculating the whole genome copy number in a single tumor sample, or a device for calculating the whole genome copy number in a single tumor sample, or a computer device, or a computer-readable storage medium, or a computer program product, is provided for application in any one or more of the following: pan-cancer HRD detection, tumor evolution analysis, or precision medicine assistance.
[0030] By applying the technical solution of this invention and utilizing the aforementioned method for calculating the whole-genome copy number in a single tumor sample, copy number variations and allele imbalances in tumor samples from high-throughput sequencing data (including but not limited to whole-genome sequencing or exome sequencing) can be analyzed. This calculation method, without requiring strict matching of normal control samples, achieves accurate analysis of tumor genome structural variations by constructing a baseline (reference) allele depth and combining it with allele frequency information. Attached Figure Description
[0031] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0032] Figure 1 A flowchart of an optional calculation method according to an embodiment of the present invention is shown.
[0033] Figure 2 A graph showing the results of single-sample allele copy number according to Embodiment 1 of the present invention is provided.
[0034] Figure 3 A graph showing the allele copy number results of paired samples according to Embodiment 1 of the present invention is displayed.
[0035] Figure 4 Allele copy number maps of a single sample according to Comparative Example 1 of the present invention are shown.
[0036] Figure 5 A graph showing the density or average number of heterozygous sites on each chromosome according to Comparative Example 1 of the present invention is provided.
[0037] Figure 6 The mutation abundance results of SNP sites on each chromosome according to Comparative Example 1 of the present invention are shown.
[0038] Figure 7The purity ploidy cross-validation results of Comparative Example 1 according to the present invention are shown in the figure.
[0039] Figure 8 A graph showing the results of single-sample allele copy number according to Embodiment 2 of the present invention is provided.
[0040] Figure 9 A graph showing the allele copy number results of paired samples according to Embodiment 2 of the present invention is displayed.
[0041] Figure 10 Allelic copy number maps of a single sample according to Comparative Example 2 of the present invention are shown.
[0042] Figure 11 A graph showing the density or average number of heterozygous sites on each chromosome according to Comparative Example 2 of the present invention is provided.
[0043] Figure 12 The mutation abundance results of SNP sites on each chromosome according to Comparative Example 2 of the present invention are shown.
[0044] Figure 13 The purity ploidy cross-validation results of Comparative Example 2 according to the present invention are shown in the figure. Detailed Implementation
[0045] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the embodiments.
[0046] Terminology Explanation:
[0047] Allele frequency (BAF): In a specific population, the proportion of a particular allele at a given gene locus out of all alleles at that locus.
[0048] Allele-specific copy number (ASCN): The actual number of physical copies of a particular allele at a specific gene locus in the genome of a tested cell.
[0049] ASPCF: Allele-specific segmental constant fitting.
[0050] Mutation abundance (VAF): The percentage of DNA fragments carrying a certain mutated sequence at a specific DNA site out of the total number of DNA fragments tested at that site.
[0051] SNP: Single nucleotide polymorphism.
[0052] LOH: Loss of heterozygosity.
[0053] cnLOH: Copy-neutral loss of heterozygosity (cnLOH).
[0054] HRD: homologous recombination deficiency.
[0055] As mentioned in the background section, existing methods perform well in calculating allele copy numbers in tumor samples with control samples, but they still have significant limitations in practical applications. First, most methods heavily rely on matched normal samples, which is often difficult to achieve in clinical and historical data. Second, under single-sample conditions, it is difficult to accurately distinguish between germline homozygous sites and tumor-driven LOH events, leading to ambiguity in the interpretation of BAF signals. Furthermore, existing single-sample methods are mostly limited to target region sequencing data, cannot operate stably across the entire genome, and are quite sensitive to tumor purity and sequencing noise.
[0056] In this application, the inventors attempt to develop a novel computational method that, by introducing a reference dataset and combining it with statistical inference strategies, accurately identifies heterozygous sites in single tumor samples, thereby restoring the effectiveness of the BAF signal and further achieving allele copy number genotyping across the entire genome. Based on this, a series of protection schemes are proposed in this application.
[0057] In a first typical embodiment of this application, a method for calculating the whole genome copy number in a single tumor sample is provided. The method includes: S1) obtaining the copy number signal and allele frequency of a single tumor sample based on a population single nucleotide polymorphism (SNP) site dataset; S2) performing a first segmentation analysis using the copy number signal and allele frequency to obtain a first segment set; S3) identifying low heterozygous density regions in the first segment set, and screening a corrected heterozygous site set from the population SNP site dataset based on the distribution of allele frequencies of the single tumor sample; S4) performing a second segmentation analysis using the corrected heterozygous site set and the copy number signal to obtain a second segment set; S5) calculating the purity and ploidy of the single tumor sample based on the second segment set, and calculating the whole genome copy number in the single tumor sample.
[0058] In the above calculation method, the population single nucleotide polymorphism (SNP) dataset consists of pre-screened high-confidence germline SNP sites from a population reference database. It does not rely on paired normal samples but instead uses a standardized analytical coordinate system constructed from population genotype frequencies and local GC content characteristics as a benchmark reference for inferring tumor genomic structural variations under no-control conditions. The copy number signal is the logarithmic ratio (LogR) signal corrected for and normalized by GC content. corr The ), reflects the relative copy number change of tumor genome regions relative to the population baseline; the allele frequency signal is the distribution of B allele frequency (BAF) calculated at identifiable heterozygous sites, used to characterize the relative abundance imbalance among alleles.
[0059] The first segmentation analysis in S2 divides the genome into several continuous segments with uniform signals by jointly modeling the copy number signal and the initial BAF signal, forming the first segment set, which is used to preliminarily determine the genome copy number status and potential heterozygous loss regions.
[0060] In S3, the "low heterozygosity density region" refers to the interval in the first segment set where the number of analyzable SNP sites per unit genome length (e.g., per Mb) is lower than a preset threshold. In such regions, the sparse heterozygous sites lead to unreliable BAF signals, which can easily cause misjudgment or missed detection of LOH events. In this type of region, this method does not rely on the preset VAF threshold. Instead, it analyzes the frequency distribution of BAF values in the tumor sample itself to identify newly emerging subsignificant peaks in addition to the normal heterozygous peak (≈0.5) and homozygous peak (≈0 or 1.0). These peaks correspond to the VAF shift caused by the monoallelic expression of germline heterozygous sites due to LOH events. Then, these sites that were misjudged as homozygous in the first round of analysis are selected and reconstructed from the population SNP dataset to form a corrected heterozygous site set.
[0061] The second segmentation analysis in S4 re-inputs the corrected set and the original copy number signal into the segmentation algorithm to achieve accurate identification and boundary correction of hidden LOHs (especially copy number neutral LOHs) in low hybrid density regions, thereby improving the completeness and accuracy of the segmentation results.
[0062] In S5, based on the more stable copy number breakpoints and allele composition patterns obtained from the second segment set, combined with grid search and model fitting for purity and ploidy, the diploid ploidy and cell purity parameters of tumor samples are inferred, and finally the absolute allele copy number (AA, AB, BB, etc.) of each segment interval is obtained, achieving genome-wide allele-level copy number genotyping relying solely on a single tumor sample. This method overcomes the dependence of traditional methods on paired normal samples, enabling systematic analysis of tumor genome allele structural variations under conditions without control samples.
[0063] In a preferred embodiment, S1) includes: S1-1) obtaining single nucleotide polymorphism (SNP) sites from a population frequency database, extracting genotype information (i.e., population genotype information) for each SNP site in the population, calculating the GC content of upstream and downstream windows for each SNP site, and obtaining a standard analysis site dataset; obtaining sequencing data of normal samples, calculating the sequencing depth of the SNP sites in the sequencing data, and denoting it as the normal sample sequencing depth; the standard analysis site dataset and the normal sample sequencing depth constitute the population SNP dataset; S1-2) obtaining sequencing data of a single tumor sample; calculating the sequencing depth of each SNP site in the sequencing data of the single tumor sample, and denoting it as the single tumor sample sequencing depth; calculating the number of supporting reads for each allele in the sequencing data of the single tumor sample; calculating the logarithmic ratio signal LogR, where LogR = log2. R R = sequencing depth of the tumor single sample ÷ sequencing depth of the normal sample; S1-3) Correct the above LogR, the correction includes: correcting the above LogR using the above GC content in the above standard analysis site dataset, and normalizing the corrected signal to obtain the corrected log ratio signal LogR. corr , which is the copy number signal mentioned above; S1-4) Use the model to judge the genotype status of each of the above single nucleotide polymorphism sites in the above standard analysis site dataset, obtain a set of credible heterozygous sites, and calculate the allele frequency (BAF) of the above set of credible heterozygous sites.
[0064] S1 above includes: S1-1) Screening high-confidence, high-polymorphism SNP loci from population frequency databases (such as 1000 Genomes), extracting their population genotype frequencies and allele distributions, and combining the GC content within upstream and downstream (including but not limited to 25bp, 50bp, 100bp, 200bp, 500bp, 1kb, 2kb, or 5kb upstream and downstream) of each locus to construct a standardized analysis locus dataset for unified coordinate system and system bias modeling; simultaneously collecting sequencing data from multiple high-quality normal samples, calculating their average sequencing depth on this locus set, forming a stable and reusable "reference depth baseline," which, together with the locus information, constitutes a population single nucleotide polymorphism locus dataset as an intrinsic reference for controlless analysis. S1-2) Performing high-throughput sequencing on single tumor samples, extracting their total sequencing depth and A / B allele support read counts on this locus set, and calculating the log ratio signal LogR = log2(tumor depth / reference depth) to quantify relative copy number changes. (S1-3) To eliminate systematic sequencing bias caused by GC content, a local regression model was constructed based on the nonlinear relationship between LogR and GC content to correct LogR. Quantile normalization was then used to eliminate sequencing depth differences between samples, yielding the corrected copy number signal LogR. corr To ensure signal comparability and stability, Hidden Markov Models (HMMs) were used in S1-4. Using population genotype frequencies as a priori and combining them with allele read distributions in tumor samples, the potential germline genotype status (AA, AB, BB) of each SNP locus was inferred. Only loci with a posterior probability ≥0.9 were retained as the "credible heterozygous locus set," and their allele frequencies (BAF) were calculated accordingly. These allele frequencies served as the core input for subsequent joint segmentation analysis, achieving high-precision allele signal reconstruction in unpaired normal samples.
[0065] The genotype information in the "population genotype information" mentioned above refers to the reference and mutated bases of SNPs, such as A / C. The "normal sample" mentioned above refers to the normal control sample, that is, the sample without copy number variation, which is generally taken from blood or normal tissue (non-tumor tissue).
[0066] In a preferred embodiment, the model in S1-4 above includes a Hidden Markov Model (HMM).
[0067] In a preferred embodiment, S2) above includes: utilizing a segmentation algorithm (including but not limited to ASPCF) and using the above LogR corr Using the allele frequencies of the aforementioned set of credible heterozygous sites as the synchronous analysis object, the genome of the aforementioned single tumor sample was subjected to the aforementioned first segmentation process to obtain the aforementioned first segment set composed of different genomic fragments.
[0068] The above S2) includes: using a segmentation algorithm (such as ASPCF) and using the above LogR corr Using the allele frequencies of the aforementioned set of credible heterozygous sites as the synchronous analysis object, the genome of the aforementioned single tumor sample was subjected to the first segmentation process described above, resulting in the aforementioned first segment set composed of different genomic fragments. ASPCF is a joint segmentation method based on a Hidden Markov Model (HMM), which simultaneously models the copy number signal (LogR). corr The algorithm infers the potential copy number state and allele composition (e.g., AA, AB, AAB) of each region in the genome space by combining the distribution of copy number amplification, deletion, and loss of heterozygosity (LOH) with the allele frequency (BAF). It uses a set of reliable heterozygous sites constructed from population references as the BAF input source, avoiding dependence on paired normal samples. Through optimization of state transition probabilities and observational likelihood functions, it achieves collaborative identification of copy number amplification, deletion, and LOH boundaries, outputting an initial genome segmentation structure and providing a basic framework for subsequent correction of low heterozygous density regions.
[0069] In a preferred embodiment, S3) includes: calculating the number N of heterozygous single nucleotide polymorphism sites in the genome segments of the first segment set; calculating the SNP density of each genome segment, wherein the SNP density = N ÷ length of the genome segment × 10 6 If the above SNP density is <5, calculate the mutation abundance (VAF) distribution of the above single nucleotide polymorphism sites, select the secondary significant peaks other than the main peaks at 0.5 or 1.0, and select the SNP sites located in the interval of the above secondary significant peaks as correction candidate sites and obtain the set of corrected heterozygous sites, and calculate the allele frequency of the above set of corrected heterozygous sites.
[0070] The above S3) includes: calculating the number N of SNP sites that have been determined to be credibly heterozygous in each genomic segment in the first segment set, and then calculating the SNP density per unit length (Mb), defined as N divided by the segment length and then multiplied by 10. 6The abundance of allele information in a region is quantified. When the SNP density is below 5, it indicates that the BAF signal is unreliable due to the sparseness of sites in the region, which may easily mask the true LOH event. In this case, this method does not rely on a fixed threshold, but is based on the mutation abundance (VAF) distribution of all detectable SNPs in the tumor sample. The main peak and the secondary significant peak are identified by kernel density estimation. The main peak is located at 0.5 (heterozygous) or 1.0 (homozygous), while the secondary significant peak usually appears in the 0.7–1.0 range, indicating that there is unilateral allele amplification in the region due to LOH. Although the corresponding site was misclassified as homozygous in the first round, it is actually germline heterozygous, but one allele has been selectively lost by the tumor clone. This method selects SNPs located in the VAF range of the secondary significant peak as "correction candidate sites", constructs a "correction heterozygous site set", and recalculates its allele frequency (BAF) to be included in the second round of segmented analysis, thereby achieving accurate recall of occult events such as copy number neutral LOH (cnLOH) and improving the integrity of whole-genome allele typing.
[0071] Optionally, machine learning models can also be introduced in S3 above to predict heterozygous sites, or haplotype information can be combined to improve the accuracy of determination.
[0072] The aforementioned "sub-significant peak interval" includes mutations within 0.1 of the VAF peak value.
[0073] In a preferred embodiment, S4) above includes: utilizing a segmentation algorithm (ASPCF) and using the above LogR corr Using the allele frequencies of the aforementioned set of corrected heterozygous sites as synchronous analysis objects, the genome of the aforementioned single tumor sample was subjected to the aforementioned second segmentation process to obtain the aforementioned second segment set composed of different genomic fragments.
[0074] The above S4) includes: using the segmented algorithm (ASPCF) and with the above LogR corr Using the allele frequencies of the aforementioned corrected heterozygous loci as synchronous analysis objects, the genome of the aforementioned single tumor sample was subjected to the aforementioned second segmentation process to obtain the aforementioned second segment set composed of different genomic fragments. ASPCF is a joint segmentation algorithm based on a hidden Markov model, which simultaneously models the corrected copy number signal (LogR). corrThe algorithm, together with optimized allele frequencies (BAF), jointly infers the copy number status and allele composition (e.g., AA, AB, AAB) of each genomic region in the state space. The second segment, starting from the results of the first round, introduces corrected heterozygous sites selected by VAF distribution-driven screening, significantly enhancing sensitivity to low-density regions and copy number-neutral LOHs (cnLOHs). The algorithm optimizes fragment boundaries and state assignments based on signal co-change patterns, enabling regions that were previously blurred due to sparse sites to obtain clear classifications. The final output is a second segment set with high resolution and biological consistency, providing a precise structural basis for subsequent absolute copy number calculations.
[0075] In a preferred embodiment, S5) includes: performing a grid search on each of the genomic fragments in the second segment set for purity and ploidy to obtain a candidate solution list; inputting the candidate solution list and the corresponding genomic fragments in the second segment set into ABSOLUTE software to determine the purity and ploidy of the tumor single sample; and using the purity and ploidy of the tumor single sample to calculate the whole genome copy number in the tumor single sample.
[0076] The above S5) includes: LogR for each genomic fragment in the second segment set. corr A purity-ploidy grid search was performed with the BAF signal. The system traversed discrete combinations of tumor purity (0–1) and genomic ploidy (1–5), fitted the expected copy number state of each fragment, and generated a candidate solution list. This list, along with the allelic composition of the fragments, was input into the ABSOLUTE software. This software, based on a Bayesian framework, models the genome-wide copy number load distribution and the BAF co-occurrence pattern, evaluates the posterior probability of each candidate solution, and screens the optimal combination of purity and ploidy parameters. Finally, using these optimal parameters as a benchmark, the absolute allelic copy number of each fragment (e.g., AA=2, AAB=3, etc.) was calculated in reverse, achieving a precise conversion from relative signals to the true copy number state within tumor cells. This constructed a control-free allelic-specific copy number map across the entire genome, providing a foundation for HRD assessment and clonal evolution analysis.
[0077] In a preferred embodiment, the above calculation method further includes: iteratively optimizing the allele frequencies by repeating steps S1-4)-S4) in the above calculation method, thereby detecting cnLOH or LOH in high tumor purity samples.
[0078] The above calculation method also includes iteratively optimizing allele frequencies by repeatedly executing steps S1-4) to S4) to improve the detection sensitivity of copy number neutral loss of heterozygosity (cnLOH) and LOH events in high-tumor-purity samples. After the initial segmentation, some heterozygous sites that were missed due to weak BAF deviation or low site density may still have biological significance when the proportion of tumor cells is high or the LOH region is narrow. The iterative process dynamically updates the set of credible heterozygous sites based on the current segmentation results, recalculates the BAF distribution, and relaxes the genotype inference threshold in subsequent rounds to introduce secondary peak signals and gradually recall germline heterozygous sites that were misclassified as homozygous. Each iteration re-estimates the posterior probability of the genotype through the HMM model, so that the BAF signal gradually converges to the true tumor allele structure, thereby achieving adaptive identification and accurate typing of occult cnLOH and high-purity LOH without normal controls, significantly enhancing the completeness and robustness of the analysis.
[0079] In a preferred embodiment, the above-mentioned high tumor purity samples include samples with a tumor purity of 80% or higher.
[0080] Figure 1 A flowchart of an optional calculation method according to the present invention is shown. Figure 1 In this context, SNP represents single nucleotide polymorphism, LogR represents the log ratio signal, BAF represents the frequency of the B allele, VAF represents the frequency of the mutation site, and ASPCF represents the allele-specific segmental constant fitting.
[0081] In a second typical embodiment of this application, a device for calculating the whole genome copy number in a single tumor sample is provided. The device includes: a copy number and allele frequency acquisition unit, a first segmentation analysis unit, a corrected heterozygous site screening unit, a second segmentation analysis unit, and an allele copy number calculation unit. The copy number and allele frequency acquisition unit is used to acquire the copy number signal and allele frequency signal of a single tumor sample based on a population single nucleotide polymorphism (SNP) site dataset. The first segmentation analysis unit is used to perform a first segmentation analysis using the copy number signal and the allele frequency signal to obtain... The first segment set; the aforementioned corrected heterozygous site screening unit is used to identify low heterozygous density regions in the aforementioned first segment set, and to screen out a corrected heterozygous site set from the aforementioned population single nucleotide polymorphism site dataset based on the distribution of allele frequencies of the aforementioned single tumor sample; the aforementioned second segment analysis unit is used to perform a second segment analysis using the aforementioned corrected heterozygous site set and the aforementioned copy number signal to obtain a second segment set; the aforementioned allele copy number calculation unit is used to calculate the purity and ploidy of the aforementioned single tumor sample based on the aforementioned second segment set, and to calculate the whole genome allele copy number in the aforementioned single tumor sample.
[0082] In a preferred embodiment, the copy number and allele frequency acquisition unit includes: a standard analysis site dataset construction module, a sequencing depth calculation module, a signal correction module, and an allele frequency calculation module; wherein, the standard analysis site dataset construction module is used to obtain single nucleotide polymorphism (SNP) sites from a population frequency database, extract the genotype information of each SNP site in the population, calculate the GC content of upstream and downstream windows of each SNP site, and obtain a standard analysis site dataset; and obtain sequencing data of normal samples, calculate the sequencing depth of the SNP sites in the normal samples, and record it as the normal sample sequencing depth, forming the population SNP site dataset; the sequencing depth calculation module is used to obtain sequencing data of a single tumor sample to be tested, calculate the sequencing depth of each SNP site in the sequencing data of the single tumor sample, and record it as the tumor single sample sequencing depth; calculate the number of supporting reads of each allele in the sequencing data of the single tumor sample; and calculate the log ratio signal LogR, where LogR = log2. R R = sequencing depth of the tumor single sample ÷ sequencing depth of the normal sample; the signal correction module is used to correct the LogR using the GC content in the standard analysis site dataset, and outputs the corrected log-ratio signal LogR after normalization. corr As the copy number signal mentioned above; the allele frequency calculation module is used to infer the genotype status of each SNP locus in the standard analysis locus dataset using the model, obtain a set of credible heterozygous loci, and calculate the allele frequency (BAF) of the set of credible heterozygous loci.
[0083] In a preferred embodiment, the model in the allele frequency calculation module includes a hidden Markov (HMM) model.
[0084] In a preferred embodiment, the first segmentation analysis unit includes a first segmentation algorithm execution module, which is used to execute a segmentation algorithm (ASPCF) to achieve the aforementioned LogR corr Using the allele frequencies of the aforementioned set of credible heterozygous sites as the synchronous analysis object, the genome of the aforementioned single tumor sample was subjected to the aforementioned first segmentation process to obtain the aforementioned first segment set composed of different genomic fragments.
[0085] In a preferred embodiment, the above-mentioned corrected heterozygous site screening unit includes: a SNP density calculation module, a subsignificant peak identification module, and a corrected candidate site screening module; wherein, the above-mentioned SNP density calculation module is used to calculate the number N of heterozygous SNP sites in each genomic fragment in the above-mentioned first segment set, and calculate SNP density = N ÷ length of the above-mentioned genomic fragment × 106 The aforementioned subsignificant peak identification module is used to calculate the mutation abundance (VAF) distribution of the aforementioned single nucleotide polymorphism sites when the aforementioned SNP density is less than 5, and to identify subsignificant peaks other than the main peaks at 0.5 or 1.0. The aforementioned correction candidate site screening module is used to select SNP sites located within the aforementioned subsignificant peak interval as correction candidate sites, form a set of correction heterozygous sites, and calculate the allele frequency of the aforementioned set of correction heterozygous sites.
[0086] In a preferred embodiment, the second segmentation analysis unit includes a second segmentation algorithm execution module, which is used to execute the segmentation algorithm (ASPCF) to achieve the aforementioned LogR corr Using the allele frequencies of the aforementioned set of corrected heterozygous sites as synchronous analysis objects, the genome of the aforementioned single tumor sample was subjected to the aforementioned second segmentation process to obtain the aforementioned second segment set composed of different genomic fragments.
[0087] In a preferred embodiment, the allele copy number calculation unit includes: a grid search module for performing a grid search on each genomic fragment in the second segment set for purity and ploidy to obtain a candidate solution list; an ABSOLUTE fitting module for inputting the candidate solution list and the corresponding genomic fragment in the second segment set into ABSOLUTE software to determine the purity and ploidy of the tumor single sample; and a copy number inference module for calculating the whole genome allele copy number of the tumor single sample based on the purity and ploidy.
[0088] In a preferred embodiment, the device is configured to cyclically execute the allele frequency calculation module, the first segmentation analysis unit, the corrected heterozygous site screening unit, and the second segmentation analysis unit to iteratively optimize the allele frequencies, thereby detecting copy number neutral loss of heterozygosity (cnLOH) or loss of heterozygosity (LOH) in high-purity samples.
[0089] In a third typical embodiment of this application, a computer device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method for calculating the whole genome copy number in a single tumor sample.
[0090] In a fourth typical embodiment of this application, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program / instructions are executed by a processor, the steps of the method for calculating the whole genome copy number in a single tumor sample are implemented.
[0091] In a fifth typical embodiment of this application, a computer program product is provided, including a computer program / instructions, which, when executed by a processor, implement the steps of the method for calculating the whole genome copy number in a single tumor sample.
[0092] In the sixth typical embodiment of this application, a method for calculating the whole genome copy number in a single tumor sample, or a device for calculating the whole genome copy number in a single tumor sample, or a computer device, or a computer-readable storage medium, or a computer program product is provided for application in any one or more of the following: pan-cancer HRD detection, tumor evolution analysis, or precision medicine assistance.
[0093] In the above applications, HRD detection can calculate the number of loss of heterozygosity (LOH), telomere allele imbalance (TAI), and large fragment migration (LST) based on the allele copy number results; tumor evolution analysis can use the copy number results to calculate the mutated cancer cell fraction (CCF); and copy number variation can serve as a marker for tumor molecular subtyping, with different subtypings having different prognoses, thus enabling its application in precision medicine.
[0094] The beneficial effects of this application will be explained in more detail below with reference to specific embodiments.
[0095] This invention has been validated using real sequencing data. Two normally paired samples were selected, and the allele copy number of each sample was analyzed using the technology of this invention. The results were then compared with those of known paired samples. The comparison results showed a high degree of consistency between the two methods in identifying loss of heterozygosity (LOH) regions, particularly in identifying copy number-neutral loss of heterozygosity (cnLOH).
[0096] Example 1
[0097] Allele copy number analysis of a single sample was performed using the technique described in this application. The sample used was named "Sample 1".
[0098] The technical process includes the following main steps:
[0099] 1) High-confidence single nucleotide polymorphism (SNP) sites were screened based on a population frequency database (1000 Genomes) to construct a standard set of analytical sites. For each site, population genotype information was extracted, and the GC content within a preset window range upstream and downstream of the site was calculated. This resulted in a standardized feature set containing site location, genotype distribution, and local sequence characteristics, which was used for a unified coordinate system and system bias modeling in subsequent analyses.
[0100] 2) Select multiple normal samples with good sequencing quality, calculate their sequencing depth at the standard site set, and summarize the depth information of multiple samples to construct a reference dataset (baseline) based on the median. This dataset will be used in subsequent steps to normalize the sequencing depth of tumor samples, thereby reducing the impact of differences in sequencing platforms, capture efficiency, and system noise.
[0101] 3) Obtain WES data of the tumor samples to be analyzed, extract sequencing depth and allele count information from the standard locus set, and obtain the total depth of each locus and the number of supporting reads for each allele. Based on this, construct a copy number signal index, that is, by calculating the logarithmic form of the depth ratio between the tumor sample and the reference sample set, obtain the logR value of each locus to reflect the relative copy number change of the genomic region.
[0102] 4) Considering that sequencing depth is affected by factors such as GC content, this invention performs systematic bias correction on the logR value. By establishing a functional relationship model between logR and GC content, the logR signal is corrected, and the corrected signal is normalized to eliminate differences in sample sequencing depth.
[0103] GC Correction for Regression Models: LogR corr = LogR - f(GC).
[0104] 5) Regarding BAF, the mutation frequency of SNP sites is calculated from the sequencing data. This invention introduces a genotype status inference method based on a probability model, using a Hidden Markov Model (HMM) to infer the genotype status of each site and identify a reliable set of heterozygous sites. For the heterozygous sites, their allele frequency (BAF) is calculated.
[0105] 6) The normalized logR signal and the BAF signal of heterozygous sites are used as joint inputs, and the genome is segmented using the ASPCF (Allele Frequency Fractional Array) algorithm. ASPCF simultaneously considers copy number signals and allele frequency signals for joint segmentation, dividing the whole genome into several continuous intervals, so that the signal in each interval presents a stable distribution, thereby corresponding to the potential copy number state and allele composition.
[0106] 7) To address the issue of analytical anomalies caused by the scarcity of heterozygous sites in certain genomic regions, this invention introduces a low-heterozygous-density region identification mechanism. Based on the first segmentation, the number of heterozygous SNP sites on each chromosome or within each segment interval is calculated. When the number of heterozygous sites per MB is less than 5, SNP sites within the second VAF peak range are selected based on the VAF distribution of germline mutations in the tumor sample (for samples with high tumor purity, the VAF of these SNP sites may be greater than 0.9).
[0107] 8) After optimizing the heterozygous site set, this invention re-executes the joint segmentation analysis based on logR and BAF to obtain more stable and accurate genome segmentation results. Through this iterative optimization process, cnLOH or LOH in high tumor purity samples can be effectively detected.
[0108] 9) This invention combines grid search for purity ploidy with ABSOLUTE software to cross-validate the purity ploidy model of the sample, select the most accurate model, and infer the allele copy number status of each interval, thereby achieving a systematic analysis of structural variations such as copy number amplification, deletion, and loss of heterozygosity in the tumor genome.
[0109] The results are highly consistent with those of paired samples, as shown below. Figure 2 The figure shows the single-sample allele copy number results corresponding to the technology in this application. Figure 2 In this context, A represents the major allele, B represents the minor allele, and cnLOH represents copy number-neutral loss of heterozygosity. Figure 3 This diagram shows the allele copy number results for paired samples corresponding to the technology described in this application. Figure 3 In this context, A represents the major allele, B represents the minor allele, and cnLOH represents copy number-neutral loss of heterozygosity.
[0110] The above Figure 2 Compared to the following comparative example Figure 4 It can accurately detect copy number-neutral heterozygous deletions on each chromosome (such as chromosomes 1-6 and 8-21), while the traditional method used in Comparative Example 1 would consider that these regions have no copy number variations.
[0111] Comparative Example 1
[0112] The allele copy number of the sample (the first sample in Example 1) was detected using conventional methods. The difference from Example 1 was that SNP sites within a specific VAF range were selected, whereas in Comparative Example 1, sites with VAF > 0.3 and VAF < 0.7 were used as heterozygous sites. The results are as follows: Figure 4 As shown, Figure 4 In this context, A represents the major allele, B represents the minor allele, and cnLOH represents copy number-neutral loss of heterozygosity.
[0113] in, Figure 4 The horizontal axis in the graph represents the genomic coordinates of each chromosome. Figure 4 The upper part of the image shows the logRatio of the SNP sites, where red indicates amplification, blue indicates deletion, green indicates cnLOH, and black indicates normal copy number. Figure 4The middle section shows the allele frequency (BAF) of the heterozygous locus, where A and B are the two genotypes at that locus, with A = 1 – B. Figure 4 The lower part of the image shows the allele copy number, with purple representing the total copy number and blue representing the copy number of the minor alleles.
[0114] Figure 5 This displays the density or average number of heterozygous loci on each chromosome. For example... Figure 5 As shown, most chromosomes (except chromosome 7) have very few heterozygous sites, less than 5 within a 1-megagen genome, and allele information is basically lost.
[0115] Figure 6 The data shows the mutation abundance (VAF) of SNP sites on each chromosome. Except for chromosome 7, the VAF of SNP sites on other chromosomes is above 0.8.
[0116] Figure 7 The results shown are the purity and ploidy cross-validation results. ABSOLUTE was used to perform cross-validation on tumor purity and ploidy to select the optimal model. Figure 7 The left image shows the process of determining purity and ploidy through grid search. Figure 7 The right figure shows the results of the purity-ploidy model fitted using ABSOLUTE for the copy number fragment loading distribution. Both show a purity of 0.76 and a ploidy of 2.09 that match the sample.
[0117] Example 2
[0118] Allele copy number analysis of a single sample (named "second sample") was performed using the technology in this application (same as in Example 1). The results were highly consistent with those of paired samples, as shown below.
[0119] in, Figure 8 The figure shows the single-sample allele copy number results corresponding to the technology in this application. Figure 8 In this context, A represents the major allele, B represents the minor allele, and cnLOH represents copy number-neutral loss of heterozygosity. Figure 9 This diagram shows the allele copy number results for paired samples corresponding to the technology described in this application. Figure 9 In this context, A represents the major allele, B represents the minor allele, and cnLOH represents copy number-neutral loss of heterozygosity.
[0120] The above Figure 8 Compared to the following comparative example 2 Figure 10 It can accurately detect copy number-neutral heterozygous deletions on each chromosome (such as chromosomes 1-6 and 8-21), while the traditional method used in Comparative Example 2 failed to detect them.
[0121] Comparative Example 2
[0122] The allele copy number of the sample (the second sample in Comparative Example 1) was detected using conventional methods. The difference from Example 2 was that SNP sites within a specific VAF range were selected, while Comparative Example 2 used sites with VAF > 0.3 and VAF < 0.7 as heterozygous sites. The results are as follows: Figure 10 As shown, Figure 10 In this context, A represents the major allele, B represents the minor allele, and cnLOH represents copy number-neutral loss of heterozygosity.
[0123] The allele copy number diagram for Comparative Example 2 is shown below. Figure 10 As shown. Figure 11 This displays the density or average number of heterozygous loci on each chromosome. For example... Figure 11 As shown, there are very few heterozygous sites on chromosomes 4, 5, 10, 13, and 22, less than 5 within a 1-megagen genome. Figure 12 The data shows the mutation abundance (VAF) of SNP sites on each chromosome. The VAF peaks of chromosomes 4, 5, 10, 13, and 22 deviate from 0.5. This is because these chromosomes have undergone copy number variations, which have led to changes in the VAF of heterozygous sites. It is necessary to reselect heterozygous sites based on the distribution of VAF.
[0124] Figure 13 The results of purity-ploidy cross-validation are shown in the left and right figures. ABSOLUTE was used to cross-validate tumor purity and ploidy to select the optimal model. The left figure illustrates the process of determining purity and ploidy through grid search, while the right figure shows the results of the purity-ploidy model fitted using ABSOLUTE for copy number fragment load distribution. Both figures show a purity of 0.87 and a ploidy of 1.75 that match this sample.
[0125] From the above description, it can be seen that the key technical points of this invention are mainly reflected in the following aspects: ① Most software uses relatively coarse selection of heterozygous sites, such as using fixed thresholds (VAF > 0.3 & VAF < 0.7). This invention uses a hidden Markov model with genotype status as the hidden variable to achieve robust identification of heterozygous sites across the entire genome. ② This invention employs a two-round segmentation strategy. In the case of abnormality detected in the first round, pseudo-heterozygous sites are selected for the second round of segmentation, which can accurately detect the copy number of alleles. ③ By analyzing the frequency of heterozygous sites on each chromosome or segmented region, abnormal chromosomes or regions are identified. Possible heterozygous sites are then reselected for these regions for the second round of segmentation analysis. ④ This invention simultaneously introduces grid search and the ABSOLUTE model, and cross-validates the purity ploidy model to obtain more accurate allele copy number genotyping results.
[0126] Compared with existing methods such as ASCAT and FACETS, this invention has significant advantages. First, this method can run without matching normal samples, greatly expanding its applicability to clinical and historical data. Second, this method supports genome-wide analysis, rather than being limited to target regions, providing more comprehensive copy number information. Furthermore, while maintaining high analytical accuracy, this method is compatible with different types of sequencing data (covering whole exomes and the entire genome), exhibiting good versatility and scalability. Simultaneously, this method can stably identify LOH and cnLOH events and supports HRD score calculation, providing important support for precision medicine.
[0127] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for calculating the whole genome copy number in a single tumor sample, characterized in that, The calculation method includes: S1) Obtain single nucleotide polymorphism sites from the population frequency database, extract the genotype information of each single nucleotide polymorphism site in the population, calculate the GC content of the upstream and downstream windows of each single nucleotide polymorphism site, and obtain a standard analysis site dataset. Obtain sequencing data from normal samples, calculate the sequencing depth of the single nucleotide polymorphism sites in the sequencing data, and denot it as the normal sample sequencing depth; The standard analysis site dataset and the sequencing depth of the normal samples constitute the population single nucleotide polymorphism site dataset; Based on the population single nucleotide polymorphism site dataset, the copy number signal and allele frequency of a single tumor sample were obtained; S2) Perform a first segmentation analysis using the copy number signal and the allele frequency to obtain a first segment set; Among them, a piecewise algorithm is used, and LogR is used. corr Using the allele frequencies of the set of credible heterozygous sites as the synchronous analysis object, the genome of the tumor single sample is first segmented to obtain a first segment set composed of different genomic fragments; S3) In the first segment set, identify low heterozygous density regions and, based on the distribution of allele frequencies of the tumor single sample, screen out a set of corrected heterozygous sites from the population single nucleotide polymorphism site dataset. S4) Perform a second segmentation analysis using the set of corrected heterozygous sites and the copy number signal to obtain a second segmentation set; Among them, a piecewise algorithm is used, and the corrected logarithmic ratio signal LogR is used. corr Using the allele frequencies of the corrected heterozygous site set as synchronous analysis objects, the genome of the tumor single sample is subjected to a second segmentation process to obtain a second segment set composed of different genomic fragments; S5) Calculate the purity and ploidy of the tumor single sample based on the second segment set, and calculate the whole genome copy number in the tumor single sample.
2. The calculation method according to claim 1, characterized in that, S1) includes: S1-1) Obtain the standard analysis site dataset and the sequencing depth of the normal sample to form the population single nucleotide polymorphism site dataset; S1-2) Obtain sequencing data of the single tumor sample; The sequencing depth of each single nucleotide polymorphism site in the sequencing data of the tumor single sample is calculated and denoted as the tumor single sample sequencing depth. Calculate the number of supporting reads for each allele in the sequencing data of the tumor single sample; Calculate the logarithmic ratio signal LogR, where LogR = log2. R R = sequencing depth of the tumor single sample ÷ sequencing depth of the normal sample; S1-3) Correcting the LogR, the correction including: correcting the LogR using the GC content in the standard analysis site dataset, and normalizing the corrected signal to obtain the corrected log-ratio signal LogR. corr This is the copy number signal; S1-4) The model is used to determine the genotype status of each single nucleotide polymorphism site in the standard analysis site dataset, to obtain the set of credible heterozygous sites, and to calculate the allele frequency of the set of credible heterozygous sites.
3. The calculation method according to claim 2, characterized in that, The models in S1-4 include Hidden Markov Models.
4. The calculation method according to claim 2 or 3, characterized in that, S3) includes: Calculate the number N of heterozygous single nucleotide polymorphism sites in the genome fragments in the first segment set; Calculate the SNP density for each of the genome segments, where SNP density = N ÷ length of the genome segment × 10. 6 ; If the SNP density is <5, calculate the mutation abundance distribution of the single nucleotide polymorphism site, select the secondary significant peaks other than the main peaks at 0.5 or 1.0, and select the SNP sites located in the interval of the secondary significant peaks as correction candidate sites to obtain the set of corrected heterozygous sites, and calculate the allele frequency of the set of corrected heterozygous sites.
5. The calculation method according to claim 1, characterized in that, S5) includes: For each genomic fragment in the second segment set, a grid search is performed for purity and ploidy to obtain a candidate solution list; The candidate solution list and the corresponding genomic fragments in the second segment set are input into ABSOLUTE software to determine the purity and ploidy of the tumor single sample; The whole genome copy number in the tumor single sample is calculated using the purity and ploidy of the tumor single sample.
6. The calculation method according to claim 2 or 3, characterized in that, The calculation method further includes: iteratively optimizing the allele frequencies by repeating steps S1-4)-S4) in the calculation method, thereby detecting copy number-neutral loss of heterozygosity or loss of heterozygosity in high tumor purity samples.
7. A device for calculating the whole genome copy number in a single tumor sample, characterized in that, The computing device includes: a copy number and allele frequency acquisition unit, a first segmentation analysis unit, a corrected heterozygous site screening unit, a second segmentation analysis unit, and an allele copy number calculation unit; The copy number and allele frequency acquisition unit is used to acquire copy number and allele frequency signals of a single tumor sample based on a population single nucleotide polymorphism site dataset. The copy number and allele frequency acquisition unit includes: a standard analysis site dataset construction module, a sequencing depth calculation module, a signal correction module, and an allele frequency calculation module. The standard analysis site dataset construction module is used to obtain single nucleotide polymorphism (SNP) sites from the population frequency database, extract the genotype information of each SNP site in the population, calculate the GC content of the upstream and downstream windows of each SNP site, and obtain the standard analysis site dataset; and obtain the sequencing data of normal samples, calculate the sequencing depth of the SNP sites in the normal samples, and record it as the normal sample sequencing depth to form the population SNP site dataset. The first segmentation analysis unit is used to perform a first segmentation analysis using the copy number signal and the allele frequency signal to obtain a first segment set; the first segmentation analysis unit includes a first segmentation algorithm execution module, which is used to execute a segmentation algorithm, using LogR corr Using the allele frequencies of the set of credible heterozygous sites as the synchronous analysis object, the genome of the tumor single sample is first segmented to obtain the first segment set composed of different genome fragments; The corrected heterozygous site screening unit is used to identify low heterozygous density regions in the first segment set and to screen a set of corrected heterozygous sites from the population single nucleotide polymorphism site dataset based on the distribution of allele frequencies of the tumor single sample. The second segmentation analysis unit is used to perform a second segmentation analysis using the corrected heterozygous site set and the copy number signal to obtain a second segment set; the second segmentation analysis unit includes a second segmentation algorithm execution module, which is used to execute the segmentation algorithm, using the LogR corr Using the allele frequencies of the corrected heterozygous site set as synchronous analysis objects, the genome of the tumor single sample is subjected to a second segmentation process to obtain a second segment set composed of different genomic fragments; The allele copy number calculation unit is used to calculate the purity and ploidy of the tumor single sample based on the second segment set, and to calculate the whole genome allele copy number in the tumor single sample.
8. The computing device according to claim 7, characterized in that, The sequencing depth calculation module is used to obtain sequencing data of a single tumor sample to be tested, calculate the sequencing depth of each single nucleotide polymorphism site in the sequencing data of the single tumor sample, and denot it as the single tumor sample sequencing depth; calculate the number of supporting reads for each allele in the sequencing data of the single tumor sample; and calculate the log ratio signal LogR, where LogR = log2. R R = sequencing depth of the tumor single sample ÷ sequencing depth of the normal sample; The signal correction module is used to correct the LogR using the GC content in the standard analysis site dataset, and outputs the corrected log-ratio signal LogR after normalization. corr As the copy number signal; The allele frequency calculation module is used to infer the genotype status of each single nucleotide polymorphism site in the standard analysis site dataset using a model, obtain a set of credible heterozygous sites, and calculate the allele frequency of the set of credible heterozygous sites.
9. The computing device according to claim 8, characterized in that, The model in the allele frequency calculation module includes a hidden Markov model.
10. The computing device according to claim 8 or 9, characterized in that, The corrected heterozygous site screening unit includes: an SNP density calculation module, a subsignificant peak identification module, and a corrected candidate site screening module; The SNP density calculation module is used to calculate the number N of heterozygous single nucleotide polymorphism sites in each genomic fragment in the first segment set, and to calculate the SNP density as N ÷ the length of the genomic fragment × 10. 6 ; The subsignificant peak identification module is used to calculate the mutation abundance distribution of the single nucleotide polymorphism site when the SNP density is less than 5, and to identify subsignificant peaks other than the main peak at 0.5 or 1.
0. The correction candidate site screening module is used to select single nucleotide polymorphism sites located within the secondary significant peak interval as correction candidate sites, form a set of correction heterozygous sites, and calculate the allele frequency of the set of correction heterozygous sites.
11. The computing device according to claim 7, characterized in that, The allele copy number calculation unit includes: The grid search module is used to perform a grid search on each genomic fragment in the second segment set for purity and ploidy to obtain a list of candidate solutions. The ABSOLUTE fitting module is used to input the candidate solution list and the corresponding genomic fragments in the second segment set into the ABSOLUTE software to determine the purity and ploidy of the tumor single sample; The copy number inference module is used to calculate the whole genome allele copy number of the tumor single sample based on the purity and ploidy.
12. The computing device according to claim 8 or 9, characterized in that, The computing device is configured to cyclically execute the allele frequency calculation module, the first segmentation analysis unit, the corrected heterozygous site screening unit, and the second segmentation analysis unit to iteratively optimize the allele frequency, thereby detecting copy number-neutral loss of heterozygosity or loss of heterozygosity in high tumor purity samples.
13. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the method for calculating the whole genome copy number in a single tumor sample according to any one of claims 1-6.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, characterized in that, when the computer program instructions are executed by a processor, they implement the steps of the method for calculating the whole genome copy number in a single tumor sample according to any one of claims 1-6.
15. A computer program product, characterized in that, The computer program product includes computer program instructions, characterized in that, when the computer program instructions are executed by a processor, they implement the steps of the method for calculating the whole genome copy number in a single tumor sample according to any one of claims 1-6.
16. The method for calculating the whole genome copy number in a single tumor sample according to any one of claims 1-6, or the apparatus for calculating the whole genome copy number in a single tumor sample according to any one of claims 7-12, or the computer apparatus according to claim 13, or the computer-readable storage medium according to claim 14, or the computer program product according to claim 15, for use in any one or more of the following: pan-cancer HRD detection or tumor evolution analysis.
Citation Information
Patent Citations
Single-sample whole-genome allele specific copy number variation prediction method
CN112802548A
Method and device for deducing tumor purity and ploidy
CN113990389A