Method for detecting gene copy number variation types and related products
By acquiring and analyzing sequencing alignment data, the corrected gene copy number of tumor cells is calculated, solving the problem that NGS detection methods cannot accurately identify homozygous deletions of genes, and realizing accurate detection of gene copy number variation types.
Patent Information
- Application Number
- CN202511142662.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing NGS detection methods cannot accurately assess the absolute copy number of genes in tumor cells and cannot identify homozygous deletions of genes, resulting in the inability to accurately identify homozygous deletions of tumor suppressor genes.
By acquiring sequencing alignment data between the test sample and the baseline sample, tumor purity and ploidy are determined. The corrected tumor cell copy number of the test gene in the test sample is calculated using a correction formula. Combined with the judgment thresholds for homozygous deletion and copy number amplification, the detection results are generated.
It enables accurate assessment of gene copy number in tumor cells, and can identify homozygous deletions and copy number amplifications of genes, thus improving the accuracy and reliability of detection.
Smart Images

Figure CN120748484B_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of bioinformatics analysis technology, specifically to methods for detecting gene copy number variation types and related products. Background Technology
[0002] A healthy human genome consists of 23 pairs of chromosomes, each pair containing two copies, one from the father and one from the mother, corresponding to a diploid genome. During the development and progression of tumors, genomic rearrangements can cause changes in gene copy numbers, such as copy number deletions or copy number amplifications.
[0003] Homozygous deletion of copy number is a type of copy number deletion where both copies of a gene are lost. For tumor suppressor genes, such as TP53, RB1, PTEN, and MTAP, the protein is normally expressed and suppresses the development of malignant tumors when both alleles are present. When a gene undergoes homozygous or heterozygous deletion, the gene is inactivated and no longer produces a suppressive effect, allowing cells to transform into cancer cells. Therefore, homozygous deletion of copy number in tumor suppressor genes is a crucial tumor marker, and its detection is essential.
[0004] Taking the tumor suppressor gene PTEN as an example, we can illustrate the significant importance of homozygous deletion of a gene. PTEN, short for human chromosome 10 deletion phosphatase, is a tumor suppressor gene located on the long arm of chromosome 10 (10q23.3). It consists of nine exons and is a phosphokinase with bispecific phosphatase activity. PTEN is located upstream of the PIK3-AKT-mTOR pathway and acts as a negative regulator of this pathway. It participates in cellular regulation through dephosphorylation, and its loss of function leads to pathway activation, driving tumor growth. Currently, various types of PTEN gene mutations have been identified, with PTEN deletion being the most common, prevalent in breast and prostate cancers. PTEN gene mutations or deletions often lead to the inactivation or non-expression of the PTEN protein, activating the PIK3-AKT-mTOR signaling pathway and thereby enhancing cell proliferation, survival, and migration. Patients with PTEN gene mutations or deletions typically have a poorer prognosis, and current evidence also shows that PTEN gene status is a potential predictive biomarker for treatment efficacy, highlighting the important clinical value of PTEN gene testing.
[0005] Currently, the detection methods for gene copy number variations (including deletions and amplifications) are mainly based on methodologies such as IHC (ImmunoHistoChemistry), FISH (Fluorescence In Situ Hybridization), and NGS (Next Generation Sequencing, also known as High-Throughput Sequencing, HTS).
[0006] IHC, as a routine clinical pathology test, is often used to detect protein expression. However, it suffers from difficulties in standardizing and unifying the procedure, and the results are easily affected by human and subjective factors, especially in the interpretation of the results. Furthermore, IHC can only perform qualitative detection of protein expression and cannot provide information on specific gene mutation types.
[0007] Compared to IHC, FISH offers the advantage of quantitative detection, but it also suffers from similar issues such as difficulty in standardizing procedures and susceptibility to human error. Furthermore, FISH is more complex to perform, requiring probe design and synthesis, and specialized equipment is needed for fluorescence signal detection. In addition, FISH can only detect known mutations and cannot detect unknown mutations. These limitations pose significant challenges to the clinical application of FISH.
[0008] NGS sequencing technology overcomes the shortcomings of IHC and FISH, and can accurately, quantitatively, and comprehensively detect all gene mutation information. At the same time, it greatly improves the detection throughput, and the operation method and data analysis are standardized and streamlined, avoiding interference from human factors.
[0009] Given the diversity and complexity of gene variations, NGS detection methods have significant advantages. However, existing NGS detection methods generally assess gene copy number variations based on sequencing depth. This method can only identify the mixed copy number of normal cells and tumor cells in tumor samples, and cannot accurately assess the absolute copy number of genes in tumor cells, thus failing to accurately identify homozygous deletions. Summary of the Invention
[0010] Embodiments of this disclosure provide methods, apparatus, electronic devices, storage media, and computer program products for detecting gene copy number variation types.
[0011] In a first aspect, embodiments of this disclosure provide a method for detecting gene copy number variation types, the method comprising:
[0012] Sequencing alignment data of gene sequencing of the sample to be tested and at least one baseline sample are obtained, wherein each sequencing alignment data includes sequencing alignment data of a preset set of gene analysis regions to be tested, and the preset set of gene analysis regions to be tested includes the region where the gene to be tested is located.
[0013] Based on the sequencing alignment data, the tumor purity and tumor ploidy of the sample to be tested are determined;
[0014] Based on the sequencing alignment data, the average mixed copy number of the gene to be tested in the sample to be tested is determined;
[0015] Based on the tumor purity and ploidy of the test sample and the average mixed copy number of the test gene in the test sample, the corrected tumor cell copy number of the test gene in the test sample is determined.
[0016] In response to the fact that the average mixed copy number of the gene to be tested in the test sample is less than 2, the test result of the test sample for the homozygous deletion of the gene to be tested is generated based on the comparison result of the corrected tumor cell copy number of the gene to be tested in the test sample and the positive judgment threshold of the homozygous deletion of the gene to be tested.
[0017] And / or,
[0018] In response to the average mixed copy number of the test gene in the test sample being greater than 2, the amplification judgment value of the test gene in the test sample is determined based on the corrected tumor cell copy number of the test gene in the test sample, and the detection result of the copy number amplification of the test gene in the test sample is generated based on the comparison result of the amplification judgment value of the test gene in the test sample and the amplification positive judgment threshold of the test gene.
[0019] In some optional implementations, determining the corrected tumor cell copy number of the test gene in the test sample based on the tumor purity and ploidy of the test sample and the average mixed copy number of the test gene in the test sample includes:
[0020] The tumor purity and ploidy of the test sample, and the average mixed copy number of the test gene in the test sample are substituted into the formula for calculating the corrected tumor cell copy number to obtain the corrected tumor cell copy number of the test gene.
[0021] In some optional embodiments, the formula for calculating the corrected tumor cell copy number is obtained by deriving a correction formula, wherein the correction formula is:
[0022]
[0023] The formula for calculating the corrected tumor cell copy number is as follows:
[0024]
[0025] in, and These represent the tumor purity and tumor ploidy of the sample to be tested, respectively. The average mixed copy number of the gene to be tested in the sample to be tested. The corrected tumor cell copy number of the gene to be tested in the sample to be tested.
[0026] In some optional embodiments, determining the amplification judgment value of the test gene in the test sample based on the corrected tumor cell copy number of the test gene in the test sample includes:
[0027] The corrected tumor cell copy number of the gene to be tested in the test sample is determined as the amplification judgment value of the gene to be tested in the test sample, or...
[0028] The difference between the corrected tumor cell copy number of the gene to be tested in the test sample and the tumor ploidy of the test sample is determined as the amplification judgment value of the gene to be tested in the test sample, or...
[0029] The amplification judgment value of the test gene in the test sample is determined by dividing the corrected tumor cell copy number of the test gene in the test sample by the tumor ploidy of the test sample.
[0030] In some optional implementations, determining the tumor purity and tumor ploidy of the sample to be tested based on the sequencing alignment data includes:
[0031] Based on the sequencing alignment data of the sample to be tested, the abundance of each SNP site in the preset whole-genome SNP site set is determined.
[0032] For each preset whole-genome SNP site, the relevant gene analysis region is divided into at least one SNP analysis region according to a first preset length.
[0033] For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region;
[0034] For each SNP analysis region, the mixed copy number of the sequencing alignment data of the sample to be tested in the SNP analysis region is determined based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the SNP analysis region and the mean of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region.
[0035] Based on the sequencing alignment data of the sample to be tested, the abundance and mixed copy number of each SNP site are used to determine the tumor purity and tumor ploidy of the sample to be tested. The mixed copy number of each SNP site is the mixed copy number of the SNP analysis region in which the SNP site is located.
[0036] In some optional implementations, determining the average mixed copy number of the gene to be tested in the sample based on each of the sequencing alignment data includes:
[0037] For each preset gene analysis region, the preset gene analysis region is divided into at least one gene analysis window according to the second preset length;
[0038] For each sequencing alignment data point, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window;
[0039] For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, the corresponding reference baseline sample sequencing alignment data of the gene analysis window, as well as the corrected average sequencing depth of the baseline sample and the deviation of the corrected average sequencing depth of the baseline sample in the gene analysis window are determined from the sequencing alignment data of each baseline sample.
[0040] Based on the corrected baseline sample average sequencing depth deviation of each of the gene analysis windows, a reference gene analysis window is determined in each of the gene analysis windows.
[0041] The analysis windows of each reference gene containing the gene to be tested are determined as the analysis windows of the gene to be tested.
[0042] For each gene analysis window, the mixed copy number of the gene analysis window in the test sample is determined based on the corrected average coverage depth of the gene analysis window and the corrected average sequencing depth of the baseline sample in the gene analysis window, according to the sequencing alignment data of the test sample.
[0043] The mean of the mixed copy number of each gene analysis window in the test sample is determined as the average mixed copy number of the gene in the test sample.
[0044] In some optional implementations, for each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determining the corresponding reference baseline sample sequencing alignment data for that gene analysis window, as well as the corrected average sequencing depth of the baseline samples and the deviation of the corrected average sequencing depth of the baseline samples in that gene analysis window, includes:
[0045] For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation:
[0046] Based on the sequencing alignment data of each baseline sample, the average coverage depth after correction within the gene analysis window is used to determine the normal range of the average sequencing depth of the baseline samples.
[0047] The sequencing alignment data of each baseline sample whose corrected average coverage depth in the gene analysis window is within the normal range of the average sequencing depth of the baseline sample are determined as the reference baseline sample sequencing alignment data corresponding to the gene analysis window.
[0048] Based on the sequencing alignment data of each reference baseline sample in the gene analysis window, and the corrected average coverage depth of the gene analysis window, the corrected average sequencing depth of the baseline samples and the deviation of the corrected average sequencing depth of the baseline samples in the gene analysis window are determined.
[0049] In some optional implementations, determining the corrected average coverage depth of each sequencing alignment data point for each gene analysis window includes:
[0050] For each sequencing alignment data point, using each gene analysis window as the target analysis region, the following target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each gene analysis window. The target analysis region sequencing depth correction operation includes:
[0051] Based on the total sequencing depth of the sequencing alignment data, the sequencing depth of each base site in each target analysis region is standardized to obtain the corresponding standardized sequencing depth.
[0052] For each target analysis region, the average sequencing depth of the sequencing alignment data in that target analysis region is determined based on the normalized sequencing depth of each base site in that target analysis region.
[0053] Based on the number of probes covered by each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region;
[0054] Based on the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.
[0055] In some optional implementations, determining the corrected average coverage depth of each sequencing alignment data point for each SNP analysis region includes:
[0056] For each sequencing alignment data, the SNP analysis region is used as the target analysis region, and the sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each SNP analysis region.
[0057] In some optional implementations, before determining the tumor purity and tumor ploidy of the sample to be tested based on the sequencing alignment data, the method further includes:
[0058] For each sequencing alignment data, the following base site sequencing depth normalization operation is performed: For each preset gene analysis region, the sequencing depth of each base site in the preset gene analysis region is determined; the sum of the sequencing depths of all base sites in each preset gene analysis region is determined as the total sequencing depth of the sequencing alignment data; for each base site in each preset gene analysis region, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.
[0059] In some optional implementations, the homozygous deletion positive threshold for the gene to be tested is predetermined through the following homozygous deletion positive threshold determination step:
[0060] A set of homozygous deletion sample data is obtained, wherein the homozygous deletion samples include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing alignment data of benign samples, and each homozygous deletion positive sample data is sequencing alignment data for homozygous deletion of the gene to be tested.
[0061] Determine the corrected tumor cell copy number for the gene to be tested for each homozygous deletion sample data;
[0062] For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, the following operation is performed to determine the positive criteria for homozygous deletion: based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the positive concordance rate corresponding to the candidate homozygous deletion copy number is determined; and / or, based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each homozygous deletion sample data, the positive predictive value corresponding to the candidate homozygous deletion copy number is determined;
[0063] The positive concordance rate and / or positive prediction value corresponding to each candidate homozygous deletion copy number are used to determine the positive judgment threshold for homozygous deletion of the gene to be tested.
[0064] In some optional implementations, obtaining the homozygous missing sample dataset includes:
[0065] Obtain sequencing alignment data for N pairs of 100% tumor cell lines and their paired cell lines, where N is a positive integer;
[0066] Based on the N pairs of sequencing alignment data, at least one homozygous deletion positive sample data for the gene to be tested is generated, wherein the tumor purity of each homozygous deletion positive sample data is the corresponding labeled tumor purity.
[0067] Obtain sequencing alignment data of M benign samples, and generate at least one negative sample data based on the sequencing alignment data of the M benign samples.
[0068] In some alternative implementations, determining the corrected tumor cell copy number for each homozygous deletion sample data for the gene to be tested includes:
[0069] For each homozygous deletion sample, determine the tumor purity and tumor ploidy corresponding to the homozygous deletion sample, as well as the average mixed copy number for the gene to be tested. Based on the tumor purity and tumor ploidy corresponding to the homozygous deletion sample, as well as the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the homozygous deletion sample.
[0070] In some optional implementations, determining the positive concordance rate corresponding to the candidate homozygous deletion copy number based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion positive sample data includes:
[0071] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number are counted.
[0072] The positive concordance rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number.
[0073] In some optional implementations, determining the positive predictive value corresponding to the candidate homozygous deletion copy number based on the comparison results of the corrected tumor cell copy number for the gene to be tested with the candidate homozygous deletion copy number in each of the homozygous deletion sample data includes:
[0074] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted.
[0075] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each negative sample data, the number of false positives corresponding to the candidate homozygous deletion copy number is counted; and
[0076] The positive predicted value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false positives corresponding to the candidate homozygous deletion copy number.
[0077] In some optional implementations, generating at least one negative sample data based on the sequencing alignment data of the M benign samples includes:
[0078] For each benign sample's sequencing alignment data, noise of varying degrees is added to the sequencing depth of the base sites in the benign sample's sequencing alignment data to obtain at least one negative sample data corresponding to the benign sample's sequencing alignment data.
[0079] In some optional implementations, the amplification positivity threshold for the gene to be tested is predetermined through the following amplification positivity threshold determination step:
[0080] A set of amplified sample data is obtained, which includes negative sample data and amplified positive sample data for the gene to be tested. Each negative sample data is obtained based on sequencing alignment data of benign samples, and each amplified positive sample data is sequencing alignment data of copy number amplification of the gene to be tested.
[0081] Determine the corrected tumor cell copy number for each amplified sample data;
[0082] The amplification judgment value for the gene to be tested is determined based on the corrected tumor cell copy number of each amplified sample data.
[0083] For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, the following amplification positive judgment basis determination operation is performed: based on the comparison result of the amplification judgment value of each amplification positive sample data for the gene to be tested with the candidate amplification positive judgment value, the positive concordance rate corresponding to the candidate amplification positive judgment value is determined; and / or, based on the comparison result of the amplification judgment value of each amplification sample data for the gene to be tested with the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined;
[0084] The amplification positivity threshold of the gene to be tested is determined based on the positive concordance rate and / or positive prediction value corresponding to each candidate amplification positivity judgment value.
[0085] In some optional implementations, obtaining the amplified sample data set includes:
[0086] Obtain sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines, where I is a positive integer;
[0087] Based on the I-pair sequencing alignment data, at least one amplified positive sample data for the gene to be tested is generated, wherein the tumor purity of each amplified positive sample data is the corresponding labeled tumor purity.
[0088] Obtain sequencing alignment data of J benign samples, and generate at least one negative sample data based on the sequencing alignment data of the J benign samples.
[0089] In some alternative implementations, determining the corrected tumor cell copy number for each amplified sample data includes:
[0090] For each amplified sample data, determine the tumor purity and tumor ploidy corresponding to the amplified sample data, as well as the average mixed copy number for the gene to be tested. Based on the tumor purity and tumor ploidy corresponding to the amplified sample data, as well as the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the amplified sample data.
[0091] In some optional embodiments, determining the positive concordance rate corresponding to the candidate amplification positive judgment value based on the comparison result between the amplification judgment value of each of the amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value includes:
[0092] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification positive judgment value of each amplification positive sample data, the number of true positives and false negatives corresponding to the candidate amplification positive judgment value are counted.
[0093] The positive concordance rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and false negatives corresponding to the candidate amplification positive judgment value.
[0094] In some optional embodiments, determining the positive prediction value corresponding to the candidate amplification positive judgment value based on the comparison result between the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value includes:
[0095] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification judgment value of each amplification positive sample data, the number of true positives corresponding to the candidate amplification positive judgment value is counted.
[0096] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification positive judgment value of each negative sample data, the number of false positives corresponding to the candidate amplification positive judgment value is counted; and
[0097] Based on the number of true positives and false positives corresponding to the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined.
[0098] In some optional implementations, determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested based on the corrected tumor cell copy number of each amplified sample data includes:
[0099] The corrected tumor cell copy number for the gene to be tested in each amplified sample is determined as the amplification judgment value for that amplified sample data, or...
[0100] The difference between the corrected tumor cell copy number for the target gene in each amplified sample and the tumor ploidy of that amplified sample is determined as the amplification judgment value for that amplified sample, or...
[0101] The amplification judgment value of each amplified sample data is determined by dividing the corrected tumor cell copy number for the gene to be tested by the tumor ploidy ratio of that amplified sample data.
[0102] Secondly, embodiments of this disclosure provide a gene copy number variation type detection device, the device comprising:
[0103] The sequencing alignment data acquisition module is configured to acquire sequencing alignment data from gene sequencing of the test sample and at least one baseline sample, respectively. Each sequencing alignment data includes sequencing alignment data of a preset set of gene analysis regions to be tested, and the preset set of gene analysis regions to be tested includes the region where the gene to be tested is located.
[0104] The purity and ploidy determination module is configured to determine the tumor purity and tumor ploidy of the test sample based on each of the sequencing alignment data.
[0105] The average mixed copy number determination module is configured to determine the average mixed copy number of the gene to be tested in the sample to be tested based on each of the sequencing alignment data.
[0106] The tumor cell copy number correction module is configured to determine the corrected tumor cell copy number of the test gene in the test sample based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the test gene in the test sample.
[0107] A homozygous deletion detection module is configured to, in response to a mean mixed copy number of the test gene in the test sample being less than 2, generate a detection result for homozygous deletion of the test gene in the test sample based on a comparison between the corrected tumor cell copy number of the test gene in the test sample and the maximum homozygous deletion positive copy number threshold of the test gene; and / or,
[0108] The copy number amplification detection module is configured to, in response to an average mixed copy number of the gene to be tested in the test sample being greater than 2, determine the amplification judgment value of the gene to be tested in the test sample based on the corrected tumor cell copy number of the gene to be tested in the test sample, and generate a detection result for the copy number amplification of the gene to be tested in the test sample based on a comparison result between the amplification judgment value of the gene to be tested in the test sample and the minimum amplification positive judgment value of the gene to be tested, the amplification positive judgment threshold.
[0109] In some optional implementations, the tumor cell copy number correction module is further configured to:
[0110] The tumor purity and ploidy of the test sample, and the average mixed copy number of the test gene in the test sample are substituted into the formula for calculating the corrected tumor cell copy number to obtain the corrected tumor cell copy number of the test gene.
[0111] In some optional embodiments, the formula for calculating the corrected tumor cell copy number is obtained by deriving a correction formula, wherein the correction formula is:
[0112]
[0113] The formula for calculating the corrected tumor cell copy number is as follows:
[0114]
[0115] in, and These represent the tumor purity and tumor ploidy of the sample to be tested, respectively. The average mixed copy number of the gene to be tested in the sample to be tested. The corrected tumor cell copy number of the gene to be tested in the sample to be tested.
[0116] In some optional embodiments, determining the amplification judgment value of the test gene in the test sample based on the corrected tumor cell copy number of the test gene in the test sample includes:
[0117] The corrected tumor cell copy number of the gene to be tested in the test sample is determined as the amplification judgment value of the gene to be tested in the test sample, or...
[0118] The difference between the corrected tumor cell copy number of the gene to be tested in the test sample and the tumor ploidy of the test sample is determined as the amplification judgment value of the gene to be tested in the test sample, or...
[0119] The amplification judgment value of the test gene in the test sample is determined by dividing the corrected tumor cell copy number of the test gene in the test sample by the tumor ploidy of the test sample.
[0120] In some optional implementations, the purity and ploidy determination module is further configured to:
[0121] Based on the sequencing alignment data of the sample to be tested, the abundance of each SNP site in the preset whole-genome SNP site set is determined.
[0122] For each preset whole-genome SNP site, the relevant gene analysis region is divided into at least one SNP analysis region according to a first preset length.
[0123] For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region;
[0124] For each SNP analysis region, the mixed copy number of the sequencing alignment data of the sample to be tested in the SNP analysis region is determined based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the SNP analysis region and the mean of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region.
[0125] Based on the sequencing alignment data of the sample to be tested, the abundance and mixed copy number of each SNP site are used to determine the tumor purity and tumor ploidy of the sample to be tested. The mixed copy number of each SNP site is the mixed copy number of the SNP analysis region in which the SNP site is located.
[0126] In some optional implementations, the average mixed copy number determination module is further configured to:
[0127] For each preset gene analysis region, the preset gene analysis region is divided into at least one gene analysis window according to the second preset length;
[0128] For each sequencing alignment data point, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window;
[0129] For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, the corresponding reference baseline sample sequencing alignment data of the gene analysis window, as well as the corrected average sequencing depth of the baseline sample and the deviation of the corrected average sequencing depth of the baseline sample in the gene analysis window are determined from the sequencing alignment data of each baseline sample.
[0130] Based on the corrected baseline sample average sequencing depth deviation of each of the gene analysis windows, a reference gene analysis window is determined in each of the gene analysis windows.
[0131] The analysis windows of each reference gene containing the gene to be tested are determined as the analysis windows of the gene to be tested.
[0132] For each gene analysis window, the mixed copy number of the gene analysis window in the test sample is determined based on the corrected average coverage depth of the gene analysis window and the corrected average sequencing depth of the baseline sample in the gene analysis window, according to the sequencing alignment data of the test sample.
[0133] The mean of the mixed copy number of each gene analysis window in the test sample is determined as the average mixed copy number of the gene in the test sample.
[0134] In some optional implementations, for each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determining the corresponding reference baseline sample sequencing alignment data for that gene analysis window, as well as the corrected average sequencing depth of the baseline samples and the deviation of the corrected average sequencing depth of the baseline samples in that gene analysis window, includes:
[0135] For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation:
[0136] Based on the sequencing alignment data of each baseline sample, the average coverage depth after correction within the gene analysis window is used to determine the normal range of the average sequencing depth of the baseline samples.
[0137] The sequencing alignment data of each baseline sample whose corrected average coverage depth in the gene analysis window is within the normal range of the average sequencing depth of the baseline sample are determined as the reference baseline sample sequencing alignment data corresponding to the gene analysis window.
[0138] Based on the sequencing alignment data of each reference baseline sample in the gene analysis window, and the corrected average coverage depth of the gene analysis window, the corrected average sequencing depth of the baseline samples and the deviation of the corrected average sequencing depth of the baseline samples in the gene analysis window are determined.
[0139] In some optional implementations, determining the corrected average coverage depth of each sequencing alignment data point for each gene analysis window includes:
[0140] For each sequencing alignment data point, using each gene analysis window as the target analysis region, the following target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each gene analysis window. The target analysis region sequencing depth correction operation includes:
[0141] Based on the total sequencing depth of the sequencing alignment data, the sequencing depth of each base site in each target analysis region is standardized to obtain the corresponding standardized sequencing depth.
[0142] For each target analysis region, the average sequencing depth of the sequencing alignment data in that target analysis region is determined based on the normalized sequencing depth of each base site in that target analysis region.
[0143] Based on the number of probes covered by each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region;
[0144] Based on the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.
[0145] In some optional implementations, determining the corrected average coverage depth of each sequencing alignment data point for each SNP analysis region includes:
[0146] For each sequencing alignment data, the SNP analysis region is used as the target analysis region, and the sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each SNP analysis region.
[0147] In some alternative implementations, the apparatus further includes a standardization module configured to determine the tumor purity and tumor ploidy of the test sample based on each of the sequencing alignment data:
[0148] For each sequencing alignment data, the following base site sequencing depth normalization operation is performed: For each preset gene analysis region, the sequencing depth of each base site in the preset gene analysis region is determined; the sum of the sequencing depths of all base sites in each preset gene analysis region is determined as the total sequencing depth of the sequencing alignment data; for each base site in each preset gene analysis region, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.
[0149] In some optional implementations, the homozygous deletion positive threshold for the gene to be tested is predetermined through the following homozygous deletion positive threshold determination step:
[0150] A set of homozygous deletion sample data is obtained, wherein the homozygous deletion samples include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing alignment data of benign samples, and each homozygous deletion positive sample data is sequencing alignment data for homozygous deletion of the gene to be tested.
[0151] Determine the corrected tumor cell copy number for the gene to be tested for each homozygous deletion sample data;
[0152] For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, the following operation is performed to determine the positive criteria for homozygous deletion: based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the positive concordance rate corresponding to the candidate homozygous deletion copy number is determined; and / or, based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each homozygous deletion sample data, the positive predictive value corresponding to the candidate homozygous deletion copy number is determined;
[0153] The positive concordance rate and / or positive prediction value corresponding to each candidate homozygous deletion copy number are used to determine the positive judgment threshold for homozygous deletion of the gene to be tested.
[0154] In some optional implementations, obtaining the homozygous missing sample dataset includes:
[0155] Obtain sequencing alignment data for N pairs of 100% tumor cell lines and their paired cell lines, where N is a positive integer;
[0156] Based on the N pairs of sequencing alignment data, at least one homozygous deletion positive sample data for the gene to be tested is generated, wherein the tumor purity of each homozygous deletion positive sample data is the corresponding labeled tumor purity.
[0157] Obtain sequencing alignment data of M benign samples, and generate at least one negative sample data based on the sequencing alignment data of the M benign samples.
[0158] In some alternative implementations, determining the corrected tumor cell copy number for each homozygous deletion sample data for the gene to be tested includes:
[0159] For each homozygous deletion sample, determine the tumor purity and tumor ploidy corresponding to the homozygous deletion sample, as well as the average mixed copy number for the gene to be tested. Based on the tumor purity and tumor ploidy corresponding to the homozygous deletion sample, as well as the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the homozygous deletion sample.
[0160] In some optional implementations, determining the positive concordance rate corresponding to the candidate homozygous deletion copy number based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion positive sample data includes:
[0161] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number are counted.
[0162] The positive concordance rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number.
[0163] In some optional implementations, determining the positive predictive value corresponding to the candidate homozygous deletion copy number based on the comparison results of the corrected tumor cell copy number for the gene to be tested with the candidate homozygous deletion copy number in each of the homozygous deletion sample data includes:
[0164] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted.
[0165] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each negative sample data, the number of false positives corresponding to the candidate homozygous deletion copy number is counted; and
[0166] The positive predicted value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false positives corresponding to the candidate homozygous deletion copy number.
[0167] In some optional implementations, generating at least one negative sample data based on the sequencing alignment data of the M benign samples includes:
[0168] For each benign sample's sequencing alignment data, noise of varying degrees is added to the sequencing depth of the base sites in the benign sample's sequencing alignment data to obtain at least one negative sample data corresponding to the benign sample's sequencing alignment data.
[0169] In some optional implementations, the amplification positivity threshold for the gene to be tested is predetermined through the following amplification positivity threshold determination step:
[0170] A set of amplified sample data is obtained, which includes negative sample data and amplified positive sample data for the gene to be tested. Each negative sample data is obtained based on sequencing alignment data of benign samples, and each amplified positive sample data is sequencing alignment data of copy number amplification of the gene to be tested.
[0171] Determine the corrected tumor cell copy number for each amplified sample data;
[0172] The amplification judgment value for the gene to be tested is determined based on the corrected tumor cell copy number of each amplified sample data.
[0173] For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, the following amplification positive judgment basis determination operation is performed: based on the comparison result of the amplification judgment value of each amplification positive sample data for the gene to be tested with the candidate amplification positive judgment value, the positive concordance rate corresponding to the candidate amplification positive judgment value is determined; and / or, based on the comparison result of the amplification judgment value of each amplification sample data for the gene to be tested with the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined;
[0174] The amplification positivity threshold of the gene to be tested is determined based on the positive concordance rate and / or positive prediction value corresponding to each candidate amplification positivity judgment value.
[0175] In some optional implementations, obtaining the amplified sample data set includes:
[0176] Obtain sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines, where I is a positive integer;
[0177] Based on the I-pair sequencing alignment data, at least one amplified positive sample data for the gene to be tested is generated, wherein the tumor purity of each amplified positive sample data is the corresponding labeled tumor purity.
[0178] Obtain sequencing alignment data of J benign samples, and generate at least one negative sample data based on the sequencing alignment data of the J benign samples.
[0179] In some alternative implementations, determining the corrected tumor cell copy number for each amplified sample data includes:
[0180] For each amplified sample data, determine the tumor purity and tumor ploidy corresponding to the amplified sample data, as well as the average mixed copy number for the gene to be tested. Based on the tumor purity and tumor ploidy corresponding to the amplified sample data, as well as the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the amplified sample data.
[0181] In some optional embodiments, determining the positive concordance rate corresponding to the candidate amplification positive judgment value based on the comparison result between the amplification judgment value of each of the amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value includes:
[0182] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification positive judgment value of each amplification positive sample data, the number of true positives and false negatives corresponding to the candidate amplification positive judgment value are counted.
[0183] The positive concordance rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and false negatives corresponding to the candidate amplification positive judgment value.
[0184] In some optional embodiments, determining the positive prediction value corresponding to the candidate amplification positive judgment value based on the comparison result between the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value includes:
[0185] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification judgment value of each amplification positive sample data, the number of true positives corresponding to the candidate amplification positive judgment value is counted.
[0186] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification positive judgment value of each negative sample data, the number of false positives corresponding to the candidate amplification positive judgment value is counted; and
[0187] Based on the number of true positives and false positives corresponding to the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined.
[0188] In some optional implementations, determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested based on the corrected tumor cell copy number of each amplified sample data includes:
[0189] The corrected tumor cell copy number for the gene to be tested in each amplified sample is determined as the amplification judgment value for that amplified sample data, or...
[0190] The difference between the corrected tumor cell copy number for the target gene in each amplified sample and the tumor ploidy of that amplified sample is determined as the amplification judgment value for that amplified sample, or...
[0191] The amplification judgment value of each amplified sample data is determined by dividing the corrected tumor cell copy number for the gene to be tested by the tumor ploidy ratio of that amplified sample data.
[0192] Thirdly, embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation of the first aspect.
[0193] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any implementation of the first aspect.
[0194] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the method described in any of the implementations of the first aspect.
[0195] To address the shortcomings of existing NGS-based gene copy number variation (GNV) detection methods, which cannot accurately identify the absolute copy number of genes in tumor cells and cannot distinguish between homozygous and heterozygous deletions, embodiments of this disclosure provide a GNV detection method, apparatus, electronic device, storage medium, and computer program product. This method first acquires sequencing alignment data from gene sequencing of the test sample and at least one baseline sample. Each sequencing alignment data includes sequencing alignment data from a preset set of regions to be analyzed for the target genes, which includes the regions where the target genes are located. Then, based on the sequencing alignment data, the tumor purity and ploidy of the test sample, as well as the average mixed copy number of the target genes in the test sample, are determined. Finally, based on the tumor purity and ploidy of the test sample, the tumor purity and ploidy of the tumor sample are determined. The tumor ploidy and the average mixed copy number of the test gene in the test sample are used to determine the corrected tumor cell copy number of the test gene in the test sample. Next, if the average mixed copy number of the test gene in the test sample is less than 2, the test result for homozygous deletion of the test gene in the test sample is determined by comparing the corrected tumor cell copy number of the test gene in the test sample with the positive judgment threshold for homozygous deletion of the test gene. And / or, if the average mixed copy number of the test gene in the test sample is greater than 2, the amplification judgment value of the test gene in the test sample is determined by the corrected tumor cell copy number of the test gene in the test sample, and the test result for copy number amplification of the test gene in the test sample is determined by comparing the amplification judgment value of the test gene in the test sample with the positive judgment threshold for amplification. Therefore, by taking into account the tumor ploidy of the test sample in the process of determining the corrected copy number of the test gene in the test sample, the accuracy of the corrected copy number of the test gene in the test sample (i.e., the absolute copy number of the test gene in the tumor cells) can be improved due to the richer range of factors considered. This, in turn, can improve the accuracy of detecting copy number variation states of the test gene. Attached Figure Description
[0196] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the scope of this disclosure. In the drawings:
[0197] Figure 1 This is an exemplary system architecture diagram to which one embodiment of this disclosure can be applied;
[0198] Figure 2A This is a flowchart of an embodiment of the gene copy number variation type detection method according to the present disclosure;
[0199] Figure 2B This is an exploded flowchart of one embodiment of step 202 of the present disclosure;
[0200] Figure 2C This is an exploded flowchart of one embodiment of step 203 of this disclosure;
[0201] Figure 2D This is a breakdown flowchart of one embodiment of the baseline sample gene analysis window depth correction operation according to this disclosure;
[0202] Figure 3A This is a flowchart of one embodiment of the steps for determining the homozygous deletion positive threshold according to this disclosure;
[0203] Figure 3B This is an exploded flowchart of an embodiment of step 301 of the present disclosure;
[0204] Figure 3C This is an exploded flowchart of one embodiment of step 3012 of this disclosure;
[0205] Figure 3D This is a flowchart of one embodiment of the operation determined based on the homozygous deletion positive criterion of this disclosure;
[0206] Figure 4A This is a flowchart of one embodiment of the amplification positive judgment threshold determination step according to the present disclosure;
[0207] Figure 4B This is an exploded flowchart of an embodiment of step 401 of the present disclosure;
[0208] Figure 4C This is a breakdown flowchart of one embodiment of the amplified positive sample data generation operation according to this disclosure;
[0209] Figure 4D This is a flowchart of one embodiment of the operation based on the amplification positive judgment criteria of this disclosure;
[0210] Figure 5 This is a Probit curve plot of the gene PTEN under two homozygous deletion variants: 1-2 exon deletion and 3-9 exon deletion.
[0211] Figure 6A , Figure 6B and Figure 6C These are the Probit curves for the genes to be tested, MTAP, CDKN2A, and CDKN2B, respectively.
[0212] Figure 7A , Figure 7B and Figure 7C These are Probit curves for diploid CCNE1 amplification, triploid CCNE1 amplification, and tetraploid CCNE1 amplification, respectively.
[0213] Figure 8 A schematic diagram of the structure of an embodiment of the gene copy number variation type detection device disclosed herein;
[0214] Figure 9 A schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Detailed Implementation
[0215] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0216] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0217] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the gene copy number variation type detection methods, apparatuses, electronic devices and storage media of the present disclosure can be applied.
[0218] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0219] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as gene copy number variation detection applications, high-throughput sequencing applications, and bioinformatics applications.
[0220] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with information input devices (e.g., keyboard, mouse, touchscreen, microphone, camera, etc.) and information output devices (e.g., display screen, speakers, etc.), including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed on the terminal devices listed above. They can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.
[0221] In some cases, the gene copy number variation type detection method provided in this disclosure can be executed by terminal devices 101, 102, and 103, and correspondingly, the gene copy number variation type detection device can be set in terminal devices 101, 102, and 103. In this case, the system architecture 100 may not include server 105.
[0222] In some cases, the gene copy number variation type detection method provided in this disclosure can be jointly executed by terminal devices 101, 102, 103 and server 105. For example, the step of "acquiring sequencing alignment data of the sample to be tested and at least one baseline sample" can be executed by terminal devices 101, 102, and 103, and the step of "determining the tumor purity and tumor ploidy of the sample to be tested based on each sequencing alignment data" can be executed by server 105. This disclosure does not limit this. Correspondingly, the gene copy number variation type detection device can also be respectively set in terminal devices 101, 102, 103 and server 105.
[0223] In some cases, the gene copy number variation type detection method provided in this disclosure can be executed by server 105, and correspondingly, the gene copy number variation type detection device can also be set in server 105. In this case, the system architecture 100 may not include terminal devices 101, 102, and 103.
[0224] It should be noted that server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules (for example, used to provide distributed services), or as a single software program or software module. No specific limitations are made here.
[0225] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0226] Continue to refer to Figure 2A The diagram illustrates a flow 200 of an embodiment of a gene copy number variation type detection method according to the present disclosure, which includes the following steps:
[0227] Step 201: Obtain sequencing alignment data of the test sample and at least one baseline sample, respectively.
[0228] Here, the sample to be tested can be, for example, tumor tissue (such as lung cancer tissue or breast cancer tissue) obtained through biopsy or surgical resection, used to detect whether copy number variations have occurred in the sample.
[0229] The baseline sample can be normal tissue cells from the same patient as the sample to be tested, or it can be normal tissue (or benign tissue) cells from a group of other samples. It is usually taken from the patient's peripheral blood leukocytes (germline DNA) or adjacent normal tissue (collected at the same time during surgery).
[0230] First, high-throughput sequencing can be performed on the sample to obtain the raw sequencing data. Then, the raw sequencing data is aligned (e.g., using tools like BWA or Bowtie2) to a human reference genome (e.g., hg38 / GRCh38) to obtain the sequencing alignment data. Here, the sequencing alignment data can be, for example, a BAM (Binary Alignment Map) file.
[0231] Alternatively, high-throughput sequencing can be performed on each baseline sample to obtain the raw sequencing data of the corresponding baseline sample. Then, the raw sequencing data of each baseline sample can be back-aligned (e.g., using tools such as BWA or Bowtie2) to the human reference genome (e.g., hg38 / GRCh38) to obtain the sequencing alignment data of the corresponding baseline sample.
[0232] As an example, suppose high-throughput sequencing is performed on each of the B baseline samples. Then, after step 201, the sequencing alignment data of one sample to be tested and the sequencing alignment data of each of the B baseline samples can be obtained, that is, a total of (B+1) sequencing alignment data are obtained. For example, a total of (B+1) BAM files are obtained.
[0233] It should be noted that the sequencing alignment data mentioned above (B+1) all include sequencing alignment data of the preset set of regions to be analyzed for the gene to be tested, which includes the region where the gene to be tested is located.
[0234] In the sequencing process described above, various library preparation and sequencing methods can be used. For example, hybridization capture, amplicon sequencing, or whole-genome sequencing can be employed, as long as the library preparation and sequencing methods can sequence each preset gene region and obtain sequencing alignment data for each preset gene region.
[0235] It should be noted that the sequencing methods for the test sample and each baseline sample can be the same, including both library preparation and sequencing methods. For example, if hybridization capture sequencing is used, the capture probe settings can be identical.
[0236] Step 202: Based on the sequencing alignment data, determine the tumor purity and tumor ploidy of the sample to be tested.
[0237] Here, the tumor purity and tumor ploidy of the test sample can be determined based on the sequencing alignment data of one test sample and the sequencing alignment data of each baseline sample obtained in step 201, such as the sequencing alignment data of one test sample and the sequencing alignment data of B baseline samples.
[0238] Here, the tumor purity of the test sample is the proportion of tumor cells in the test sample, and its value can be between 0 and 100%.
[0239] The tumor ploidy of the sample to be tested is the average copy number of the tumor cell genome (usually 2 (diploid), 3, 4, etc.).
[0240] In some alternative implementations, step 202 may include, for example: Figure 2B The following steps 2021 to 2025 are shown:
[0241] Step 2021: Based on the sequencing alignment data of the sample to be tested, determine the abundance of each SNP site in the preset whole-genome SNP site set of the sample to be tested.
[0242] Specifically, step 2021 can be performed as follows:
[0243] First, variant calling involves detecting SNP sites in the sequencing alignment data of the sample and generating a VCF file.
[0244] For example, you can use GATK (the gold standard), BCFtools (lightweight), or FreeBayes to generate VCF files.
[0245] Then, SNP abundance (Allele Frequency) is calculated, which is the allele frequency (AF) of each SNP locus extracted from the VCF file.
[0246] For example, it can be achieved using BCFtools, GATK VariantsToTable, or through a custom script (using scripting languages such as Python / R).
[0247] Here, the preset set of whole-genome SNP sites can be SNP sites that are pre-set according to actual detection needs.
[0248] Step 2022: For each preset whole-genome SNP site, the relevant gene analysis region is divided into at least one SNP analysis region according to the first preset length.
[0249] Here, various methods can be used to first determine the relevant gene analysis region for each genome-wide SNP locus. As an example, the genome-wide SNP locus can be used as a reference location, and the relevant gene analysis region for the genome-wide SNP locus can be determined according to a pre-defined gene analysis region determination method. For instance, the region between the first base site located three pre-defined lengths forward from the reference location and the second base site located four pre-defined lengths backward from the reference location can be considered as the relevant gene analysis region for the genome-wide SNP locus.
[0250] Then, the relevant gene analysis region of the whole genome SNP site is divided into at least one SNP analysis region according to the first preset length.
[0251] Here, the first preset length can be pre-set according to the actual detection needs. For example, when sequencing the test sample and the baseline sample using hybridization capture, the first preset length can be determined based on the probe length during the hybridization capture sequencing process. Generally, the probe length of each capture region is the same, such as 120 BP, and the length of each capture probe can be used as the first preset length. That is, as an example, the first preset length can be 120 BP.
[0252] Step 2023: For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region.
[0253] Here, for each sequencing alignment data point obtained in step 201 from the sequencing alignment data of one sample to be tested and the sequencing alignment data of B baseline samples, the corrected average coverage depth of the sequencing alignment data for each SNP analysis region can be determined. Specifically, the average coverage depth of each SNP analysis region in the sequencing alignment data can be determined first, and then various coverage depth correction methods can be used to correct the average coverage depth of each SNP analysis region.
[0254] Alternatively, step 2023 can be performed as follows:
[0255] For each sequencing alignment dataset, each SNP analysis region is used as the target analysis region. Sequencing depth correction is performed on the target analysis region to obtain the corrected average coverage depth of each SNP analysis region. The target analysis region sequencing depth correction operation can be performed as follows:
[0256] First, the sequencing depth of each base site in each target analysis region is standardized based on the total sequencing depth of the sequencing alignment data to obtain the corresponding standardized sequencing depth.
[0257] For example, the normalized sequencing depth of a specific base site within a target analysis region can be calculated as: the sequencing depth of that base site divided by the total sequencing depth of the sequencing alignment data, multiplied by a global scaling factor. Here, the global scaling factor is typically the median or mean of the total depths of all samples (e.g., 100×) to make the normalized sequencing depth more intuitive. The global scaling factor is related to the size of the capture region. Specifically, it is generally based on the size of the panel (gene pack / detection package) used for sequencing; for example, a large panel uses 10 to the power of 7, and a small panel uses 10 to the power of 6, with the aim of keeping the normalized sequencing depth within 100.
[0258] Standardized sequencing depth can eliminate differences in sequencing data volume between different samples.
[0259] Then, for each target analysis region, the average sequencing depth of the sequencing alignment data in that target analysis region is determined based on the normalized sequencing depth of each base site in that target analysis region.
[0260] That is, the average sequencing depth of the sequencing alignment data in the target analysis region can be obtained by dividing the sum of the standardized sequencing depths of all base sites in the target analysis region by the number of base sites in the target analysis region.
[0261] Next, based on the number of probes covered by each target analysis region, the average coverage depth of the sequencing alignment data in each target analysis region was locally weighted by regression correction.
[0262] Here, in order to eliminate the difference in coverage depth caused by the difference in the number of probes covered by different target analysis regions, the average coverage depth of the sequencing alignment data in each target analysis region can be locally weighted by regression correction based on the number of probes covered by each target analysis region.
[0263] Finally, based on the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.
[0264] Similarly, in order to eliminate the difference in coverage depth caused by the difference in GC content ratio in different target analysis regions, the average coverage depth of the sequencing alignment data in each target analysis region can be locally weighted by regression correction based on the GC content ratio of the sequencing alignment data in each target analysis region, so as to obtain the corrected average coverage depth of the corresponding target analysis region.
[0265] After step 2023, the corrected average coverage depth of each sequencing data for each SNP analysis region can be obtained from (B+1) sequencing alignment data of 1 test sample and B baseline samples.
[0266] Step 2024: For each SNP analysis region, determine the mixed copy number of the sequencing alignment data of the sample to be tested in the SNP analysis region based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the SNP analysis region and the mean of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region.
[0267] In other words, for each SNP analysis region, firstly, the corrected average coverage depth of the test sample obtained in step 2023 in that SNP analysis region can be obtained. Then, the mean corrected coverage depth of the sequencing alignment data of the B baseline samples obtained in step 2023 in the SNP analysis region is calculated. Finally, based on and Determine the mixed copy number of the sequencing alignment data of the sample to be tested in the SNP analysis region. Specifically, the corrected average coverage depth of the sample in the SNP analysis region can be used. Divide by the mean of the corrected average coverage depth of the SNP analysis region based on the sequencing alignment data of each baseline sample. The product of the ratio and 2 is used to determine the mixed copy number of the sequencing alignment data of the sample to be tested in that SNP analysis region. That is, the following formula can be used for calculation:
[0268] If there is Each SNP analysis region can be obtained through step 2024. The number of mixed copies of each SNP analysis region in the SNP analysis region can be obtained. Number of mixed copies .
[0269] Step 2025: Based on the abundance and mixed copy number of each SNP site in the sequencing alignment data of the sample to be tested, determine the tumor purity and tumor ploidy of the sample to be tested.
[0270] Here, the mixed copy number of each SNP site is the mixed copy number of the SNP analysis region where that SNP site is located.
[0271] If the sequencing alignment data of the sample to be tested contains There are several SNP sites, and various implementation methods can be used here, based on... The abundance and mixed copy number of each SNP site in the sample were used to determine the tumor purity and ploidy of the sample.
[0272] As an example, this can be achieved using existing mature bioinformatics analysis tools or equivalent techniques. For instance, using bioinformatics analysis tools with similar functions (ABSOLUTE, Sequenza, ASCAT, FACETS, etc.), and adapting and implementing them based on their principles and processes to meet the specific needs of this disclosure, it is possible to determine the tumor purity of the sample by analyzing the abundance and mixed copy number of each SNP site based on the sequencing alignment data of the sample. and tumor ploidy Understandably, other methods can also be used to determine the tumor purity of the test sample. and tumor ploidy This is not the focus of this invention and will not be elaborated upon here.
[0273] The optional approach in steps 2021 to 2025 above involves dividing the relevant gene analysis region of each preset whole-genome SNP locus into at least one SNP analysis region, and then correcting the average coverage depth of each SNP analysis region. This can improve the accuracy of the average coverage depth of each SNP analysis region. Furthermore, by obtaining the mixed copy number of each SNP analysis region based on the corrected average coverage depth of the test sample and the baseline sample, and then using the mixed copy number of each SNP analysis region as the mixed copy number of all SNP loci within that region, the computational workload of calculating the mixed copy number of all SNP loci can be significantly reduced. Finally, based on the abundance and mixed copy number of each SNP locus, the tumor purity of the test sample is obtained. and tumor ploidy .
[0274] Step 203: Based on the sequencing alignment data, determine the average mixed copy number of the gene to be tested in the sample to be tested.
[0275] If the gene sequencing process involving the test sample and at least one baseline sample is targeted sequencing (such as panel sequencing or exon sequencing), the test gene can be a gene selected during the experimental design (e.g., cancer-related genes such as PTEN, TP53, ERBB2, EGFR, BRCA1 / 2). In targeted sequencing, the test gene is determined by the coverage of the probe or capture kit (such as Illumina TruSeq or Agilent SureSelect). For example, tumor panels (such as MSK-IMPACT) contain hundreds of cancer driver genes, while clinical diagnostic panels (such as FoundationOne CDx) cover key genes specific to certain cancer types.
[0276] If the gene sequencing process for the test sample and at least one baseline sample is whole-genome sequencing (WGS), theoretically, any gene can be analyzed, but usually priority is given to genes known to be associated with diseases (such as tumors).
[0277] The mean mixed copy number of the gene to be tested in the test sample can be considered as the weighted average of the copy number of the gene to be tested in all cells (tumor cells + normal cells) in the test sample (such as tumor tissue). Since tumor samples usually contain tumor cells and mixed normal cells (such as stromal cells and immune cells), the mean mixed copy number is the result of the mixture of copy numbers of the two cell populations.
[0278] In some alternative implementations, step 203 may include, for example: Figure 2C The following steps 2031 to 2037 are shown:
[0279] Step 2031: For each preset gene analysis region, divide the preset gene analysis region into at least one gene analysis window according to the second preset length.
[0280] Here, the second preset length can be set according to the actual detection needs. For example, when the sequencing of the test sample and each baseline sample is performed using the hybridization capture method, the second preset length can be determined based on the average probe deployment multiplier and the average probe length of each capture region.
[0281] Here, executing step 2031 can achieve two technical effects:
[0282] First, generally speaking, the length of the pre-defined gene analysis region is 120 BP. If the pre-defined gene analysis region is not divided, each pre-defined gene analysis region corresponds to only one data point as an analysis interval, making it impossible to perform the statistical tests in subsequent steps 2036. Therefore, it is necessary to cut it into smaller regions for subsequent statistical tests.
[0283] Second, in the subsequent process of correcting the average coverage depth using the number of probes covered, the accuracy of correcting the average coverage depth can be improved by dividing the preset gene analysis region to be tested into gene analysis regions with smaller granularity.
[0284] As an example, assuming the average probe deployment multiplier for each capture area is 5 and the average probe length is 120BP, then the second preset length can be the ratio of the average probe length of 120BP to the average probe deployment multiplier of 5, that is, the probe step size of 24BP.
[0285] Step 2032: For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window.
[0286] Here, for each sequencing alignment data point obtained in step 201 from the sequencing alignment data of one test sample and the sequencing alignment data of B baseline samples, the corrected average coverage depth of the sequencing alignment data for each gene analysis window can be determined. Specifically, the average coverage depth of each gene analysis window in the sequencing alignment data can be determined first, and then various coverage depth correction methods can be used to correct the average coverage depth of each gene analysis window.
[0287] Optionally, step 2032 can be performed as follows: for each sequencing alignment data, take each gene analysis window as the target analysis region, perform the sequencing depth correction operation of the target analysis region, and obtain the corrected average coverage depth of each gene analysis window.
[0288] Specifically, the execution process of the sequencing depth correction operation for the target analysis region can be found in the previous description of step 2023, and will not be repeated here.
[0289] After step 2032, the corrected average coverage depth of each sequencing data for each gene analysis window can be obtained from (B+1) sequencing alignment data of 1 test sample and B baseline samples.
[0290] Step 2033: For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, determine the corresponding reference baseline sample sequencing alignment data, the corrected average sequencing depth of the baseline sample and the deviation of the corrected average sequencing depth of the baseline sample in the sequencing alignment data of each baseline sample.
[0291] Step 2032 has yielded the average coverage depth of each baseline sample after correction for sequencing alignment data in each gene analysis window.
[0292] Optionally, step 2033 can be performed as follows: For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation, which includes, for example: Figure 2D The following steps 20331 to 20333 are shown:
[0293] Step 20331: Based on the sequencing alignment data of each baseline sample and the corrected average coverage depth within the gene analysis window, determine the normal range of the average sequencing depth of the baseline samples.
[0294] In practice, factors such as low capture efficiency, high GC content, and repetitive sequences can cause a sudden drop / increase in the coverage depth of some gene analysis windows in certain baseline samples. Therefore, the normal range of average sequencing depth for baseline samples can be determined first based on the corrected average coverage depth of the sequencing alignment data for that gene analysis window. For example, the median or mean of the corrected average coverage depth of the sequencing alignment data for each baseline sample in that gene analysis window can be determined first. Then, the absolute value of the difference between this median or mean and the mean is less than a specified threshold (e.g., 3 standard deviations) is taken as the normal range of average sequencing depth for baseline samples.
[0295] Step 20332: Sequencing alignment data of each baseline sample whose corrected average coverage depth within the normal range of the average sequencing depth of the baseline sample in the gene analysis window is identified as the reference baseline sample sequencing alignment data corresponding to the gene analysis window.
[0296] In other words, the baseline sample sequencing data corresponding to the gene analysis window excludes those baseline sample sequencing data whose corrected average coverage depth is in an abnormal range (i.e., the absolute value of the difference between the mean and the median or mean is greater than or equal to a specified threshold). The remaining baseline samples that can represent the gene analysis window can be used as the reference baseline sample sequencing data corresponding to the gene analysis window.
[0297] Step 20333: Based on the sequencing alignment data of each reference baseline sample in the gene analysis window and the corrected average coverage depth in the gene analysis window, determine the corrected average sequencing depth of the baseline samples and the deviation of the corrected average sequencing depth of the baseline samples in the gene analysis window.
[0298] After step 20332, the sequencing data of each reference baseline sample in the gene analysis window have been removed, eliminating baseline samples with corrected average coverage depth values outside the normal range. Here, the mean of the corrected average coverage depth of each reference baseline sample in the gene analysis window can be calculated, and this mean is used as the corrected average sequencing depth of the baseline samples in the gene analysis window. Then, based on the corrected average coverage depth of each reference baseline sample in the gene analysis window and the corrected average sequencing depth of the baseline samples in the gene analysis window, the deviation of the corrected average sequencing depth is determined. The deviation of the corrected average sequencing depth can be, for example, the standard deviation, coefficient of variation, absolute median difference, Z-score (standardized deviation), etc.
[0299] Step 2034: Based on the corrected baseline sample average sequencing depth deviation of each gene analysis window, determine the reference gene analysis window in each gene analysis window.
[0300] After step 2033, the corrected baseline average sequencing depth deviation of each gene analysis window can be obtained. Here, various anomaly determination methods can be used to identify statistically abnormal corrected baseline average sequencing depth deviations among the corrected baseline average sequencing depth deviations of each gene analysis window. For example, it can be a corrected baseline average sequencing depth deviation with a Z-score greater than 3. Then, the gene analysis windows in each gene analysis window whose corrected baseline average sequencing depth deviations do not belong to the above-mentioned statistically abnormal corrected baseline average sequencing depth deviations are determined as reference gene analysis windows.
[0301] Through steps 2031 to 2034, the sequencing alignment data based on each baseline sample were screened in each gene analysis window to determine the reference gene analysis window, and for each reference gene analysis window, the sequencing alignment data of the baseline sample was screened to obtain the reference baseline sample sequencing alignment data.
[0302] Step 2035: Determine the analysis window of each reference gene containing the gene to be tested as the analysis window of the gene to be tested.
[0303] To determine the mixed copy number of the gene to be tested, the region where the gene to be tested is located is first needed, that is, the analysis window of each reference gene where the gene to be tested is located is determined as the analysis window of the gene to be tested.
[0304] Step 2036: For each gene analysis window to be tested, the mixed copy number of the gene analysis window to be tested in the test sample is determined based on the corrected average coverage depth of the sequencing alignment data of the test sample in the gene analysis window and the corrected average sequencing depth of the baseline sample in the gene analysis window.
[0305] Specifically, the corrected average coverage depth of the sample to be tested within the analysis window of the gene to be tested can be used as a reference. Divide by the average sequencing depth of the corrected baseline samples within the analysis window of the gene to be tested. The product of the ratio and 2 is used to determine the mixed copy number of the gene to be tested in the sample. That is, the following formula can be used for calculation:
[0306]
[0307] If there is The analysis window for the genes to be tested, after step 2036, can obtain the genes in the sample to be tested. The mixed copy number of each gene analysis window in the gene analysis window can be obtained. The mixed copy number of each gene analysis window to be tested.
[0308] Step 2037: The mean of the mixed copy number of each gene in the test sample is determined as the average mixed copy number of the gene in the test sample.
[0309] The optional approach in steps 2031 to 2037 above can improve the accuracy of the average coverage depth of each gene analysis window by dividing each preset capture region into smaller gene analysis windows and then correcting the average coverage depth of each gene analysis window in each sequencing alignment data. Furthermore, by filtering the sequencing alignment data of each baseline sample within each gene analysis window to determine the reference gene analysis window, and by filtering the baseline sample sequencing alignment data for each reference gene analysis window to obtain reference baseline sample sequencing alignment data, abnormal gene analysis windows and abnormal baseline samples can be removed. Then, based on the corrected average coverage depth of the test sample and the baseline sample, the mixed copy number of the reference gene analysis window containing each gene to be tested in the test sample is obtained, and the average mixed copy number of the gene to be tested in the test sample is finally determined. By combining the above steps, the average mixed copy number of the target gene in the test sample can be increased. The accuracy of the calculation.
[0310] Step 204: Based on the tumor purity and ploidy of the test sample and the average mixed copy number of the test gene in the test sample, determine the corrected tumor cell copy number of the test gene in the test sample.
[0311] Typically, to determine the corrected tumor cell copy number of the target gene in a test sample, the tumor purity of the sample and the average mixed copy number of the target gene in the sample are required. The corrected tumor cell copy number of the target gene in the test sample can be obtained by solving the traditional correction formula. The traditional correction formula can be expressed as follows:
[0312]
[0313] in, The tumor purity of the sample to be tested. This represents the average mixed copy number of the gene to be tested in the sample. Let be the corrected tumor cell copy number of the gene to be tested in the sample. Solving the above traditional correction formula yields the following formula:
[0314]
[0315] The corrected tumor cell copy number of the target gene in the sample can be calculated using the above formula. However, the above calculation method does not take into account the tumor ploidy of the test sample, which will lead to interference from polyploidy in gene copy number calculation, thus failing to accurately reveal the genomic variation characteristics of tumor cells. Therefore, in determining the corrected tumor cell copy number of the test gene in the test sample, in addition to considering the tumor purity and the average mixed copy number of the test gene in the test sample, the tumor ploidy of the test sample must also be considered. By using tumor ploidy and tumor purity to correct the tumor cell copy number of the test gene in the test sample, the accuracy of copy number variation detection can be improved, and whole-genome ploidy (such as triploidy) can be distinguished from amplification of local chromosomal regions.
[0316] Optionally, step 204 can be performed as follows: substitute the tumor purity and ploidy of the test sample, and the average mixed copy number of the test gene in the test sample into the formula for calculating the corrected tumor cell copy number, and solve to obtain the corrected tumor cell copy number of the test gene.
[0317] Optionally, the formula for calculating the corrected tumor cell copy number can be obtained by deriving the correction formula, which is shown below:
[0318]
[0319] in, The tumor purity of the sample to be tested. Tumor ploidy of the sample to be tested This represents the average mixed copy number of the gene to be tested in the sample. This represents the corrected tumor cell copy number of the gene to be tested in the sample.
[0320] The above correction formula is explained below:
[0321] The numerator on the left side of the equal sign in the corrected formula is: .in, The tumor purity of the sample to be tested. This represents the corrected tumor cell copy number of the gene to be tested in the sample. This represents the theoretical average mixed copy number of the gene to be tested in tumor cells within the tested sample. This represents the theoretical average mixed copy number of the gene to be tested in normal cells within the tested sample. This can theoretically represent the average mixed copy number of the gene to be tested in the sample.
[0322] The denominator on the left side of the equal sign in the corrected formula is: .in, and These represent the tumor purity and tumor ploidy of the sample being tested, respectively. This represents the theoretical average copy number of tumor cells in the sample being tested. This represents the theoretical average copy number of normal cells in the sample being tested. This represents the theoretical average mixed copy number of the sample to be tested.
[0323] Therefore, the left side of the equation in the correction formula is: This represents the ratio of the average mixed copy number of the gene to be tested in the test sample to the average mixed copy number of the test sample.
[0324] Correcting the numerator on the right side of the equation in the formula 1 represents the average mixed copy number of the gene to be tested in the actual test sample, and 2 represents the average mixed copy number of the actual test sample.
[0325] Therefore, the right side of the equation in the correction formula is: This represents the ratio of the average mixed copy number of the gene to be tested in the actual test sample to the average mixed copy number of the test sample.
[0326] In other words, the numerators on both sides of the equal sign correspond to the average mixed copy number of the gene to be tested in the sample, and the denominators correspond to the average mixed copy number of the sample. The left side of the equal sign corresponds to theoretical data, and the right side of the equal sign corresponds to actual data.
[0327] Because it is known , and By deriving the above correction formula, we can obtain the formula for calculating the corrected tumor cell copy number, as follows:
[0328] In this way, the corrected tumor cell copy number of the gene to be tested in the sample can be calculated.
[0329] Here, after completing step 204, you can proceed to step 205 or step 206.
[0330] Step 205: In response to the average mixed copy number of the test gene in the test sample being less than 2, the test result for homozygous deletion of the test gene in the test sample is generated based on the comparison between the corrected tumor cell copy number of the test gene in the test sample and the positive judgment threshold for homozygous deletion of the test gene.
[0331] Here, if the average mixed copy number of the gene to be tested in the sample is less than 2, then it can be determined first whether the corrected tumor cell copy number of the gene to be tested in the sample is greater than the positive threshold for homozygous deletion of the gene to be tested.
[0332] If it is determined that the value is not greater than a certain value, a test result can be generated to indicate a positive result for homozygous deletion of the gene to be tested in the sample.
[0333] If the value is determined to be greater than 1, a test result can be generated to indicate a negative result for homozygous deletion of the gene to be tested in the sample.
[0334] Here, the threshold for determining a positive result for homozygous deletion of the gene to be tested can be predetermined using various implementation methods. For example, it can be manually set by professional technicians.
[0335] Optionally, the threshold for determining a homozygous deletion positivity of the gene to be tested can be determined by, for example... Figure 3A The homozygous deletion positive threshold determination step 300 shown is predetermined and may include the following steps 301 to 304:
[0336] Step 301: Obtain the dataset of homozygous missing samples.
[0337] Here, homozygous deletion samples can include negative sample data and homozygous deletion positive sample data for the gene to be tested. Each negative sample data is obtained based on sequencing alignment data of benign samples, and each homozygous deletion positive sample data is sequencing alignment data for the gene to be tested that has a homozygous deletion.
[0338] Optionally, step 301 may include, for example: Figure 3B The following steps 3011 to 3013 are shown:
[0339] Step 3011: Obtain sequencing alignment data of N pairs of 100% tumor cell lines and their paired cell lines.
[0340] Here, N is a positive integer.
[0341] A 100% tumor cell line refers to a cell line composed of pure tumor cells, without any other non-tumor cells (such as stromal cells, immune cells, etc.).
[0342] A 100% tumor cell line paired with a control cell line refers to a cell line that originates from the same individual as the tumor cell line but has a normal phenotype.
[0343] Here, high-throughput sequencing can be performed on each 100% tumor cell line and each paired cell line separately to obtain the corresponding raw sequencing data. The raw sequencing data is then aligned back to the human reference genome (e.g., using tools like BWA or Bowtie2) to obtain the corresponding sequencing alignment data. Here, the sequencing alignment data can be, for example, a BAM file.
[0344] Step 3012: Based on N pairs of sequencing alignment data, generate at least one homozygous deletion positive sample data for the gene to be tested.
[0345] Here, the tumor purity corresponding to each homozygous missing positive sample data is the corresponding labeled tumor purity.
[0346] Optionally, step 3012 can be performed as follows: For the sequencing alignment data of each tumor cell line, a homozygous deletion positive sample data generation operation is performed. The homozygous deletion positive sample data generation operation may include, for example: Figure 3C The following steps 30121 and 30122 are shown:
[0347] Step 30121: Delete the corresponding reads of the gene to be tested from the sequencing alignment data of the tumor cell line.
[0348] The corresponding reads of the gene to be tested in the sequencing alignment data of this tumor cell line may include, but are not limited to, reads that cover a specific region of the gene to be tested (such as exons (generally), introns or regulatory regions) in the sequencing alignment data.
[0349] That is, homozygous deletion of the gene to be tested is simulated by deleting the corresponding reads of the gene to be tested.
[0350] Step 30122: For each preset tumor purity gradient in the preset tumor purity gradient set, based on the sequencing alignment data of the tumor cell line and the sequencing alignment data of the corresponding paired cell line, generate homozygous deletion positive sample data of the tumor cell line under the preset tumor purity gradient for the gene to be tested, and determine the preset tumor gradient as the labeled tumor purity of the generated homozygous deletion positive sample data.
[0351] That is, through step 302, the sequencing alignment data of each tumor cell line are used to simulate homozygous deletion simulated positive data with different tumor purity gradients, and the corresponding tumor purity is labeled for the generated homozygous deletion positive sample data.
[0352] For example, the preset tumor purity gradient is Random selection can be made from the sequencing alignment data of this tumor cell line. The reads were randomly selected from the sequencing alignment data of paired cell lines of this tumor cell line. The selected two reads are then combined to generate homozygous deletion positive sample data, thus yielding tumor purity. Data on homozygous missing positive samples.
[0353] Step 3013: Obtain sequencing alignment data of M benign samples, and generate at least one negative sample data based on the sequencing alignment data of M benign samples.
[0354] Here, a benign sample can be benign tissue or adjacent tissue. M is a positive integer.
[0355] Specifically, each benign sample can be subjected to high-throughput sequencing to obtain the corresponding raw sequencing data. This raw sequencing data is then aligned back to the human reference genome (e.g., hg38 / GRCh38) using tools such as BWA or Bowtie2, yielding the corresponding sequencing alignment data. Here, the sequencing alignment data can be, for example, a BAM file.
[0356] To increase the number of benign samples, various methods can be used to generate at least one negative sample data based on the sequencing alignment data of M benign samples.
[0357] In practice, due to various noises during the sequencing process, to adapt to real-world situations, step 3013 can optionally be performed as follows:
[0358] For each benign sample sequencing alignment data, noise of different proportions is added to the sequencing depth of the base sites in the benign sample sequencing alignment data to obtain at least one negative sample data corresponding to the benign sample sequencing alignment data.
[0359] As an example, noise can be added using the following method:
[0360] Assuming the sequencing depth of each base site follows a log-normal distribution, we first take the logarithmic value of the sequencing depth of each base site in the sequencing alignment data of this benign sample. Then, we calculate the standard deviation (sd) of the logarithms of the sequencing depths of all base sites. Next, we randomly add sigma to the logarithmic values of the corresponding proportion of base site depths in the sequencing alignment data of this benign sample, where sigma ~ rnorm(0, sd). To prevent excessively uniform noise across all base sites, we assume sd... 2 It follows an inverse gamma distribution. Finally, by taking the exponential value of the logarithmic sequencing depth for each base site, different levels of noise can be added to the sequencing depth of the base sites in the sequencing alignment data of this benign sample.
[0361] Step 302: Determine the corrected tumor cell copy number for the gene to be tested for each homozygous deletion sample data.
[0362] Here, the same method as described in steps 202, 203 and 204 for determining the corrected tumor cell copy number of the gene to be tested in the sample can be used to determine the corrected tumor cell copy number of each homozygous deletion sample for the gene to be tested.
[0363] Specifically, for each homozygous deletion sample data, the tumor purity and tumor ploidy of the homozygous deletion sample data can first be determined based on the sequencing alignment data of the homozygous deletion sample data and at least one baseline sample obtained in step 201 (using the method described in step 202); then, based on the sequencing alignment data of the homozygous deletion sample data and at least one baseline sample obtained in step 201, the average mixed copy number of the gene to be tested in the homozygous deletion sample data can be determined; finally, based on the tumor purity and tumor ploidy of the homozygous deletion sample data and the average mixed copy number of the gene to be tested in the homozygous deletion sample data, the corrected tumor cell copy number of the gene to be tested in the homozygous deletion sample data can be determined.
[0364] Step 303: For each candidate homozygous missing copy number in the candidate homozygous missing copy number set, perform the operation to determine the criteria for homozygous missing positive results.
[0365] Here, the candidate homozygous missing copy number set can be preset by professionals based on the possible range of values for the actual homozygous copy number in practice.
[0366] Here, the procedures for determining the criteria for a homozygous deletion positive result may include, for example: Figure 3D The following steps 3031 and / or 3032 are shown:
[0367] Step 3031: Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, determine the positive concordance rate corresponding to the candidate homozygous deletion copy number.
[0368] Specifically, step 3031 can be performed as follows:
[0369] First, based on the comparison between the corrected tumor cell copy number for the target gene and the candidate homozygous deletion copy number of each homozygous deletion-positive sample, the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number are counted. Then, for each homozygous deletion-positive sample, it is determined whether the corrected tumor cell copy number for the target gene is greater than the candidate homozygous deletion-positive copy number. If it is determined not to be greater, the homozygous deletion-positive sample is determined to be a homozygous deletion-positive sample. If it is determined to be greater, the homozygous deletion-positive sample is determined to be a homozygous deletion-negative sample.
[0370] Then, the number of homozygous deletion positive samples that were detected as homozygous deletion positive in each homozygous deletion positive sample data was counted, and this number was taken as the true positive count corresponding to the candidate homozygous deletion copy number. The number of homozygous deletion positive samples that were detected as homozygous deletion negative samples was counted, and this number was taken as the false negatives corresponding to that candidate homozygous deletion copy number. .
[0371] Finally, the positive concordance rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number.
[0372] Here, PPA (Positive Percent Agreement) can be calculated using the following formula:
[0373]
[0374] in, and These represent the number of true positives and false negatives corresponding to the homozygous deleted copy number of the candidate, respectively. It is the calculated positive concordance rate corresponding to the homozygous missing copy number of the candidate.
[0375] Step 3032: Based on the comparison results of the corrected tumor cell copy number of the gene to be tested for each homozygous deletion sample data with the candidate homozygous deletion copy number, determine the positive predictive value corresponding to the candidate homozygous deletion copy number.
[0376] Specifically, step 3032 can be performed as follows:
[0377] First, based on the comparison between the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted.
[0378] For details, please refer to the relevant description in step 3031, which will not be repeated here.
[0379] Then, based on the comparison between the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number of each negative sample, the number of false positives corresponding to the candidate homozygous deletion copy number is counted.
[0380] First, for each negative sample data point, determine whether the corrected tumor cell copy number for the target gene is greater than the candidate homozygous deletion positive copy number. If it is not greater, the negative sample data point is determined to be a homozygous deletion positive. If it is greater, the negative sample data point is determined to be a homozygous deletion negative. Then, count the number of negative sample data points that are detected as homozygous deletion positive, which is taken as the number of false positive data points corresponding to the candidate homozygous deletion copy number. .
[0381] Finally, based on the number of true positives and false positives corresponding to the candidate homozygous deletion copy number, the positive predicted value corresponding to the candidate homozygous deletion copy number is determined.
[0382] Specifically, for a candidate homozygous deletion copy number, the positive predictive value corresponding to that candidate homozygous deletion copy number can be calculated using the following formula. (Positive Predictive Value):
[0383]
[0384] in, and These represent the number of true positives and false positives corresponding to the homozygous deleted copy number of the candidate, respectively. It is the calculated positive prediction value corresponding to the homozygous missing copy number of the candidate.
[0385] Step 304: Determine the positive judgment threshold for homozygous deletion of the gene to be tested based on the positive concordance rate and / or positive prediction value corresponding to each candidate homozygous deletion copy number.
[0386] Here, we can first determine the preferred homozygous deletion copy number that meets the preset detection criteria in the candidate homozygous deletion copy number set, based on the corresponding positive concordance rate and / or positive predictive value. The preset detection criteria can be pre-set according to the needs of the specific detection application scenario. As an example, the preset detection criteria can be: PPV ≥ 99% and PPA ≥ 95%. Then, based on the aforementioned determined preferred homozygous deletion copy numbers, we determine the positive threshold for homozygous deletion of the gene to be tested. For example, the maximum value among the determined preferred homozygous deletion copy numbers can be used as the positive threshold for homozygous deletion of the gene to be tested.
[0387] The optional methods of steps 301 to 304 above can achieve, but are not limited to, the following technical effects:
[0388] First, by using sequencing alignment data from a small number of 100% tumor cell lines and their paired cell lines, homozygous deletion positive sample data is generated for the gene to be tested, reducing the sequencing cost required to generate homozygous deletion positive sample data.
[0389] Secondly, by generating a large amount of negative sample data based on sequencing alignment data of a small number of benign samples, the sequencing cost of benign samples is reduced.
[0390] In addition, by taking into account the tumor ploidy of homozygous deletion positive and negative sample data respectively in determining the corrected tumor cell copy number of the gene to be tested, the interference of aneuploidy (polyploid or haploid) on the test results can be eliminated.
[0391] Step 206: In response to the average mixed copy number of the test gene in the test sample being greater than 2, the amplification judgment value of the test gene in the test sample is determined based on the corrected tumor cell copy number of the test gene in the test sample, and the detection result of the copy number amplification of the test gene in the test sample is generated based on the comparison result between the amplification judgment value of the test gene in the test sample and the amplification positive judgment threshold of the test gene.
[0392] Here, if the average mixed copy number of the gene to be tested in the sample is greater than 2, various implementation methods can be used to determine the amplification judgment value of the gene to be tested in the sample based on the corrected tumor cell copy number of the gene to be tested in the sample.
[0393] Optionally, the corrected tumor cell copy number of the gene to be tested in the sample can be used. The amplification threshold for the gene to be tested in the test sample is determined as follows: Alternatively, the difference between the corrected tumor cell copy number of the gene to be tested in the test sample and the tumor ploidy of the test sample is used. The amplification threshold for the target gene in the test sample is determined as follows: Alternatively, the ratio of the corrected tumor cell copy number of the target gene in the test sample to the tumor ploidy of the test sample is used. The amplification judgment value of the gene to be tested in the sample to be tested is determined.
[0394] Preferably, the difference between the corrected tumor cell copy number of the gene to be tested in the test sample and the tumor ploidy of the test sample can be used. The amplification threshold for the target gene in the test sample is determined. This is to account for cases where the entire genome of a tumor patient has been amplified, for example, to tetraploid, hexaploid, or octaploid, then the corrected tumor cell copy number of the target gene in the test sample... The amplification is also performed accordingly, for example, amplified to 8. This is similar to the correction of the tumor cell copy number of the gene to be tested in the sample. If the amplification value of the test gene in the sample is determined, it is impossible to distinguish whether the test gene itself has been amplified. However, if the amplification value is determined as the amplification threshold, it is impossible to distinguish whether the test gene itself has been amplified. Determining the amplification threshold of the target gene in the test sample solves the problem of determining whether the target gene has amplified when amplification occurs throughout the entire genome. This is because... It targets the gene to be tested in the sample, and It applies to the entire sample being tested. This can indicate whether the gene to be tested in the sample has undergone copy number amplification relative to the overall genome of the sample.
[0395] With adoption For the same reason as the amplification judgment value of the gene to be tested in the sample to be tested, it can also be As a criterion for the amplification of the gene to be tested in the sample to be tested.
[0396] After obtaining the amplification judgment value of the gene to be tested in the sample, it can be determined whether the amplification judgment value of the gene to be tested in the sample is greater than the amplification positive judgment threshold.
[0397] If the result is determined to be greater than the specified value, a test result can be generated to indicate a positive amplification of the test gene copy number in the test sample.
[0398] If it is determined that the value is not greater than the specified value, a test result can be generated to indicate that the amplification of the test gene copy number in the test sample is negative.
[0399] In some alternative implementations, the threshold for determining amplification positivity of the gene to be tested can be determined by, for example... Figure 4A The amplification positive judgment threshold determination step 400 shown is predetermined and may include, for example, steps such as: Figure 4A Steps 401 to 405 are shown below:
[0400] Step 401: Obtain the amplified sample data set.
[0401] Here, the amplified sample data includes negative sample data and amplified positive sample data for the gene to be tested. Each negative sample data is obtained based on sequencing alignment data of benign samples, and each amplified positive sample data is sequencing alignment data of copy number amplification of the gene to be tested.
[0402] Optionally, step 401 may include, for example: Figure 4B The following steps 4011 and 4012 are shown:
[0403] Step 4011: Obtain sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines.
[0404] Here, I is a positive integer.
[0405] Here, the specific operation of step 4011 and its resulting technical effects are related to... Figure 3A The operation and effect of step 3011 in the illustrated embodiments are basically the same, and will not be repeated here.
[0406] Step 4012: Based on the I-pair sequencing alignment data, generate at least one amplified positive sample data for the gene to be tested.
[0407] Here, the tumor purity of each amplified positive sample data generated is the corresponding labeled tumor purity.
[0408] Optionally, step 4012 can be performed as follows: For the sequencing alignment data of each tumor cell line, perform an amplified positive sample data generation operation, which may include, for example... Figure 4C The following steps 40121 and 40122 are shown:
[0409] Step 40121: For each preset copy number gradient in the preset copy number gradient set, add the corresponding read of the gene to be tested to the sequencing alignment data of the tumor cell line to generate amplified positive sample data of the tumor cell line under the preset copy number gradient.
[0410] Here, the tumor cell copy number of the amplified positive sample data of the tumor cell line under the preset copy number gradient is the preset copy number gradient.
[0411] The corresponding reads of the gene to be tested in the sequencing alignment data of this tumor cell line may include, but are not limited to, reads that cover a specific region of the gene to be tested (such as exons (generally), introns or regulatory regions) in the sequencing alignment data.
[0412] That is, by adding different proportions of the corresponding reads of the target gene to the sequencing alignment data of this tumor cell line, the copy number amplification of the target gene under different copy numbers is simulated.
[0413] For example, the sequencing alignment data of this tumor cell line originally contained the following reads corresponding to the gene to be tested: Segment, if the preset copy number gradient is Here, you can add sequencing alignment data for this tumor cell line. The corresponding reads for each gene to be tested. That is, the number of reads to be added. Satisfy the following formula:
[0414]
[0415] Here, the default copy number of the gene to be tested in the tumor cell line is 2, and a preset copy number gradient is simulated based on this. .
[0416] Step 40122: For each preset tumor purity gradient in the preset tumor purity gradient set, based on the sequencing alignment data of the tumor cell line and the sequencing alignment data of the corresponding paired cell line, generate amplified positive sample data of the tumor cell line under each tumor cell copy number gradient and the preset tumor purity gradient, and determine the preset tumor gradient as the labeled tumor purity of the generated amplified positive sample data.
[0417] That is, through step 4012, the sequencing alignment data of each tumor cell line is used to simulate amplified positive data with different tumor purity gradients, and the corresponding tumor purity is labeled for the generated amplified positive sample data.
[0418] For example, the preset tumor purity gradient is Random selection can be made from the sequencing alignment data of this tumor cell line. The reads were randomly selected from the sequencing alignment data of paired cell lines of this tumor cell line. The selected two reads are then mixed to generate amplified positive sample data, thus yielding tumor purity. The amplified positive sample data.
[0419] Step 4013: Obtain sequencing alignment data of J benign samples, and generate at least one negative sample data based on the sequencing alignment data of J benign samples.
[0420] Here, J is a positive integer.
[0421] Here, the specific operation of step 4013 and its resulting technical effects are related to... Figure 3B The operation and effect of step 3013 in the illustrated embodiments are basically the same, and will not be repeated here.
[0422] Step 402: Determine the corrected tumor cell copy number for each amplified sample data.
[0423] Here, the same method as described in steps 202, 203 and 204 for determining the corrected tumor cell copy number of the test gene in the test sample can be used to determine the corrected tumor cell copy number of each amplified sample data for the test gene.
[0424] Specifically, for each amplified sample data, the tumor purity and tumor ploidy of the amplified sample data can first be determined based on the sequencing alignment data of the amplified sample data and at least one baseline sample obtained in step 201 (using the method described in step 202); then, based on the sequencing alignment data of the amplified sample data and at least one baseline sample obtained in step 201, the average mixed copy number of the gene to be tested in the amplified sample data can be determined; finally, based on the tumor purity and tumor ploidy of the amplified sample data and the average mixed copy number of the gene to be tested in the amplified sample data, the corrected tumor cell copy number of the gene to be tested in the amplified sample data can be determined.
[0425] Step 403: Determine the amplification judgment value of the corresponding amplified sample data for the gene to be tested based on the corrected tumor cell copy number of each amplified sample data.
[0426] Here, various methods can be used to determine the amplification judgment value of each amplified positive sample data for the gene to be tested based on the corrected tumor cell copy number for the gene to be tested.
[0427] It is understandable that the specific operation and technical effect of determining the amplification judgment value of the amplified positive sample data for the target gene based on the corrected tumor cell copy number of the target gene in the target sample in step 206 are basically the same as the operation and effect of determining the amplification judgment value of the target gene in the target sample based on the corrected tumor cell copy number of the target gene in the target sample, and will not be repeated here.
[0428] Step 404: For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, perform the operation of determining the amplification positive judgment basis.
[0429] Here, the set of candidate amplification positive judgment values can be preset by professional technicians based on the possible range of values for amplification positive judgment values in practice.
[0430] Here, the procedures for determining the criteria for amplification positivity may include, for example: Figure 4D The following steps 4041 and / or 4042 are shown:
[0431] Step 4041: Based on the comparison results of the amplification judgment value of each amplified positive sample data for the gene to be tested and the candidate amplified positive judgment value, determine the positive concordance rate corresponding to the candidate amplified positive judgment value.
[0432] Specifically, step 4041 can be performed as follows:
[0433] First, based on the comparison between the positive amplification judgment value of each positive amplification sample data for the gene to be tested and the positive amplification judgment value of the candidate amplification, the number of true positives and false negatives corresponding to the positive amplification judgment value of the candidate amplification are counted.
[0434] First, for each positive amplification sample, determine whether the positive amplification threshold for the target gene is greater than the candidate positive amplification threshold. If it is greater, the positive amplification sample is considered positive. If it is not greater, the positive amplification sample is considered negative.
[0435] Then, the number of amplified positive samples that were detected as amplified positive in each amplified positive sample data set is counted, and this number is taken as the true positive count corresponding to the candidate amplified positive judgment value. The number of samples that tested negative for amplification in each positive amplification sample data point is counted, and this number is taken as the false negative number corresponding to the candidate positive amplification value. .
[0436] Finally, based on the number of true positives and false negatives corresponding to the candidate amplification positive judgment value, the positive concordance rate corresponding to the candidate amplification positive judgment value is determined.
[0437] Here, PPA (Positive Percent Agreement) can be calculated using the following formula:
[0438]
[0439] in, and These represent the number of true positives and false negatives corresponding to the candidate amplification positive judgment value, respectively. It is the calculated positive concordance rate corresponding to the positive judgment value of the amplification.
[0440] Step 4042: Based on the comparison results of the amplification judgment value of each amplification sample data for the gene to be tested and the positive judgment value of the candidate amplification, determine the positive prediction value corresponding to the positive judgment value of the candidate amplification.
[0441] Specifically, step 4042 can be performed as follows:
[0442] First, based on the comparison between the positive amplification judgment value of each positive amplification sample for the gene to be tested and the positive amplification judgment value of the candidate amplification, the number of true positives corresponding to the positive amplification judgment value of the candidate amplification is counted.
[0443] For details, please refer to the relevant description in step 4041, which will not be repeated here.
[0444] Then, based on the comparison between the positive amplification judgment value for the target gene in each negative sample and the candidate positive amplification judgment value, the number of false positives corresponding to the candidate positive amplification judgment value is counted. First, for each negative sample, it is determined whether the positive amplification judgment value for the target gene in that negative sample is greater than the candidate positive amplification judgment value. If it is greater, the negative sample is determined to be positive for amplification. If it is not greater, the negative sample is determined to be negative for amplification. Then, the number of negative samples that are detected as positive for amplification is counted, and this number is taken as the number of false positives corresponding to the candidate positive amplification judgment value. .
[0445] Finally, based on the number of true positives and false positives corresponding to the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined.
[0446] Specifically, for a candidate amplification positive judgment value, the positive predicted value corresponding to that candidate amplification positive judgment value can be calculated according to the following formula. (Positive Predictive Value):
[0447] in, and These represent the number of true positives and the number of false positives corresponding to the candidate amplification positive judgment value, respectively. It is the calculated positive predictive value corresponding to the positive judgment value of the candidate amplification.
[0448] Step 405: Determine the positive amplification threshold of the gene to be tested based on the positive prediction value corresponding to each candidate positive amplification judgment value.
[0449] Here, we can first determine the preferred amplification positive judgment values from the candidate amplification positive judgment value set, which meet the preset detection criteria and have the corresponding positive concordance rate and / or positive predictive value. The preset detection criteria can be pre-set according to the needs of the specific detection application scenario. As an example, the preset detection criteria can be: PPV ≥ 99% and PPA ≥ 95%.
[0450] Then, based on the aforementioned preferred positive amplification values, the positive amplification threshold for the gene to be tested is determined. For example, the minimum value among the preferred positive amplification values can be used as the positive amplification threshold for the gene to be tested.
[0451] The optional methods of steps 401 to 405 above can achieve, but are not limited to, the following technical effects:
[0452] First, by using sequencing alignment data from a small number of 100% tumor cell lines and their paired cell lines, amplified positive sample data is generated for the gene to be tested, reducing the sequencing cost required to generate amplified positive sample data.
[0453] Secondly, by generating a large amount of negative sample data based on sequencing alignment data of a small number of benign samples, the sequencing cost of benign samples is reduced.
[0454] In addition, by taking into account the tumor ploidy of amplified positive and negative sample data separately when determining the corrected tumor cell copy number of the amplified sample data for the gene to be tested, the interference of aneuploidy (polyploid or haploid) on the test results can be eliminated.
[0455] In some alternative implementations, prior to step 202, the method may also perform the following base site sequencing depth normalization operation for each sequencing alignment data:
[0456] First, for each preset capture region, determine the sequencing depth of the sequencing alignment data at each base site within that preset capture region.
[0457] Then, the sum of the sequencing depths of all base sites in each preset capture region of the sequencing alignment data is determined as the total sequencing depth of the sequencing alignment data.
[0458] Finally, for each base site in each preset capture region, the sequencing depth of that base site is normalized based on the total sequencing depth of the sequencing alignment data.
[0459] Specifically, for each base site in each preset capture region, the product of the ratio of the sequencing depth of that base site to the total sequencing depth of the sequencing alignment data and the global scaling factor can be used as the sequencing depth after normalization of that base site.
[0460] By using the above optional implementation methods, the difference in sequencing data volume between different samples can be eliminated.
[0461] The gene copy number variation type detection method provided in the above embodiments of this disclosure first acquires sequencing alignment data of gene sequencing performed on the test sample and at least one baseline sample, wherein the capture region of the gene sequencing process includes a preset capture region set; then, based on each sequencing alignment data, the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the test gene in the test sample are determined; then, based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the test gene in the test sample, the corrected tumor cell copy number of the test gene in the test sample is determined; next, if If the average mixed copy number of the target gene is less than 2, the detection result for homozygous deletion of the target gene in the test sample is determined by comparing the corrected tumor cell copy number of the target gene in the test sample with the positive threshold for homozygous deletion of the target gene. If the average mixed copy number of the target gene is greater than 2, the amplification judgment value of the target gene in the test sample is determined by comparing the amplification judgment value of the target gene in the test sample with the positive threshold for amplification, thus determining the detection result for copy number amplification of the target gene in the test sample. Therefore, by considering the tumor ploidy of the test sample in the process of determining the corrected tumor cell copy number of the target gene in the test sample, the accuracy of the corrected tumor cell copy number of the target gene in the test sample can be improved due to the richer range of factors considered. This, in turn, can improve the accuracy of detecting copy number variation states of the target gene.
[0462] Example 1:
[0463] use Figure 3A The steps for determining the homozygous deletion positivity threshold shown are as follows:
[0464] 1. Select two cell lines: a tumor cell line with 100% tumor percentage and its paired cell line. Perform high-throughput sequencing on both. Align the raw sequencing data with the human hg19 genome and obtain the sequencing alignment data (BAM file). This results in four BAM files: two tumor cell line sequencing alignment data files and two paired cell line sequencing alignment data files.
[0465] 2. Generate homozygous missing positive sample data:
[0466] (1) For the 9 exons of the PTEN gene, a total of 45 homozygous deletion types of consecutive exons to be investigated were simulated, including: 9 types of single exon deletion, 8 types of consecutive deletion of two exons, 7 types of consecutive deletion of three exons, ..., 1 type of whole gene deletion.
[0467] In other words, for each tumor cell line sequencing alignment data, the following methods can be used: removing one consecutive exon yields nine different homozygous deletion positive samples; removing two consecutive exons yields eight different homozygous deletion positive samples, and so on, until removing nine consecutive exons yields one homozygous deletion positive sample. This results in a total of 45 homozygous deletion positive samples with different consecutive exon types to be investigated. Since there are two tumor cell line sequencing alignment data sets, 90 homozygous deletion positive samples (BAM file) can be obtained (2 tumor cell lines * 45 homozygous deletion types).
[0468] (2) Simulating different tumor percentage gradients: For each homozygous deletion positive sample data from the above 90 homozygous deletion positive samples, reads were randomly extracted from the sequencing alignment data of its paired cell line and mixed to form different tumor percentage gradients, specifically including the following 8 tumor percentage gradients: 20%, 30%, 40%, 50%, 60%, 70%, 80%, and 90%. Moreover, for different consecutive deletion exon homozygous deletion types, the above tumor gradient mixing operation was repeated a different number of times, generating a total of 3744 homozygous deletion positive sample data. For the specific distribution of PTEN homozygous deletion positive sample data, please refer to Table 1.
[0469] Table 1. Data on PTEN-deleted positive samples
[0470]
[0471] As shown in Table 1, ten tumor gradient mixing operations were performed on each homozygous deletion-positive sample data with consecutive deletion exon numbers of 1, 3, 6, and 9. That is, mixing was performed 10 times according to the proportion of each gradient. However, since reads were randomly extracted from the sequencing alignment data of paired cell lines each time, the 10 mixed homozygous deletion-positive sample data resulting from the 10 tumor gradient mixing operations for the same consecutive deletion exon type and the same tumor proportion gradient are also different.
[0472] For the remaining homozygous deletion types, a single tumor gradient fusion operation was performed. This was because repeating the tumor gradient fusion operation ten times for all deletion types would result in an excessively large dataset.
[0473] 3. Generation of negative sample data:
[0474] Eighty-five benign tissue samples with normal PTEN expression and no tumor cells were selected from pathological assessments. High-throughput sequencing was performed, and the raw sequencing data was compared with the human hg19 genome to obtain sequencing alignment data (BAM files). In other words, 85 sequencing alignment data were obtained.
[0475] Then, noise was added to the sequencing depths of the aforementioned 85 sequencing alignment data points at base sites with different noise gradients to simulate different noise levels. Specifically, there were four noise gradients: 25%, 33%, 50%, and 100%, resulting in 340 negative samples with added noise. Combined with the original 85 negative samples without added noise, a total of 425 negative samples were obtained. Detailed information about the negative sample data can be found in Table 2.
[0476] Table 2 PTEN Missing Negative Data
[0477]
[0478] 4. Using the above Figure 3A In the illustrated embodiment, steps 304 and 305 determine the corrected tumor cell copy number for PTEN for each of the 3744 homozygous deletion positive samples and 425 negative samples.
[0479] 5. Using the above Figure 3A In the illustrated embodiment, step 306 involves performing a homozygous deletion positive prediction value determination operation for each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, ultimately determining the homozygous deletion positive judgment threshold for the gene PTEN to be tested to be 0.6. Table 3 shows the PPA and PPV values of 425 negative sample data corresponding to different tumor proportions and different homozygous deletion types when the candidate homozygous deletion copy number is 0.6.
[0480] Table 3. Detection performance of different PTEN gene deletion types and tumor purity
[0481]
[0482] As can be seen from Table 3, with a tumor percentage (also known as tumor purity) of 0.3 as the dividing point, when the tumor percentage is lower than 0.3 (in the 3744 homozygous deletion positive samples, the 20% tumor percentage gradient consists of 3744÷8=468 homozygous deletion positive samples), all deletion types (i.e., the 468 homozygous deletion positive samples) cannot meet the requirements of PPV≥99% and PPA≥95%.
[0483] When the tumor percentage is greater than or equal to 0.3 (including 3276 homozygous deletion positive samples), the positive cutoff criteria (PPV ≥ 99% and PPA ≥ 95%) cannot be met only when one exon is missing (including 2 (2 tumor cell lines) * 9 (9 types of missing exons) * 10 (10 repeated tumor percentage mixing operations) * 7 (7 tumor percentage gradients for tumor percentages greater than or equal to 0.3) = 1260 homozygous deletion positive samples). For other deletion types (including 3276 - 1260 = 2016 homozygous deletion positive samples), the criteria of PPV ≥ 99% and PPA ≥ 95% can be met.
[0484] 6. Limit of Detection Confirmation: The corrected tumor cell copy number (GCN) ≤ 0.6 was used as the threshold for homozygous deletion positivity in PTEN. Probit analysis was performed for 1-2 exon deletion types and 3-9 exon deletion types, and Probit curves were plotted to obtain the results. Figure 5 The two Probit curves in the image. Figure 5 The left-hand plot corresponds to 1-2 exon deletion types, and the right-hand plot corresponds to 3-9 exon deletion types. The horizontal axis of each Probit curve represents the tumor percentage, and the vertical axis represents the percentage of per-tumor area (PPA). Figure 5 As shown, the limit of detection (LOD) is 43.61% when the PTEN gene has 1-2 exon deletions; and 31.71% when the PTEN gene has 3-9 exon deletions. The LODs for different deletion types are shown in Table 4.
[0485] Table 4. Limits of detection for different deletion types of the PTEN gene
[0486]
[0487] 7. Third-party IHC verification:
[0488] Ninety-seven cancer tissue samples were selected and used as follows: Figure 2A The method described above detects the homozygous deletion status of the PTEN gene according to the established threshold for judging homozygous deletion positivity (i.e., 0.6). Simultaneously, a third-party IHC method is used for verification. The IHC evaluation criteria are: an IHC staining index of 0 indicates IHC positivity, i.e., PTEN homozygous deletion positivity; otherwise, it indicates IHC negativity, i.e., PTEN homozygous deletion negativity. In this embodiment, the evaluation criteria are: a corrected tumor cell copy number greater than or equal to 0.6 indicates PTEN homozygous deletion positivity; otherwise, it indicates PTEN homozygous deletion negativity. The results show that the detection method in this embodiment has a consistency of 85.57% with IHC, indicating high consistency. The consistency results are shown in Table 5.
[0489] Table 5. Consistency between the test results disclosed herein and IHC
[0490]
[0491] Example 2:
[0492] use Figure 3A The steps for determining the positive threshold for homozygous deletion are as follows: The specific steps for determining the positive threshold for homozygous deletion of the tested genes MTAP, CDKN2A, and CDKN2B are as follows:
[0493] 1. Obtain sequencing data: Select 13 cell lines, namely tumor cell lines with 100% tumor content and their paired cell lines, and perform high-throughput sequencing. The raw sequencing data is compared with the human hg19 genome, and the sequencing alignment data (BAM file) is obtained.
[0494] 2. Generate homozygous missing positive sample data
[0495] (1) The sequencing data of the above 13 tumor cell lines were used to remove the gene reads of the target genes MTAP, CDKN2A and CDKN2B respectively, and homozygous deletion positive sample data for MTAP, CDKN2A and CDKN2B were generated respectively, that is, 13*3=39 homozygous deletion positive sample data were obtained.
[0496] (2) Simulating different tumor percentage gradients: Reads were randomly extracted from the sequencing alignment data (BAM file) of the paired cell lines of the above 39 homozygous deletion positive samples and mixed to form different tumor percentage gradients, specifically including 5 tumor percentage gradients: 20%, 30%, 35%, 40%, and 50%. The above tumor percentage gradient mixing operation was repeated 12 times for each homozygous deletion positive sample, resulting in 780*3=2340 homozygous deletion positive sample data. Please refer to Table 6 for the distribution of homozygous deletion positive sample data of MTAP, CDKN2A, and CDKN2B genes.
[0497] Table 6. Data of positive samples with deletions of the gene MTAP / CDKN2A / CDKN2B to be tested
[0498]
[0499] 3. Generation of negative sample data:
[0500] Eighty-six benign tissue samples with normal expression of MTAP, CDKN2A, and CDKN2B genes in all pathologically assessed tumor-free tissues were selected. These samples, along with paired cell lines from the aforementioned 13 cell lines, were subjected to high-throughput sequencing. The raw sequencing data were compared with the human hg19 genome, and sequencing alignment data (BAM file) was obtained. This yielded sequencing alignment data for 99 negative samples, or data for 99 negative samples. The distribution of the negative sample data is shown in Table 7.
[0501] Table 7. Negative data on gene deletions of MTAP / CDKN2A / CDKN2B.
[0502]
[0503] 4. Using the above Figure 3A In the illustrated embodiments, steps 304 and 305 determine the corrected tumor cell copy number for MTAP, CDKN2A, and CDKN2B for each of the 2340 homozygous deletion positive samples and 99 negative samples.
[0504] 5. Using the above Figure 3A In the illustrated embodiment, step 306 involves performing a homozygous deletion positive prediction value determination operation for each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, ultimately determining the homozygous deletion positive judgment threshold for the test genes MTAP, CDKN2A, and CDKN2B to be 0.6.
[0505] Table 8 shows the PPA and PPV values of 99 negative samples for different test genes and different tumor proportions when the candidate homozygous deletion copy number is 0.6.
[0506] Table 8. Detection performance of MTAP and CDKN2A genes with a positive cutoff value of 0.6
[0507]
[0508] As can be seen from Table 8, with a tumor percentage (also known as tumor purity) of 0.3 as the dividing point, when the tumor percentage is lower than 0.3, the three test genes MTAP, CDKN2A and CDKN2B cannot meet the positive judgment value criteria of PPV ≥ 99% and PPA ≥ 95%.
[0509] When the tumor percentage is greater than or equal to 0.3%, all three genes to be tested satisfy PPV ≥ 99% and PPA ≥ 95%. Therefore, the corrected tumor cell copy number GCN ≤ 0.6 was selected as the threshold for homozygous deletion positivity of MTAP, CDKN2A, and CDKN2B.
[0510] 6. Limit of Detection Confirmation: The corrected tumor cell copy number (GCN) ≤ 0.6 was used as the threshold for homozygous deletion positivity of MTAP, CDKN2A, and CDKN2B. Probit analysis was performed on the three target genes MTAP, CDKN2A, and CDKN2B, and Probit curves were plotted to obtain the results. Figure 6A , Figure 6B and Figure 6C The Probit curve is used to evaluate the limit of detection for sample data. Figure 6A , Figure 6B and Figure 6C The Probit curves in the figure correspond to the Probit curves of the three genes to be tested: MTAP, CDKN2A, and CDKN2B. The horizontal axis of the Probit curve is the tumor percentage, and the vertical axis is PPA.
[0511] like Figure 6A , Figure 6B and Figure 6C As shown, the limit of detection (LOD) for MTAP is 27%, for CDKN2A it is 21%, and for CDKN2B it is 21%.
[0512] 7. Confirmation of MTAP homozygous deletion rate: Sequencing alignment data from 2401 cancer tissue samples were selected and used... Figure 2A The method described above detects homozygous deletion of the MTAP gene according to the established threshold for positive homozygous deletion (i.e., 0.6). The detection rate is then compared with the incidence rates of different cancers in the cBiPortal public database to examine whether the detection rate of this embodiment meets expectations. The evaluation criteria in this embodiment are: a corrected tumor cell copy number greater than or equal to 0.6 indicates a positive homozygous deletion of MTAP; otherwise, it indicates a negative homozygous deletion of MTAP. The results show that the incidence rate of homozygous deletion of MTAP detected by this embodiment is basically consistent with that in the public database, meeting expectations. The consistency results are shown in Table 9.
[0513] Table 9. Incidence of homozygous deletion of the MTAP gene.
[0514]
[0515] In Table 9, the meaning of 27% (28 / 105) in the homozygous deletion rate of the MTAP gene in this embodiment is as follows:
[0516] The denominator 105 indicates that a total of 105 samples diagnosed with bladder urothelial carcinoma were tested using the method of this embodiment;
[0517] Mole 27 indicates that 28 samples were found to have homozygous deletion positivity for MTAP using the method of this embodiment.
[0518] As can be seen from Examples 1 and 2, the gene copy number variation type detection method provided in this disclosure performs well in detecting homozygous deletions of the gene to be tested.
[0519] Example 3:
[0520] use Figure 4A The steps for determining the amplification positivity threshold shown are as follows:
[0521] 1. Obtaining Sequencing Data: Seven cell lines were selected, representing tumor cell lines with 100% tumor percentage and their paired cell lines. High-throughput sequencing was performed, and the raw sequencing data was compared with the human hg19 genome to obtain sequencing alignment data (BAM file). The seven cell lines included three diploid (i.e., Ploidy 2), two triploid (i.e., Ploidy 3), and four tetraploid (i.e., Ploidy 4).
[0522] 2. Generate amplified positive sample data
[0523] (1) Sequencing alignment data of the above 7 tumor cell lines were used to insert reads of the target gene CCNE1 according to different amplification ratios, generating amplified positive sample data for the target gene CCNE1, that is, 7 amplified positive sample data were obtained. Specifically, there are 4 different amplification copy numbers: 5, 6, 7, 8, thus obtaining 7*4=28 amplified positive sample data.
[0524] (2) Simulating different tumor percentage gradients: Reads were randomly extracted from the sequencing alignment data (BAM file) of the paired cell lines of the above 28 amplified positive sample data and mixed into different gradients of tumor percentage, specifically including 4 tumor percentage gradients: 15%, 20%, 25%, and 30%. The above tumor percentage gradient mixing operation was repeated 7 times for each amplified positive sample data, resulting in 7 (7 tumor cell lines) * 4 (4 simulated copy number gradients) * 4 (4 tumor percentage gradients) * 7 (repeated tumor percentage mixing 7 times) = 784 amplified positive sample data. Please refer to Table 10 for the distribution of CCNE1 gene amplified positive sample data.
[0525] Table 10. Data on positive samples for CCNE1 copy number amplification.
[0526]
[0527] 3. Generation of negative sample data:
[0528] Seventy-eight benign tissue samples with normal CCNE1 gene expression and no tumor cells were selected from pathological assessment. High-throughput sequencing was performed, and the raw sequencing data was compared with the human hg19 genome to obtain sequencing alignment data (BAM file). This yielded data from 78 negative samples.
[0529] 4. Using the above Figure 4A In the illustrated embodiment, steps 404 and 405 determine an amplification judgment value for the target gene CCNE1 for each of the 784 homozygous deletion positive samples and 78 negative samples. The amplification judgment value is the difference between the corrected tumor cell copy number and the tumor ploidy, i.e., (GCN-ploidy).
[0530] 5. Using the above Figure 4A In the illustrated embodiment, step 406 involves performing an amplification positive prediction value determination operation for each candidate amplification positive judgment value in the candidate amplification positive judgment value set, ultimately determining the amplification positive judgment threshold of the test gene CCNE1 as 2.5.
[0531] Table 11 shows the number of data points that were detected as true positives and false negatives in 784 amplified positive samples with different tumor cell copy numbers and different tumor proportions when the candidate amplification positive judgment value was 2.5, as well as the number of data points that were detected as true negatives and false positives in 78 negative samples, and the corresponding PPA and PPV values.
[0532] Table 11 Detection performance of the test gene CCNE1 with a positive cutoff value of 2.5
[0533]
[0534] As shown in Table 11, when the candidate amplification positive judgment value is 2.5 and the tumor proportion is greater than or equal to 0.3%, the amplification positive detection of the test gene CCNE1 meets the requirements of PPV ≥ 99% and PPA ≥ 95%. Therefore, the amplification positive judgment threshold for the test gene CCNE1 can be 2.5.
[0535] 6. Limit of Detection Confirmation: The threshold for confirming the amplification positivity of the test gene CCNE1 is set at a difference of ≥2.5 between the corrected tumor cell copy number and the tumor ploidy (GCN-ploidy). Probit analysis is performed on the test gene CCNE1, and a Probit curve is plotted to obtain the detection limit. Figure 7A , Figure 7B and Figure 7C The Probit curve is used to evaluate the limit of detection for sample data. Figure 7A , Figure 7B and Figure 7C The Probit curves in the figures correspond to the Probit curves of diploid CCNE1 amplification, triploid CCNE1 amplification, and tetraploid CCNE1 amplification, respectively. The horizontal axis of the Probit curves represents the tumor percentage, and the vertical axis represents the PPA.
[0536] like Figure 7A , Figure 7B and Figure 7C As shown, the limits of detection for the diploid, triploid, and tetraploid versions of the gene CCNE1 to be tested are 24%, 20%, and 16%, respectively. Therefore, the limit of detection for the gene CCNE1 to be tested can be determined to be 24%.
[0537] 7. Confirmation of CCNE1 copy number amplification rate: Sequencing alignment data from 1528 breast cancer samples were selected and used... Figure 2A The method described above detects the copy number amplification status of the CCNE1 gene according to the established positive threshold for CCNE1 gene amplification (i.e., 2.5). The detection rate is then compared with the incidence of breast cancer in the cBiPortal public database to examine whether the detection rate of this embodiment's method meets expectations. The results show that the CCNE1 gene copy number amplification rate detected by this embodiment's method is 5.24%, which is basically consistent with the public database (3.04%) and meets expectations.
[0538] As can be seen from Example 3, the gene copy number variation type detection method provided in this disclosure performs well in detecting copy number amplification of the gene to be tested.
[0539] Further reference Figure 8 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a gene copy number variation type detection device, which is similar to... Figure 2A Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0540] like Figure 8As shown, the gene copy number variation type detection device 800 of this embodiment includes: a sequencing alignment data acquisition module 801, a purity and ploidy determination module 802, an average mixed copy number determination module 803, a tumor cell copy number correction module 804, a homozygous deletion detection module 805, and a copy number amplification detection module 806. The sequencing alignment data acquisition module 801 is configured to acquire sequencing alignment data from gene sequencing of the test sample and at least one baseline sample, respectively. Each sequencing alignment data includes sequencing alignment data of a preset set of gene analysis regions, which includes the region where the gene to be tested is located. The purity and ploidy determination module 802 is configured to determine the tumor purity and tumor ploidy of the test sample based on each sequencing alignment data. The average mixed copy number determination module 803 is configured to determine the average mixed copy number of the gene to be tested in the test sample based on each sequencing alignment data. The tumor cell copy number correction module 804 is configured to determine the corrected tumor cell copy number of the gene to be tested in the test sample based on the tumor purity and tumor ploidy of the test sample and the average mixed copy number of the gene to be tested in the test sample. The homozygous deletion detection module 805 is configured to, in response to an average mixed copy number of the test gene in the test sample being less than 2, generate a detection result for homozygous deletion of the test gene in the test sample based on a comparison between the corrected tumor cell copy number of the test gene in the test sample and the maximum homozygous deletion positive copy number threshold of the test gene; and / or, the copy number amplification detection module 806 is configured to, in response to an average mixed copy number of the test gene in the test sample being greater than 2, determine the amplification judgment value of the test gene in the test sample based on the corrected tumor cell copy number of the test gene in the test sample, and generate a detection result for copy number amplification of the test gene in the test sample based on a comparison between the amplification judgment value of the test gene in the test sample and the minimum amplification positive judgment value threshold of the test gene.
[0541] In this embodiment, the specific processing and technical effects of the sequencing alignment data acquisition module 801, purity and ploidy determination module 802, average mixed copy number determination module 803, tumor cell copy number correction module 804, homozygous deletion detection module 805, and copy number amplification detection module 806 of the gene copy number variation type detection device 800 can be found in the following references. Figure 2A The relevant descriptions of steps 201, 202, 203, 204, 205, and 206 in the corresponding embodiments will not be repeated here.
[0542] In some alternative implementations, the tumor cell copy number correction module 804 may be further configured as follows:
[0543] The tumor purity and ploidy of the test sample, and the average mixed copy number of the test gene in the test sample are substituted into the formula for calculating the corrected tumor cell copy number to obtain the corrected tumor cell copy number of the test gene.
[0544] In some optional embodiments, the formula for calculating the corrected tumor cell copy number can be obtained by deriving a correction formula, wherein the correction formula is:
[0545] The formula for calculating the corrected tumor cell copy number is as follows:
[0546]
[0547] in, and These represent the tumor purity and tumor ploidy of the sample to be tested, respectively. The average mixed copy number of the gene to be tested in the sample to be tested. The corrected tumor cell copy number of the gene to be tested in the sample to be tested.
[0548] In some optional embodiments, determining the amplification judgment value of the test gene in the test sample based on the corrected tumor cell copy number of the test gene in the test sample may include:
[0549] The corrected tumor cell copy number of the gene to be tested in the test sample is determined as the amplification judgment value of the gene to be tested in the test sample, or...
[0550] The difference between the corrected tumor cell copy number of the gene to be tested in the test sample and the tumor ploidy of the test sample is determined as the amplification judgment value of the gene to be tested in the test sample, or...
[0551] The amplification judgment value of the test gene in the test sample is determined by dividing the corrected tumor cell copy number of the test gene in the test sample by the tumor ploidy of the test sample.
[0552] In some alternative implementations, the purity and ploidy determination module 802 may be further configured as follows:
[0553] Based on the sequencing alignment data of the sample to be tested, the abundance of each SNP site in the preset whole-genome SNP site set is determined.
[0554] For each preset whole-genome SNP site, the relevant gene analysis region is divided into at least one SNP analysis region according to a first preset length.
[0555] For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region;
[0556] For each SNP analysis region, the mixed copy number of the sequencing alignment data of the sample to be tested in the SNP analysis region is determined based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the SNP analysis region and the mean of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region.
[0557] Based on the sequencing alignment data of the sample to be tested, the abundance and mixed copy number of each SNP site are used to determine the tumor purity and tumor ploidy of the sample to be tested. The mixed copy number of each SNP site is the mixed copy number of the SNP analysis region in which the SNP site is located.
[0558] In some optional implementations, the average mixed copy number determination module 803 may be further configured to:
[0559] For each preset gene analysis region, the preset gene analysis region is divided into at least one gene analysis window according to the second preset length;
[0560] For each sequencing alignment data point, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window;
[0561] For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, the corresponding reference baseline sample sequencing alignment data of the gene analysis window, as well as the corrected average sequencing depth of the baseline sample and the deviation of the corrected average sequencing depth of the baseline sample in the gene analysis window are determined from the sequencing alignment data of each baseline sample.
[0562] Based on the corrected baseline sample average sequencing depth deviation of each of the gene analysis windows, a reference gene analysis window is determined in each of the gene analysis windows.
[0563] The analysis windows of each reference gene containing the gene to be tested are determined as the analysis windows of the gene to be tested.
[0564] For each gene analysis window, the mixed copy number of the gene analysis window in the test sample is determined based on the corrected average coverage depth of the gene analysis window and the corrected average sequencing depth of the baseline sample in the gene analysis window, according to the sequencing alignment data of the test sample.
[0565] The mean of the mixed copy number of each gene analysis window in the test sample is determined as the average mixed copy number of the gene in the test sample.
[0566] In some optional implementations, for each gene analysis window, determining the reference baseline sample sequencing data corresponding to the gene analysis window, the corrected average coverage depth of the baseline samples in the sequencing alignment data of each baseline sample, and the corrected average sequencing depth and the deviation of the corrected average sequencing depth of the baseline samples in the gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, may include:
[0567] For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation:
[0568] Based on the sequencing alignment data of each baseline sample, the average coverage depth after correction within the gene analysis window is used to determine the normal range of the average sequencing depth of the baseline samples.
[0569] The sequencing alignment data of each baseline sample whose corrected average coverage depth in the gene analysis window is within the normal range of the average sequencing depth of the baseline sample are determined as the reference baseline sample sequencing alignment data corresponding to the gene analysis window.
[0570] Based on the sequencing alignment data of each reference baseline sample in the gene analysis window, and the corrected average coverage depth of the gene analysis window, the corrected average sequencing depth of the baseline samples and the deviation of the corrected average sequencing depth of the baseline samples in the gene analysis window are determined.
[0571] In some optional implementations, determining the corrected average coverage depth of each sequencing alignment data point for each gene analysis window may include:
[0572] For each sequencing alignment data point, using each gene analysis window as the target analysis region, the following target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each gene analysis window. The target analysis region sequencing depth correction operation includes:
[0573] Based on the total sequencing depth of the sequencing alignment data, the sequencing depth of each base site in each target analysis region is standardized to obtain the corresponding standardized sequencing depth.
[0574] For each target analysis region, the average sequencing depth of the sequencing alignment data in that target analysis region is determined based on the normalized sequencing depth of each base site in that target analysis region.
[0575] Based on the number of probes covered by each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region;
[0576] Based on the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.
[0577] In some optional implementations, determining the corrected average coverage depth of each sequencing alignment data point for each SNP analysis region may include:
[0578] For each sequencing alignment data, the SNP analysis region is used as the target analysis region, and the sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each SNP analysis region.
[0579] In some alternative embodiments, the device 800 further includes a standardization module ( Figure 8 (not shown in the image), configured to determine the tumor purity and tumor ploidy of the test sample based on each of the said sequencing alignment data:
[0580] For each sequencing alignment data, the following base site sequencing depth normalization operation is performed: For each preset gene analysis region, the sequencing depth of each base site in the preset gene analysis region is determined; the sum of the sequencing depths of all base sites in each preset gene analysis region is determined as the total sequencing depth of the sequencing alignment data; for each base site in each preset gene analysis region, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.
[0581] In some optional implementations, the positive threshold for homozygous deletion of the gene to be tested can be predetermined through the following steps for determining the positive threshold for homozygous deletion:
[0582] A set of homozygous deletion sample data is obtained, wherein the homozygous deletion samples include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing alignment data of benign samples, and each homozygous deletion positive sample data is sequencing alignment data for homozygous deletion of the gene to be tested.
[0583] Determine the corrected tumor cell copy number for the gene to be tested for each homozygous deletion sample data;
[0584] For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, the following operation is performed to determine the positive criteria for homozygous deletion: based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the positive concordance rate corresponding to the candidate homozygous deletion copy number is determined; and / or, based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each homozygous deletion sample data, the positive predictive value corresponding to the candidate homozygous deletion copy number is determined;
[0585] The positive concordance rate and / or positive prediction value corresponding to each candidate homozygous deletion copy number are used to determine the positive judgment threshold for homozygous deletion of the gene to be tested.
[0586] In some optional implementations, obtaining the homozygous missing sample dataset may include:
[0587] Obtain sequencing alignment data for N pairs of 100% tumor cell lines and their paired cell lines, where N is a positive integer;
[0588] Based on the N pairs of sequencing alignment data, at least one homozygous deletion positive sample data for the gene to be tested is generated, wherein the tumor purity of each homozygous deletion positive sample data is the corresponding labeled tumor purity.
[0589] Obtain sequencing alignment data of M benign samples, and generate at least one negative sample data based on the sequencing alignment data of the M benign samples.
[0590] In some alternative implementations, determining the corrected tumor cell copy number for each homozygous deletion sample data for the gene to be tested may include:
[0591] For each homozygous deletion sample, determine the tumor purity and tumor ploidy corresponding to the homozygous deletion sample, as well as the average mixed copy number for the gene to be tested. Based on the tumor purity and tumor ploidy corresponding to the homozygous deletion sample, as well as the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the homozygous deletion sample.
[0592] In some optional implementations, determining the positive concordance rate corresponding to the candidate homozygous deletion copy number based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each of the homozygous deletion positive sample data may include:
[0593] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number are counted.
[0594] The positive concordance rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number.
[0595] In some optional implementations, determining the positive predictive value corresponding to the candidate homozygous deletion copy number based on the comparison results between the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each of the homozygous deletion sample data may include:
[0596] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted.
[0597] Based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each negative sample data, the number of false positives corresponding to the candidate homozygous deletion copy number is counted; and
[0598] The positive predicted value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false positives corresponding to the candidate homozygous deletion copy number.
[0599] In some optional implementations, generating at least one negative sample data based on the sequencing alignment data of the M benign samples may include:
[0600] For each benign sample's sequencing alignment data, noise of varying degrees is added to the sequencing depth of the base sites in the benign sample's sequencing alignment data to obtain at least one negative sample data corresponding to the benign sample's sequencing alignment data.
[0601] In some optional implementations, the amplification positivity threshold for the gene to be tested can be predetermined through the following amplification positivity threshold determination steps:
[0602] A set of amplified sample data is obtained, which includes negative sample data and amplified positive sample data for the gene to be tested. Each negative sample data is obtained based on sequencing alignment data of benign samples, and each amplified positive sample data is sequencing alignment data of copy number amplification of the gene to be tested.
[0603] Determine the corrected tumor cell copy number for each amplified sample data;
[0604] The amplification judgment value for the gene to be tested is determined based on the corrected tumor cell copy number of each amplified sample data.
[0605] For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, the following amplification positive judgment basis determination operation is performed: based on the comparison result of the amplification judgment value of each amplification positive sample data for the gene to be tested with the candidate amplification positive judgment value, the positive concordance rate corresponding to the candidate amplification positive judgment value is determined; and / or, based on the comparison result of the amplification judgment value of each amplification sample data for the gene to be tested with the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined;
[0606] The amplification positivity threshold of the gene to be tested is determined based on the positive concordance rate and / or positive prediction value corresponding to each candidate amplification positivity judgment value.
[0607] In some optional implementations, obtaining the amplified sample dataset may include:
[0608] Obtain sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines, where I is a positive integer;
[0609] Based on the I-pair sequencing alignment data, at least one amplified positive sample data for the gene to be tested is generated, wherein the tumor purity of each amplified positive sample data is the corresponding labeled tumor purity.
[0610] Obtain sequencing alignment data of J benign samples, and generate at least one negative sample data based on the sequencing alignment data of the J benign samples.
[0611] In some alternative implementations, determining the corrected tumor cell copy number for each amplified sample data may include:
[0612] For each amplified sample data, determine the tumor purity and tumor ploidy corresponding to the amplified sample data, as well as the average mixed copy number for the gene to be tested. Based on the tumor purity and tumor ploidy corresponding to the amplified sample data, as well as the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the amplified sample data.
[0613] In some optional embodiments, determining the positive concordance rate corresponding to the candidate amplification positive judgment value based on the comparison result between the amplification judgment value of each of the amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value may include:
[0614] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification positive judgment value of each amplification positive sample data, the number of true positives and false negatives corresponding to the candidate amplification positive judgment value are counted.
[0615] The positive concordance rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and false negatives corresponding to the candidate amplification positive judgment value.
[0616] In some optional embodiments, determining the positive prediction value corresponding to the candidate amplification positive judgment value based on the comparison result between the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value may include:
[0617] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification judgment value of each amplification positive sample data, the number of true positives corresponding to the candidate amplification positive judgment value is counted.
[0618] Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification positive judgment value of each negative sample data, the number of false positives corresponding to the candidate amplification positive judgment value is counted; and
[0619] Based on the number of true positives and false positives corresponding to the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined.
[0620] In some optional implementations, determining the amplification judgment value of the corresponding amplified sample data for the gene to be tested based on the corrected tumor cell copy number of each amplified sample data may include:
[0621] The corrected tumor cell copy number for the gene to be tested in each amplified sample is determined as the amplification judgment value for that amplified sample data, or...
[0622] The difference between the corrected tumor cell copy number for the target gene in each amplified sample and the tumor ploidy of that amplified sample is determined as the amplification judgment value for that amplified sample, or...
[0623] The amplification judgment value of each amplified sample data is determined by dividing the corrected tumor cell copy number for the gene to be tested by the tumor ploidy ratio of that amplified sample data.
[0624] It should be noted that the implementation details and technical effects of each module in the gene copy number variation type detection device provided in the embodiments of this disclosure can be referred to the descriptions of other embodiments in this disclosure, and will not be repeated here.
[0625] The following is for reference. Figure 9It shows a schematic diagram of the structure of a computer system 900 suitable for implementing the electronic device of the present disclosure. Figure 9 The computer system 900 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0626] like Figure 9 As shown, the computer system 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the computer system 900. The processing device 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0627] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows computer system 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 A computer system 900 with various electronic devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0628] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0629] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0630] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0631] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the following functions: Figure 2A The illustrated embodiments and their alternative implementations demonstrate a method for detecting gene copy number variation types.
[0632] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Python, Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0633] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0634] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not necessarily limiting in certain circumstances; for example, a sequencing alignment data acquisition module can also be described as "a module for acquiring sequencing alignment data of the test sample and at least one baseline sample, respectively."
[0635] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A method for detecting gene copy number variation types, comprising: Sequencing alignment data of gene sequencing of the sample to be tested and at least one baseline sample are obtained, wherein each sequencing alignment data includes sequencing alignment data of a preset set of gene analysis regions to be tested, and the preset set of gene analysis regions to be tested includes the region where the gene to be tested is located. Based on the sequencing alignment data, the tumor purity and tumor ploidy of the sample to be tested are determined; Based on the sequencing alignment data, the average mixed copy number of the gene to be tested in the sample to be tested is determined; The tumor purity and ploidy of the test sample, and the average mixed copy number of the test gene in the test sample, are substituted into the formula for calculating the corrected tumor cell copy number to obtain the corrected tumor cell copy number of the test gene. The formula for calculating the corrected tumor cell copy number is obtained by deriving the correction formula, which is: The formula for calculating the corrected tumor cell copy number is as follows: in, and These represent the tumor purity and tumor ploidy of the sample to be tested, respectively. The average mixed copy number of the gene to be tested in the sample to be tested. The corrected tumor cell copy number of the gene to be tested in the sample to be tested; In response to the fact that the average mixed copy number of the gene to be tested in the test sample is less than 2, the test result of the test sample for the homozygous deletion of the gene to be tested is generated based on the comparison result of the corrected tumor cell copy number of the gene to be tested in the test sample and the positive judgment threshold of the homozygous deletion of the gene to be tested. And / or, In response to the average mixed copy number of the test gene in the test sample being greater than 2, the difference between the corrected tumor cell copy number of the test gene in the test sample and the tumor ploidy of the test sample is determined as the amplification judgment value of the test gene in the test sample; or, the ratio of the corrected tumor cell copy number of the test gene in the test sample to the tumor ploidy of the test sample is determined as the amplification judgment value of the test gene in the test sample. Based on the comparison result between the amplification judgment value of the test gene in the test sample and the amplification positive judgment threshold of the test gene, a detection result of copy number amplification of the test gene in the test sample is generated.
2. The method according to claim 1, wherein, The determination of tumor purity and tumor ploidy of the sample to be tested based on the sequencing alignment data includes: Based on the sequencing alignment data of the sample to be tested, the abundance of each SNP site in the preset whole-genome SNP site set is determined. For each preset whole-genome SNP site, the relevant gene analysis region is divided into at least one SNP analysis region according to a first preset length. For each sequencing alignment data, determine the corrected average coverage depth of the sequencing alignment data for each SNP analysis region; For each SNP analysis region, the mixed copy number of the sequencing alignment data of the sample to be tested in the SNP analysis region is determined based on the corrected average coverage depth of the sequencing alignment data of the sample to be tested in the SNP analysis region and the mean of the corrected average coverage depth of the sequencing alignment data of each baseline sample in the SNP analysis region. Based on the sequencing alignment data of the sample to be tested, the abundance and mixed copy number of each SNP site are used to determine the tumor purity and tumor ploidy of the sample to be tested. The mixed copy number of each SNP site is the mixed copy number of the SNP analysis region in which the SNP site is located.
3. The method according to claim 2, wherein, The determination of the average mixed copy number of the gene to be tested in the sample to be tested based on the sequencing alignment data includes: For each preset gene analysis region, the preset gene analysis region is divided into at least one gene analysis window according to the second preset length; For each sequencing alignment data point, determine the corrected average coverage depth of the sequencing alignment data for each gene analysis window; For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in the gene analysis window, the corresponding reference baseline sample sequencing alignment data of the gene analysis window, as well as the corrected average sequencing depth of the baseline sample and the deviation of the corrected average sequencing depth of the baseline sample in the gene analysis window are determined from the sequencing alignment data of each baseline sample. Based on the corrected baseline sample average sequencing depth deviation of each of the gene analysis windows, a reference gene analysis window is determined in each of the gene analysis windows. The analysis windows of each reference gene containing the gene to be tested are determined as the analysis windows of the gene to be tested. For each gene analysis window, the mixed copy number of the gene analysis window in the test sample is determined based on the corrected average coverage depth of the gene analysis window and the corrected average sequencing depth of the baseline sample in the gene analysis window, according to the sequencing alignment data of the test sample. The mean of the mixed copy number of each gene analysis window in the test sample is determined as the average mixed copy number of the gene in the test sample.
4. The method according to claim 3, wherein, For each gene analysis window, based on the corrected average coverage depth of the sequencing alignment data of each baseline sample in that gene analysis window, the corresponding reference baseline sample sequencing alignment data for that gene analysis window, as well as the corrected average sequencing depth of the baseline samples and the deviation of the corrected average sequencing depth of the baseline samples in that gene analysis window, are determined from the sequencing alignment data of each baseline sample, including: For each gene analysis window, perform the following baseline sample gene analysis window depth correction operation: Based on the sequencing alignment data of each baseline sample, the average coverage depth after correction within the gene analysis window is used to determine the normal range of the average sequencing depth of the baseline samples. The sequencing alignment data of each baseline sample whose corrected average coverage depth in the gene analysis window is within the normal range of the average sequencing depth of the baseline sample are determined as the reference baseline sample sequencing alignment data corresponding to the gene analysis window. Based on the sequencing alignment data of each reference baseline sample in the gene analysis window, and the corrected average coverage depth of the gene analysis window, the corrected average sequencing depth of the baseline samples and the deviation of the corrected average sequencing depth of the baseline samples in the gene analysis window are determined.
5. The method according to claim 3, wherein, For each sequencing alignment data point, determining the corrected average coverage depth for each gene analysis window includes: For each sequencing alignment data point, using each gene analysis window as the target analysis region, the following target analysis region sequencing depth correction operation is performed to obtain the corrected average coverage depth of each gene analysis window. The target analysis region sequencing depth correction operation includes: Based on the total sequencing depth of the sequencing alignment data, the sequencing depth of each base site in each target analysis region is standardized to obtain the corresponding standardized sequencing depth. For each target analysis region, the average sequencing depth of the sequencing alignment data in that target analysis region is determined based on the normalized sequencing depth of each base site in that target analysis region. Based on the number of probes covered by each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region; Based on the GC content ratio of the sequencing alignment data in each target analysis region, a local weighted regression correction is performed on the average coverage depth of the sequencing alignment data in each target analysis region to obtain the corrected average coverage depth of the corresponding target analysis region.
6. The method according to claim 5, wherein, For each sequencing alignment data point, determining the corrected average coverage depth for each SNP analysis region includes: For each sequencing alignment data, the SNP analysis region is used as the target analysis region, and the sequencing depth correction operation of the target analysis region is performed to obtain the corrected average coverage depth of each SNP analysis region.
7. The method according to claim 1, wherein, Before determining the tumor purity and tumor ploidy of the sample to be tested based on the sequencing alignment data, the method further includes: For each sequencing alignment data, the following base site sequencing depth normalization operation is performed: For each preset gene analysis region, the sequencing depth of each base site in the preset gene analysis region is determined; the sum of the sequencing depths of all base sites in each preset gene analysis region is determined as the total sequencing depth of the sequencing alignment data; for each base site in each preset gene analysis region, the sequencing depth of the base site is normalized based on the total sequencing depth of the sequencing alignment data.
8. The method according to claim 1, wherein, The positive threshold for homozygous deletion of the gene to be tested is predetermined through the following steps for determining the positive threshold for homozygous deletion: A set of homozygous deletion sample data is obtained, wherein the homozygous deletion samples include negative sample data and homozygous deletion positive sample data for the gene to be tested, wherein each negative sample data is obtained based on sequencing alignment data of benign samples, and each homozygous deletion positive sample data is sequencing alignment data for homozygous deletion of the gene to be tested. Determine the corrected tumor cell copy number for the gene to be tested for each homozygous deletion sample data; For each candidate homozygous deletion copy number in the candidate homozygous deletion copy number set, the following operation is performed to determine the positive criteria for homozygous deletion: based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the positive concordance rate corresponding to the candidate homozygous deletion copy number is determined; and / or, based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each homozygous deletion sample data, the positive predictive value corresponding to the candidate homozygous deletion copy number is determined; The positive concordance rate and / or positive prediction value corresponding to each candidate homozygous deletion copy number are used to determine the positive judgment threshold for homozygous deletion of the gene to be tested.
9. The method according to claim 8, wherein, The acquisition of the homozygous missing sample data set includes: Obtain sequencing alignment data for N pairs of 100% tumor cell lines and their paired cell lines, where N is a positive integer; Based on the N pairs of sequencing alignment data, at least one homozygous deletion positive sample data for the gene to be tested is generated, wherein the tumor purity of each homozygous deletion positive sample data is the corresponding labeled tumor purity. Obtain sequencing alignment data of M benign samples, and generate at least one negative sample data based on the sequencing alignment data of the M benign samples.
10. The method according to claim 8, wherein, The determination of the corrected tumor cell copy number for each homozygous deletion sample data for the gene to be tested includes: For each homozygous deletion sample, determine the tumor purity and tumor ploidy corresponding to the homozygous deletion sample, as well as the average mixed copy number for the gene to be tested. Based on the tumor purity and tumor ploidy corresponding to the homozygous deletion sample, as well as the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the homozygous deletion sample.
11. The method according to claim 8, wherein, The step of determining the positive concordance rate corresponding to the candidate homozygous deletion copy number by comparing the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number based on the data from each homozygous deletion positive sample includes: Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number are counted. The positive concordance rate corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false negatives corresponding to the candidate homozygous deletion copy number.
12. The method according to claim 8, wherein, The step of determining the positive predictive value corresponding to the candidate homozygous deletion copy number by comparing the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number based on the homozygous deletion sample data includes: Based on the comparison results of the corrected tumor cell copy number of the gene to be tested and the candidate homozygous deletion copy number for each homozygous deletion positive sample data, the number of true positives corresponding to the candidate homozygous deletion copy number is counted. Based on the comparison results of the corrected tumor cell copy number of the gene to be tested with the candidate homozygous deletion copy number for each negative sample data, the number of false positives corresponding to the candidate homozygous deletion copy number is counted; and The positive predicted value corresponding to the candidate homozygous deletion copy number is determined based on the number of true positives and false positives corresponding to the candidate homozygous deletion copy number.
13. The method according to claim 9, wherein, The generation of at least one negative sample data based on the sequencing alignment data of the M benign samples includes: For each benign sample's sequencing alignment data, noise of varying degrees is added to the sequencing depth of the base sites in the benign sample's sequencing alignment data to obtain at least one negative sample data corresponding to the benign sample's sequencing alignment data.
14. The method according to claim 1, wherein, The amplification positivity threshold for the gene to be tested is predetermined through the following amplification positivity threshold determination steps: A set of amplified sample data is obtained, which includes negative sample data and amplified positive sample data for the gene to be tested. Each negative sample data is obtained based on sequencing alignment data of benign samples, and each amplified positive sample data is sequencing alignment data of copy number amplification of the gene to be tested. Determine the corrected tumor cell copy number for each amplified sample data; The amplification judgment value for the gene to be tested is determined based on the corrected tumor cell copy number of each amplified sample data. For each candidate amplification positive judgment value in the candidate amplification positive judgment value set, the following amplification positive judgment basis determination operation is performed: based on the comparison results of the amplification judgment value of each amplification positive sample data for the gene to be tested with the candidate amplification positive judgment value, the positive concordance rate corresponding to the candidate amplification positive judgment value is determined; and / or, based on the comparison results of the amplification judgment value of each amplification sample data for the gene to be tested with the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined; The amplification positivity threshold of the gene to be tested is determined based on the positive concordance rate and / or positive prediction value corresponding to each candidate amplification positivity judgment value.
15. The method according to claim 14, wherein, The acquisition of the amplified sample data set includes: Obtain sequencing alignment data of I pairs of 100% tumor cell lines and their paired cell lines, where I is a positive integer; Based on the I-pair sequencing alignment data, at least one amplified positive sample data for the gene to be tested is generated, wherein the tumor purity of each amplified positive sample data is the corresponding labeled tumor purity. Obtain sequencing alignment data of J benign samples, and generate at least one negative sample data based on the sequencing alignment data of the J benign samples.
16. The method of claim 14, wherein, Determining the corrected tumor cell copy number for each amplified sample data includes: For each amplified sample data, determine the tumor purity and tumor ploidy corresponding to the amplified sample data, as well as the average mixed copy number for the gene to be tested. Based on the tumor purity and tumor ploidy corresponding to the amplified sample data, as well as the average mixed copy number for the gene to be tested, determine the corrected tumor cell copy number for the gene to be tested in the amplified sample data.
17. The method according to claim 14, wherein, The step of determining the positive concordance rate corresponding to the candidate amplification positive judgment value based on the comparison results of the amplification judgment value of each of the amplification positive sample data for the gene to be tested and the candidate amplification positive judgment value includes: Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification positive judgment value of each amplification positive sample data, the number of true positives and false negatives corresponding to the candidate amplification positive judgment value are counted. The positive concordance rate corresponding to the candidate amplification positive judgment value is determined based on the number of true positives and false negatives corresponding to the candidate amplification positive judgment value.
18. The method according to claim 14, wherein, The step of determining the positive prediction value corresponding to the candidate amplification positive judgment value based on the comparison results between the amplification judgment value of each amplification sample data for the gene to be tested and the candidate amplification positive judgment value includes: Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification judgment value of each amplification positive sample data, the number of true positives corresponding to the candidate amplification positive judgment value is counted. Based on the comparison results of the amplification judgment value of the gene to be tested and the candidate amplification positive judgment value of each negative sample data, the number of false positives corresponding to the candidate amplification positive judgment value is counted; and Based on the number of true positives and false positives corresponding to the candidate amplification positive judgment value, the positive prediction value corresponding to the candidate amplification positive judgment value is determined.
19. The method of claim 14, wherein, The step of determining the amplification judgment value for the gene to be tested for the corresponding amplified sample data based on the corrected tumor cell copy number of each amplified sample data includes: The corrected tumor cell copy number for the gene to be tested in each amplified sample is determined as the amplification judgment value for that amplified sample data, or... The difference between the corrected tumor cell copy number for the target gene in each amplified sample and the tumor ploidy of that amplified sample is determined as the amplification judgment value for that amplified sample, or... The amplification judgment value of each amplified sample data is determined by dividing the corrected tumor cell copy number for the gene to be tested by the tumor ploidy ratio of that amplified sample data.
20. A device for detecting gene copy number variation types, comprising: The sequencing alignment data acquisition module is configured to acquire sequencing alignment data of gene sequencing performed on the sample to be tested and at least one baseline sample, wherein each sequencing alignment data includes sequencing alignment data of a preset set of gene analysis regions to be tested, and the preset set of gene analysis regions to be tested includes the region where the gene to be tested is located. The purity and ploidy determination module is configured to determine the tumor purity and tumor ploidy of the test sample based on each of the sequencing alignment data. The average mixed copy number determination module is configured to determine the average mixed copy number of the gene to be tested in the sample to be tested based on each of the sequencing alignment data. The tumor cell copy number correction module is configured to substitute the tumor purity and ploidy of the test sample, and the average mixed copy number of the test gene in the test sample, into the corrected tumor cell copy number calculation formula to obtain the corrected tumor cell copy number of the test gene. The corrected tumor cell copy number calculation formula is obtained by deriving the correction formula, wherein the correction formula is: The formula for calculating the corrected tumor cell copy number is as follows: in, and These represent the tumor purity and tumor ploidy of the sample to be tested, respectively. The average mixed copy number of the gene to be tested in the sample to be tested. The corrected tumor cell copy number of the gene to be tested in the sample to be tested; A homozygous deletion detection module is configured to, in response to a mean mixed copy number of the test gene in the test sample being less than 2, generate a detection result for homozygous deletion of the test gene in the test sample based on a comparison between the corrected tumor cell copy number of the test gene in the test sample and a positive threshold for homozygous deletion of the test gene; and / or, The copy number amplification detection module is configured to, in response to an average mixed copy number of the test gene in the test sample being greater than 2, determine the amplification judgment value of the test gene in the test sample as the difference between the corrected tumor cell copy number of the test gene in the test sample and the tumor ploidy of the test sample, or determine the amplification judgment value of the test gene in the test sample as the ratio of the corrected tumor cell copy number of the test gene in the test sample to the tumor ploidy of the test sample, and generate a detection result of copy number amplification of the test gene in the test sample based on a comparison between the amplification judgment value of the test gene in the test sample and the amplification positive judgment threshold of the test gene.
21. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1-19.
22. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by one or more processors, it implements the method as described in any one of claims 1-19.
23. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method as described in any one of claims 1-19.
Citation Information
Patent Citations
Single sample allele copy number variation detection method, probe set and kit
CN113889187A